The State of Play

A living index of AI adoption across industries — where established practice meets the bleeding edge
UPDATED DAILY
← ✨ Personal Effectiveness

Skill acquisition & deliberate practice support

LEADING EDGE— Steady

181 evidence items

AI that supports individuals in acquiring new skills through structured practice, feedback, and progression tracking. Includes practice exercise generation and personalised feedback; distinct from L&D recommendations which suggest resources rather than structuring practice.

Overview

AI-powered deliberate practice works in controlled settings but September 2026 evidence confirms the performance-learning paradox remains a systemic design challenge dependent on pedagogical guardrails and behavioral adoption. Efficacy research remains robust: Google's Gemini study (Sierra Leone, Italy) showed +0.26 to +0.38 SD gains with learning-science grounding; Wharton's adaptive sequencing RCT demonstrated +0.15 SD improvements; Harvard physics RCT found AI tutoring d=0.73-1.3. Yet unguarded AI enables visible short-term task completion while degrading durable skill acquisition: Liu et al.'s 1,222-participant RCT documents unguided AI assistance reduces persistence within ~10 minutes; World Bank analysis of 26,000 Chinese students shows autonomous AI for homework raises practice 18% but reduces exam performance 20% (1.4 SD effect). The mechanism is identified (cognitive offloading, illusion of competence, effort reduction) and well-documented. Recent evidence (Aug-Sep 2026) clarifies critical boundaries: (1) Design determines outcome: early September NBER RCT (6,000+ students, 18 Tennessee schools) confirms AI tutoring effectiveness depends entirely on guardrails and student willingness—without compulsion, 96% try the tool but median usage drops to 1/3 of practice days; paired evidence documents 64-67% of students with high AI-assisted practice scores later fail unaided exams (illusion of competence). (2) Substitution vs. complement: Washington University study of competitive programming shows AI substitute-mode (AI does the work) harms skill while complement-mode (AI supports the learner) preserves performance on proctored evaluation. (3) Workforce deskilling risk: analysis of junior engineers and workers shows 39% report AI weakened their skill sets, with weakened debugging and conceptual understanding—evidence that AI can actively undermine rather than support skill development when substituting for practice. (4) Vendor strategy persists: Google's Sept 2026 rollout of 30+ Classroom AI tools and Khan-Google partnership signal ecosystem maturation but 86% global student adoption coexists with <10% institutional governance, creating an adoption-without-policy crisis. El Salvador's large-scale deployment (1,198 students, 171 schools, World Bank backing) demonstrates named-institution scale, validating deployment feasibility in Global South while North American institutional confidence remains fragile. Properly designed systems (adaptive support, sequenced scaffolding, effort preservation, Socratic structure) can preserve learning but require pedagogical expertise most organizations lack. Duolingo (58.7M DAU +23% Q2 2026, 12.7M paid subscribers) and Khan Academy (700K+ users, 380+ U.S. districts) prove platform infrastructure scale; selective adoption and unresolved skill transfer determine tier positioning rather than capability alone.

Current Landscape

August 2026 market dynamics show vendor strategy alignment on pedagogical design combined with persistent adoption and transfer barriers. Consumer platforms (Duolingo 58.7M DAU +23% YoY with 12.7M subscribers, Q2 2026; Khan Academy 700K+ across 380+ U.S. districts) demonstrate infrastructure scale but limited engagement: only 15% of Khanmigo users actively engage despite access. Two-year Tennessee RCT (18 schools, N~6,900 student-terms) documents critical engagement barrier: 96% of students tried Khanmigo at least once, but median engagement was only 1/3 of practice days and 17% of error-correction sessions; gains matched Khan Academy without AI, confirming that practice structure—not AI capability—drives achievement. Stanford SCALE Initiative synthesis confirms engagement crisis: AI-led tutoring shows 40–47% non-use rates; AI-assisting human tutors shows emerging evidence of equal/better effectiveness, establishing human involvement as non-negotiable for scaled deployment. Employer skill demand shows rapid acceleration: American University survey (2024-2026) documents employers requesting AI skills increased from 11.6% to 42.6% (285% increase), with 80%+ student adoption within 6 months, signaling shift from supplementary to baseline competency requirement. Pedagogical pivot evident: Google, OpenAI, and Anthropic announced Q3 2026 strategy shift from quick-answer efficiency to scaffolded learning with Socratic questioning, step-by-step progression, and guided reasoning—with Google-Khan partnership releasing interactive Gemini diagrams for August 2026 back-to-school, demonstrating product-level maturation in response to engagement barriers. Duolingo implementation reveals adoption gap: structured deliberate practice using AI tutors (2–3 scheduled sessions/week with repair drills and generalization exercises) generates skill gains; unfocused exploration yields minimal transfer. Institutional adoption barriers remain unchanged: (1) teacher involvement requirement—Brookings research confirms gains only when educators remain present, not autonomous AI; (2) pedagogical calibration gap—Allen Institute TutorMoments framework shows LLMs default to over-helping rather than productive struggle, requiring explicit design constraints; (3) skill transfer disconnect—85% of enterprise learners cannot apply AI training to work contexts, and medical training research documents "never-skilling" risk when AI introduced before foundational mastery; (4) population-level evidence from 26,811 students over 30 months shows homework gains (+18%) reverse to exam losses (-20%), establishing performance-learning paradox as systemic risk requiring design intervention. Global deployment evidence: Pakistan's Sindh Department trained 3,503 teachers (91% pass rate) in Khanmigo, with AI lesson-planning adoption rising from 0% to 79% and assessment design ease rising from 51% to 92%, demonstrating feasibility of teacher-scale professional skill development in Global South contexts. Harvard/MIT validation demonstrates solution exists: reinforcement learning-based adaptive tutoring (N=1,280) that calibrates help-level by learner state preserves skill retention (0.70 effect) while maintaining accuracy gains, showing mechanism for balancing performance and learning is technically feasible. Market segmentation persists: teacher-facing tools (MagicSchool 6M, Brisk 1M) gained 2024-25 market share from student tutors; K-12 districts prioritize teacher augmentation over student-facing AI. Regional variation shows LATAM higher education at 92% student and 79% faculty AI engagement but U.S. K-12 adoption remains constrained by integration friction, infrastructure gaps (40% of primary schools globally lack internet), and absence of rigorous independent evaluation in most rollouts. Critical barrier remains institutional design: pedagogical expertise in skill-preservation mechanisms, teacher training burden, environmental alignment between practice and deployment contexts, and human accountability for engagement determine outcome, not technology capability or platform access alone.

Tier History

ResearchMar-2023 → Jul-2023
Bleeding EdgeJul-2023 → Oct-2024
Leading EdgeOct-2024 → present
Open on full timeline →

Evidence (181)

— Brandon Hall Group analysis finds only 17% of organisations reach deliberate-practice maturity; 39% lack foundational infrastructure (personalised pathways, gap analytics, scaled access)—shows capability and structure, not technology, are the binding constraints at scale.

— Workera survey of 1,000 U.S. professionals quantifies adoption–practice inversion: 67.8% use AI tools regularly but 56.4% receive no work time for upskilling—evidence of 'adoption-without-policy crisis' that the practice's landscape already identifies as systemic.

— Scoping review of 153 GenAI medical education reports shows encouraging short-term signals for history-taking and clinical reasoning, but only 13 tracked retention and zero documented patient-level outcomes—confirming the practice's transfer and durability challenge.

— Narrative review of 20 studies identifies AI maths tutoring benefits only when valid learner modelling combines with adaptive tasks, calibrated scaffolding and teacher mediation—confirms increased platform interaction does not guarantee conceptual understanding or transfer.

— Mixed-methods study of 302 Turkish primary maths teachers shows moderate positive perceptions masked by serious concerns: data privacy, reliability and 'students' possible over-reliance on ready-made responses'—evidence of cognitive offloading mechanism identified in the practice.

176 more · latest 2026-09-18 →

— Multiple surveys (Talogy 207 HR leaders, PwC, Gartner 12,004 employees) document erosion of entry-level roles as deliberate-practice ground: 78% concerned about skill loss, junior roles most AI-exposed are seven times more likely to demand senior capabilities—validates practice's deskilling risk.

— PRISMA systematic review of 73 EFL studies (2015–Feb 2026) identifies AI chatbots supporting willingness to communicate via private practice spaces and personalised feedback—rare positive evidence on deliberate-practice efficacy in language learning.

— Wharton analysis citing Microsoft Research (319 knowledge workers), BetterUp (1,150 desk workers) and MIT Media Lab documents how heavy AI use erodes independent thinking—mechanism explaining why deliberate-practice support can actively undermine learning if poorly designed.

— Independent Instruction Partners evaluation of 20 AI learning tools (16 deep dives) found general-purpose chatbots reduce effortful thinking and leave learners 'isolating'; only purpose-built tutors showed promise—confirms tool design determines whether AI supports or undermines skill acquisition.

When the Tutor Leaves | EuraStudyResearch Paper

— Secondary analysis of RCT with 2,899 student sessions showing 64–67% of students with high AI-assisted practice scores later scored below 50% on unaided exams—critical evidence of assisted-performance illusion in AI-supported deliberate practice.

— NBER RCT with 6,000+ students across 18 Tennessee schools shows AI tutoring effectiveness depends entirely on design, guardrails, and student willingness; without compulsion, students largely avoid the tutor despite availability.

— Compilation of adoption metrics from primary sources (RAND, UNESCO, Gallup, HolonIQ) showing 86% global student AI use, 53% K-12 teacher use, +42% market CAGR, and <10% institutional guidance—signals mainstream adoption outpacing governance.

— Case study of El Salvador's national AI tutoring expansion with World Bank backing: 1,198 volunteer students at 171 schools scored above national average on PISA for Schools; expansion planned to 5,000+ schools and 1M+ students over 18 months.

— Washington University empirical study of competitive programming distinguishing AI substitute (harmful) vs. complement (beneficial) modes, with proctored evaluation gates separating skill preservation from scaffolding dependence in real-world elite skill contexts.

— Major vendor (Google) general availability of 30+ integrated AI tutoring tools across Classroom, Gemini, Search, and Chromebook—signals ecosystem maturity and enables broad institutional deployment with existing district infrastructure.

— Critical assessment of New Mexico's AI tutoring mandate documenting data-privacy concerns, scoring inaccuracy, and acknowledged lack of efficacy evidence—important negative signal on adoption barriers and implementation risks in governance contexts.

— Documents the 'deskilling problem': 39% of workers report AI weakened their skill set; junior engineers with AI access show measurably weaker code comprehension and debugging despite working faster—revealing how AI substitution harms skill development.

— Duolingo Q2 2026: DAU +23% to 58.7M, paid subscribers +17% to 12.7M, with CEO guidance for sustained >20% DAU growth—confirming platform adoption breadth and sustained engagement with AI-powered language practice mechanics at consumer scale.

— Google and Khan Academy launched interactive Gemini-powered diagrams for Khanmigo (August 2026): students manipulate visual elements, receive contextual feedback; early signals show increased engagement and improved answer correctness—representing product evolution in response to identified barriers.

— Pakistan's Sindh Department trained 3,503 teachers (91% pass rate) in Khanmigo; AI lesson-planning adoption rose 0% to 79%, assessment design ease 51% to 92%, with teachers saving 5+ hours/week—demonstrating professional skill acquisition deployment at scale in Global South context.

— Stanford SCALE Initiative authoritative synthesis: AI-led tutoring shows 40–47% non-use rates and insufficient evidence base; AI-assisted human tutors shows emerging promise; high-impact tutoring requires human relationships and structured engagement.

— SEC Form 8-K disclosed Duolingo's 27.4% YoY DAU growth on Aug 17, 2026, confirming sustained user engagement with AI-powered language practice despite industry disruption narratives; primary regulatory source at scale.

— Quasi-experimental study (13,037 students) comparing AI feedback workflows: 'Enacted Feedback' (structured student dialogue) achieved 26.2% uptake vs 14.1% directed and 0.1% self-directed—demonstrating workflow design as critical to feedback utility in deliberate practice.

— Two-year RCT across 18 Tennessee schools (N~6,900 student-terms) found Khanmigo gains (0.06–0.08 SD/year) matched Khan Academy alone; 96% trial rate but median engagement only 1/3 of days—identifying engagement as binding constraint to skill acquisition deployment.

— Large RCT (6,145 students, Hamilton County Schools) shows AI tutoring value depends on mastery-based workflow structure: AI improved post-error accuracy (+8.5pp) and recovery time, with 3.2pp delayed-test gains when embedded in workflow—design determines outcome.

— Peer-reviewed research documents never-skilling risk in medical training; proposes three-stage approach (establish competency without AI, teach critical evaluation, supervised integration) to preserve deliberate practice in clinical skill acquisition.

— Implementation guide for structuring deliberate practice routines with Duolingo Max: 2–3 Roleplay sessions/week, 5–15 minute sessions, turning wrong answers into repair drills via reconstruction and generalization. Identifies adoption barrier: gap between exploration and disciplined repetition determines skill gains.

— Google, OpenAI, and Anthropic pivoting learning products from quick answers to scaffolded Socratic approaches: asking questions, providing hints, breaking content into steps. Represents strategic shift by major AI vendors away from efficiency gains toward active learning and skill preservation.

— Physicians Bajaj and Sakran propose pedagogical interventions grounded in Bjork's desirable difficulties: require independent reasoning before AI consultation, periodic no-AI cases (parallel to FAA pilot currency), and sequential ordering (unaided diagnosis first) to preserve deliberate practice in clinical training.

— Allen Institute TutorMoments framework (replay-based evaluation on math tutoring transcripts): LLMs when instructed to 'tutor well' over-help (excessive scaffolding, rarely push for reasoning), diverging from pedagogically sound tutoring; identifies calibration gap as critical design challenge for AI-supported skill development.

— American University survey (N=483, tracked 2024–2026): employers requesting AI skills rose 11.6% to 42.6% (285% increase); student adoption 80%+ within 6 months; shows rapid shift from supplementary to baseline skill requirement, with growing demand for coding/technical instruction.

— Harvard/MIT reinforcement learning study (N=1,280) shows adaptive tutoring that calibrates help-level by learner state and confidence preserves skill retention while improving accuracy, validating mechanism to balance immediate performance with long-term learning.

— AWS Developer Advocate documents deliberate practice methodology using AI assistants: goal-based learning, spec-driven development, strategic collaboration. Applies Pearson & Gallagher's 'I do, we do, you do' (gradual release of responsibility) with zone-of-proximal-development scaffolding.

— Q1 2026 deployment metrics: 56.5M DAU (+21% YoY), 20,500 AI-generated course units per quarter (10× acceleration), new speaking-focused features (spoken tokens, speaking adventures); demonstrates sustained deliberate practice investment.

— AEFP policy synthesis of 20 high-quality causal studies: tools prompting reasoning and active engagement reinforce cognitive work for deeper learning; systems generating complete answers reduce cognitive work, directly undermining skill acquisition.

— Meta-analysis of 35 experimental studies: knowledge acquisition improved with AI assistance, but higher-order thinking and critical analysis showed smaller/reversed gains; demonstrates cognitive domain determines efficacy, not tool capability alone.

— 454-student semester-long RCT: accurate, grounded, popular RAG chatbot produced zero learning gains despite high student satisfaction, separating perceived helpfulness from efficacy in skill acquisition.

— RCT with 90 learners: personalized feedback yielded highest long-term mastery but temporarily reduced next-attempt performance due to cognitive load, demonstrating temporal trade-off requiring dynamic assessment of learner state.

— Deep behavioral analysis documenting Duolingo's deliberate practice engagement mechanics: 34.7% DAU/MAU ratio, 10M+ users with 1+ year streaks via loss-aversion streaks and variable rewards, driving 41% YoY revenue growth at scale.

— Rigorous analysis distinguishing implementation factors from platform effects, comparing Harvard RCT (custom-built, tightly scaffolded) vs general platforms; meta-analysis finding: effects larger on tutor-aligned assessments than transfer tasks.

— RCT with 275 programming students: both AI conditions improved task completion and reduced frustration, but neither produced knowledge gains, demonstrating the 'comfort trap' where easier work masks absent learning transfer.

— Technical playbook identifies five architectural pillars required for effective AI tutors (curriculum RAG, mastery model, pedagogy, conversation, engagement); cites ~70% day-30 collapse when pillars missing; Khanmigo 2025 RCT +0.34 SD.

— Metro Nashville DEC deployment: 19.50% faster time-to-competence, 10.95% higher mastery, 95.45% trainer alignment. Real named-org deployment scaling skill training in high-stakes domain.

The quest to build a better AI tutorResearch Paper

— University of Pennsylvania adaptive sequencing study with 800 Taiwanese high school students: personalized difficulty achieved 6-9 month equivalent improvement over fixed-difficulty control; demonstrates zone-of-proximal-development value.

— Empirical study of 16,851 real tutoring interactions identifying concrete elaboration (analogies, examples) significantly aids understanding; response length inversely associated with learning. Informs deliberate practice design.

— VC market analysis: Khanmigo grew 17x to 700K users but only 15% use it regularly (269K weekday interactions). Key finding: 'AI tutoring in 2026 has an adoption problem, not a technology problem.'

— North Carolina cut Khan Academy contract 95% (from $10M to $500K) citing Sal Khan's admission Khanmigo 'has been a non-event for students'; reflects institutional skepticism despite infrastructure investment.

— UC San Diego ASPIRE system across 12 courses (2,200+ students): Foundations of Precalculus failure rate dropped 35.3% to 11.4% through embedded AI tutoring preserving productive struggle. Demonstrates scale with named institution.

— Course restructuring scrambled learner progress, leaving users confused; 46% say annoying but continuing, 9% quitting. Reveals platform maturity gap when redesigns erase user investment.

— 30-month longitudinal study of 26,811 Chinese secondary students: AI assistance raised homework scores 18% but reduced exam performance 20% (1.4 SD effect) within 6 months, establishing cognitive offloading as population-scale mechanism undermining durable skill development.

— RCT (N=90 students): Hybrid AI+teacher co-design outperformed AI-only and teacher-only conditions on motor-skill gains (d=1.42 vs 1.01 and 0.78), lesson quality, and efficiency (67% planning-time reduction), validating human-AI collaboration as effective deliberate practice model.

— Stanford SCALE Initiative RCTs across two districts: AI tutors alone produced near-zero engagement and learning gains; human support increased usage 71-80% but remained insufficient without baseline engagement, establishing implementation and human involvement as binding constraints.

— Synthesis of Pew, RAND, OECD, and University of Chicago research on AI learning outcomes: productivity gains consistently fade or reverse when tool removed (exam performance advantage vanishes), establishing tool-removal decay as systematic limitation of unguarded AI assistance.

— PNAS study (Bastani et al. 2025): Randomized trial comparing AI tutoring with and without pedagogical guardrails showed GPT Base group achieved +48% on assisted exercises but regressed 17% on unassisted exams, while guardrailed version preserved learning, establishing desirable difficulty as critical design parameter.

— UK government procurement document and Education Endowment Foundation research call explicitly identify cognitive offloading as measurable harm; DfE Safety Standards now require tracking 'cognitive offloading rate,' establishing pedagogical design (cognitive preservation) as procurement-level requirement.

— World Bank analysis of 26,000 Chinese students: autonomous AI use for homework improved practice 18% but reduced exam performance 20% (1.4 SD effect), entrance exam scores drop 18-24%, establishing metacognitive laziness at population scale.

— Peer-reviewed study of 1,498 undergraduates: gap between AI-assisted output quality (7.62/10) and independent knowledge mastery (5.55/10) documents illusion of competence, showing AI assistance improves perceived but not actual skill acquisition.

— Multi-site RCT (1,222 participants) by Liu, Christian, Bakker, Dubey: unguided AI assistance reduces persistence and undermines independent problem-solving; effect occurs within ~10 minutes, establishing behavioral mechanism critical to deliberate practice design.

— Multi-site RCT (1,222 participants) shows unguided AI assistance reduces persistence and impairs independent problem-solving within ~10 minutes, establishing behavioral dependency as intrinsic risk directly undermining durable skill acquisition in deliberate practice contexts.

Will AI in Education Succeed?Industry Report

— Brookings synthesis of 20+ years EdTech research: success requires three conditions (teacher involvement, infrastructure readiness, rigorous evaluation); cites RCTs from Nigeria and Ghana showing gains only when teachers remain present; establishes deployment prerequisites.

— Peer-reviewed framework directly addressing AI-supported deliberate practice: EFFORT-AI specifies six phases to preserve cognitive effort and retrieve, explain, monitor, and transfer as learner-owned processes rather than delegated to AI.

— Meta-analysis of 4 RCTs (268 surgical trainees): AI tutoring showed small OSATS skill gain (0.20 SD) but significantly higher cognitive load, with authors concluding findings do not support replacing human tutors and indicating hybrid model required.

— Khan Academy founder candid assessment: AI tutoring outcomes are 'mixed bag' / 'neutral-to-marginal positive'; disconfirms 2023 leading-edge claims; acknowledges cheating concerns and low adoption despite infrastructure scaling.

— Large independent survey (1,660 teachers, 10 countries): gap between adoption (50% positive) and meaningful impact (13% 'very positive'); teacher training critical (75% well-trained vs 38% untrained report positive impact).

— Large-scale RCT (770 students, 10 Spanish schools) shows realistic outcomes scaling tutoring: math grades +0.15σ, standardized tests +0.11σ—one-third of pilot effects. Critical voltage-drop evidence for real-world deployment.

— RCT of Tutor CoPilot (AI augmenting human tutors) shows largest gains for less-experienced tutors, leveling expertise gaps. Frames AI as instructor augmentation not replacement, demonstrating collaborative model for deliberate practice support.

— EdWeek opinion cites UPenn finding (48% grade boost then 17% crash post-AI removal) and Stanford report (gains vanish in independent assessment). Sets four conditions for durable learning: domain knowledge, productive struggle, critical thinking, human interaction.

— Gallup survey: fewer than 1 in 5 teachers (18%) received formal AI guidance; 48% informal; 33% none. RAND data: majority use AI but fewer than half of principals have AI policies. Signals institutional barriers to structured deployment.

— US Department of Education IES grant ($3.75M, 3-year 2024–2027) funding CAIT system development with deliberate practice design: adaptive assessment, personalized sequencing. Federal R&D commitment signals institutional confidence; pilot RCT 2026-27.

— Critical practitioner analysis: Khanmigo adoption remains ~15% despite 2-year pilot and free launch; text-based bot limitations; system architecture deficit ('anterograde amnesia'—no cross-session memory); cost barrier and competitive displacement persist.

— Independent analysis of 502K+ customer reviews: Duolingo's April 'AI-first' announcement triggered trust erosion (0.27%→3.71%) and quality concerns persisting 11 months. Customers interpreted AI pivot as quality-control removal, predicting later bookings deceleration.

— Enterprise learning survey (2,000+ learners): 85% cannot apply AI training to their roles; 78% trained in systems disconnected from work context; evidence that skill acquisition fails without environmental alignment.

— Large-scale longitudinal study (3.2M math problem interactions): ChatGPT release caused 26.9% collapse in study time on AI-susceptible problems and 25% decline in learning gains on proctored assessment, establishing behavioral cognitive surrender.

— OECD 2026 finding: practice performance gains (+127% with generic AI) paradoxically reduce exam scores (-17%), establishing 'metacognitive laziness' as core mechanism undermining durable skill acquisition.

— RCT in Sierra Leone (1,800 students) and Italy (9,000 students) demonstrated Gemini grounded in learning science improved math mastery +0.26 to +0.38 SD; positive signal that well-designed AI tutoring with pedagogical grounding supports skill development.

— RCT (770 students): adaptive problem sequencing (not better answers) improved exam performance +0.15 SD; proactive pedagogical design outperforms reactive AI assistance; strongest gains for beginners, narrowing skill gaps.

— Critical market assessment: Sal Khan admitted Khanmigo became 'a non-event' for most students; 95% engagement exclusion rate and engagement cliff (60% drop after 3 weeks) documented; teacher tools bifurcated market from student tutors.

— Turkey RCT (1,000 students): unguarded AI raised practice accuracy 48% but lowered exam performance 17%; safeguarded AI version preserved skills; OECD identifies cognitive offloading as mechanism and documents pedagogical guardrails required.

— Research from Carnegie Mellon, MIT, Oxford, UCLA shows 10 minutes of AI assistance impairs analytical thinking and problem-solving persistence; identifies cognitive dependency as foreseeable risk directly undermining durable skill development.

— AIED 2026 peer-reviewed analysis of 10,235 code submissions from deployed AI tutors reveals engagement-based behavioral signals more strongly predict student learning than pedagogical quality alone; identifies substantial differences in tutor effectiveness in driving student action.

— Practitioner documentation of classroom Khanmigo implementation: students skip guiding prompts, seek direct answers, avoid productive struggle; reveals behavioral adoption gaps and weak knowledge transfer despite system design for deliberate practice.

— 6-month A/B testing (Oct 2025–April 2026) on 1.35M+ tutoring threads shows: learning history +3.4%, prerequisite review +2.7%, conversation context +5.09% engagement; demonstrates iterative refinement driving measurable skill transfer gains.

— Critical deployment signal: Khan Academy's own Chief Learning Officer admits only 15% of students with Khanmigo access actually engage with it, triggering redesign from passive to proactive AI; reveals adoption gap despite 108M interactions.

Duolingo Q1 Earnings Call HighlightsAdoption Metric

— Official metrics: 56.5M DAU (+21% YoY), 12.5M subscribers (+21%); product pivot to Speaking Adventures, Video Call, and spoken tokens demonstrates sustained deliberate practice investment; 20,500 AI-generated course units show content velocity for skill progression.

— Harvard RCT (Kestin et al., Scientific Reports 2025) documents AI tutors outperforming active learning on engagement and motivation; emphasizes effectiveness contingent on pedagogical alignment and is primarily demonstrated in STEM-structured domains.

— Frontiers in Psychology meta-analysis of 72 studies shows AI-enabled teaching yields effect size g=0.586 but is fundamentally conditional on implementation type and application context; emphasizes importance of appropriate AI deployment over presence alone.

— Engineering firm documentation of production ITS deployments (Khan Academy, Carnegie Learning, Duolingo) at scale (600K–1M+ students). Documents required five-layer architecture (learner model, curriculum, pedagogy, interface, evaluation) and effect sizes d=0.66–0.79 only when guardrails are applied.

— Khanmigo deployment case study documenting Khan Academy's four-phase measurement evolution from intuition-based to production A/B testing, revealing iterative improvement in pedagogical design.

— Working paper establishing cognitive harm from AI tutoring as foreseeable documented risk, with neuroimaging evidence and institutional liability framework; distinguishes adult cognitive atrophy from childhood cognitive foreclosure.

— Khanmigo deployment data: 108M+ interactions since 2023, but only 15% student engagement despite access; introduces 'next-item correctness' metric measuring independent learning transfer, not AI-assisted performance.

— Quasi-experimental study (60 EFL learners, Indian university) showing AI-mediated writing gains sustained at 5-week follow-up without AI, indicating skill internalization not tool dependency.

— Large-scale RCT (14,892 students across 4 countries, 62 schools) found MathMentor-GPT produced measurable math achievement gains with teacher integration amplifying effects; named mechanism.

— OECD's authoritative 2026 policy framework establishes pedagogical design (augmentation, Socratic questioning) as determinant of learning gains, not tool access alone; defines effectiveness conditions.

— Meta-analysis contrasting Harvard RCT (pedagogically-designed AI tutor, 0.73-1.3 SD gains) with Wharton RCT (unguarded ChatGPT access, -17% exam decline); demonstrates design determines outcome.

— Synthesis of peer-reviewed studies documenting critical limitation: AI assistance creates cognitive dependency, reversing 48% practice gains to -17% exam failure when tool removed.

— Peer-reviewed umbrella review synthesizing 102 systematic reviews on AI in K-12: confirms diverse applications but identifies critical gaps in teacher training, student AI literacy, ethical frameworks, and theoretical grounding of effectiveness research.

— Systematic review of 28 studies (4,597 K-12 students): ITSs show medium-to-large positive effects vs traditional instruction, but mixed results vs non-intelligent tutoring; effectiveness contingent on pedagogical design and teacher guidance, not technology alone.

— Khan Academy founder admits Khanmigo deployment 'was a non-event' for most students; uptake and learning gains fell short of 2023 promises; teacher adoption at featured schools waned despite administrator enthusiasm.

— RCT (N=1,222) shows AI assistance improves short-term performance but reduces persistence and impairs unassisted task performance, directly undermining skill acquisition; effect emerges after only 10 minutes of AI interaction.

— Independent comparative testing: Khanmigo (Socratic method) achieved 23% learning outcome improvement vs ChatGPT's 16% with direct answers; pedagogical approach significantly affects skill acquisition effectiveness in math and science.

— Critical assessment of Khanmigo deployment: effective for guided problem-solving when students are in 'productive zone' but ineffective for foundational instruction; weak for open-ended writing; named Hobart HS teacher reduced usage citing student frustration.

— University of Sheffield analysis documents fundamental AI tutoring limitation: models deliver errors with same confidence as correct answers; 45% of AI responses contain significant inaccuracy but users cannot distinguish; pedagogical mismatch: immediate resolution of confusion prevents learning.

— Stanford preprint identifies systematic bias in deployed AI tutor feedback: high-achieving/White students receive development-focused critique; Hispanic/ELL students receive grammar-only feedback; low-achieving students receive withholding bias.

— Harvard RCT with 194 physics students: AI tutor median learning gain (2.75→4.5/6) more than doubled vs active learning (3.5); completion 18% faster; pedagogical guardrails (active problem-solving first, cognitive load management) were critical design features.

— African Educational Research Journal (2026): quasi-experimental study showed AI personalized tutoring with adaptive pathways and learning analytics significantly improved academic performance and reduced at-risk student proportion; enables early detection of learning difficulties.

— 5-month RCT in Taipei high schools (770 Python students): reinforcement learning-enhanced AI problem sequencing improved final exam performance 0.15 SD with larger gains for beginners; isolates adaptive sequencing as scalable AI enhancement.

— Khan Academy 2022-23 annual report documents organizational deployment scale: Khanmigo launched; added 13.7M new learners; 945K US K-12 students in Districts Program; 40% YoY growth in very-active learners (18+ hours/year); 5-year target to accelerate 5M students.

— PRISMA systematic review (Frontiers Education, 2026) meta-analyzing 2013-2025 literature: AI adaptive learning systems yield medium-to-large effects (g=0.50-0.70) moderated by study quality, implementation duration, learner characteristics; identifies professional development as essential.

— Georgia State University's Jill Watson AI teaching assistant answered 10K+ questions in single semester with 95%+ satisfaction; response time <3 seconds; demonstrates infrastructure scalability for one-on-one tutoring to thousands of students per educator.

— RCT with 334 university economics students: unrestricted GPT-4 access raised exam performance 0.23 SD vs textbook control, 0.21 SD more than restricted access; contradicts premature AI reliance concerns when students adopt scaffolding strategies.

— Peer-reviewed controlled experiments (Chua & Pan, 2026) with adult learners: pretesting on word-picture matching yielded recall improvements (d=0.18-0.40) and multiple-choice gains (d=0.25-0.67); extends deliberate practice methodology to digital learning platforms.

— Quasi-experimental study (Frontiers Psychology, 2026): 383 primary students in AI self-study rooms vs traditional study showed significant effects on motivation (η²=0.171), self-regulation (η²=0.238), engagement (η²=0.220) over 10 weeks.

— Interview with Sal Khan and district admin: India study showed 0.44 effect size over 7 months; implementation threshold ~18 hours/year student engagement and 60-skill mastery; explicit requirement for teacher oversight and intervention—not autonomous AI.

— IZA Discussion Paper: RCT of 334 students shows unrestricted AI tutor access raises exam performance 0.23 SD; unrestricted outperforms restricted access 0.21 SD; benefits concentrated among low-baseline-knowledge students with strong self-regulation.

— AI for Cause case study shows Khanmigo growth from 68K (Mar 2023) to 700K users across 380+ US school districts, using Socratic method pedagogy; Harvard/Stanford RCT confirms effect size favoring AI tutor over traditional instruction.

— Duolingo's AI-first strategy drove growth but stock correction reveals risks: revenue growth slowing to 18-20% for 2026, margin compression from R&D spending, and competitive threats from free LLMs; signals unsustainability concerns in AI-powered platform economics.

— Rose Luckin critique argues AI tutors address only ~16% of learning (knowledge delivery), neglecting metacognition and social intelligence; cites research showing AI-assisted learners exhibit metacognitive laziness; documents narrow scope of current AI tutoring approach.

— Nationally representative AmeriSpeak survey finds 57% personal AI use and 20% professional use, with adoption sharply increasing by education level; documents broad AI tool penetration including learning support among educated demographics.

— Industry synthesis of 2,400+ enterprise AI initiatives documents 80.3% overall failure rate and 95% GenAI pilot failure-to-scale; identifies leadership and organizational barriers; critical evidence of adoption challenges constraining AI skill tool deployment.

— Digital Education Council survey of 30,000+ LATAM higher education respondents shows 92% student and 79% faculty AI engagement, with 50% of students supporting AI-assisted feedback; demonstrates rapid regional adoption surpassing global trends.

— SIAI analysis examines AI tutoring scaling risks: cites 2025 Nature study showing AI tutor outperformed traditional active learning, but documents persistent inequality risks and error concerns; balances efficacy gains with critical adoption barriers.

— SIAI research finds AI course bots handle 60-90% of routine questions in first year; Tutor CoPilot study shows 4pp higher mastery achievement with AI-augmented tutors; validates scaled AI tutoring infrastructure for deliberate practice support.

— Anthropic study of developers learning new library shows heavy AI assistance decreased library-specific skill mastery by 17% (two grade points); critical evidence that AI delegation reduces conceptual learning despite modest productivity gains.

— Critical analysis cites PwC survey: 56% of CEOs report no effect from AI investments on costs/revenue; 87% of workers report <10% productivity savings; documents persistent gap between AI adoption and real-world benefit realization.

— Duolingo case study of 25,000 test-takers shows AI-powered practice test improves confidence and certified test performance through adaptive item generation; demonstrates deliberate practice efficacy at scale in language assessment.

— Khan Academy announced Khanmigo deployment across New Mexico districts with features for personalized student support and teacher insights; signals continued geographic expansion of AI tutoring infrastructure into additional U.S. regions.

— McNorton analysis of Khanmigo efficacy across 350,000 students shows 20% higher learning gains for students meeting recommended usage; validates AI tutoring effectiveness and establishes baseline engagement-outcome linkage.

— Critical assessment examining limitations of Duolingo's grammar-drill mechanics in preparing learners for real-world conversational deployment, highlighting gaps between structured practice outcomes and authentic communication readiness.

— Review of June 2025 Harvard physics study showing AI tutoring with verified solutions outperformed human instruction (effect sizes d=0.73-1.3) and achieved 70% efficiency gain, addressing hallucination concerns in pedagogically-designed tutors.

— Systematic literature review of 48 peer-reviewed studies comparing AI tutors against human tutors and traditional instruction, establishing comparative effectiveness and global adoption patterns across diverse educational contexts.

— Empirical classroom trial validating safe and effective AI tutoring implementation with rigorous safety audits, providing evidence-based support for pedagogically-sound AI tutoring deployment in educational settings.

— Editorial review establishing that properly-designed AI tutoring surpasses traditional pedagogical best practices in controlled settings, with RCT evidence published in Scientific Reports providing rigorous validation for deployment efficacy.

ResultsAdoption Metric

— Multi-university qualitative adoption study across 4 Australian institutions with 79 focus groups and 8,000+ student survey responses analyzing student experience and institutional integration of AI tutoring tools.

— Duolingo achieved 51% year-over-year DAU growth to 47M users and $1B revenue forecast through AI-powered conversational practice, lesson generation, and video tutoring; demonstrates production-scale deployment of deliberate practice mechanics.

— Survey of 2,000+ US K-12 teachers finds 60% use AI weekly reclaiming 6 hours on grading, but 42% of untrained users report anxiety over errors; highlights adoption acceleration paired with critical training and confidence gaps.

— LinkedIn survey: 51% of professionals report AI training feels excessive; 33% embarrassed about understanding, 35% nervous discussing AI; linked to broader enterprise AI ROI challenges (95% pilot failure rate).

— Khan Academy's Khanmigo user base scaled from 40K to 700K in one year with 1M+ expected in 2025-26; critical assessment notes lack of RCT efficacy data and MIT study documenting cognitive risks from AI tool reliance.

— University of Iowa pilot of Khanmigo Teacher Tools found faculty used tool <1/week with no reported teaching impact due to manual copy-pasting friction; provides negative deployment signal on institutional integration barriers.

— Educator Elliott Hedman documents that Khanmigo fails underrepresented students lacking self-efficacy; argues AI tutors only work for already-motivated students and raise equity concerns about skill development access.

— Critical analysis questioning whether Duolingo's gamification and habit-formation mechanics drive actual language fluency or primarily psychological dependence on streaks and rewards.

— Preregistered experimental study demonstrating that personalised AI dialogue effectively corrects persistent misconceptions in psychology and education, validating dialogue-based deliberate practice for conceptual learning.

— Critical assessment of Duolingo's AI-accelerated course generation (150 new courses in one year) and job displacement of human language teachers; raises concerns about quality trade-offs and labor impact.

— University of Nebraska Medical Center integrated Khanmigo Teacher Tools into Canvas LMS with faculty training; demonstrates institutional adoption of AI tutoring infrastructure across higher education.

— Duolingo engineering deep-dive on Lily AI video call feature design for conversational practice: structured guardrails, predictable patterns, and adaptive difficulty to enable realistic speaking partner interaction.

— Research review showing 1,000-student Penn study where guarded AI tutor improved mastery 127% vs unrestricted ChatGPT's 17% loss on retention; validates deliberate practice guardrails including hints-not-answers methodology.

— Production engineering improvements to Khanmigo's math tutoring—calculator integration, GPT-4 Omni upgrade, improved error detection—demonstrate active refinement of AI deliberate practice infrastructure addressing numerical reasoning.

— Coverage of Class Companion and Tutor CoPilot adoption in K-12 classrooms, with Tutor CoPilot showing 4pp higher mastery in math for 1,800 underserved students; signals real-world deliberate practice deployment at scale.

— Product manager interview reveals Duolingo's 'Lily' AI video call feature design for conversational speaking practice with personality and adaptive difficulty; demonstrates production deployment of realistic speaking partner mechanics.

— Critical analysis by Dan Meyer documenting specific Khanmigo interaction failures (2-second premature intervention) and argues AI tutors fail to approximate human teaching; provides negative signal on current AI tutor effectiveness.

— Named school deployment of Khanmigo in Geometry and broader math classes with documented improvements in engagement and achievement, including benefits for ELL and special education students via data-driven grouping.

— Critical analysis by microschool educator argues AI tutors fail to address core learning motivation and meaningful context; highlights limitation of curriculum-aligned problem-solving in fostering durable skill development.

— Philippines Department of Education partnered with Khan Academy and Smart Communications for nationwide free AI-powered learning access; 5-10GB weekly data allocation demonstrates government-scale commitment to skill acquisition infrastructure.

— CBS 60 Minutes coverage documents Khanmigo deployment across 266 school districts including Hobart High School; pilot students and teachers report effectiveness; $15/student annual cost demonstrates scaling economics.

— Large-scale efficacy study of ~350K students in grades 3-8 found 30+ minutes weekly use associated with ~20% greater learning gains (effect size .36), with longitudinal analysis linking skill practice progression to incremental gains.

— Stanford RCT with ~1,000 students and 900 tutors found AI-augmented tutoring improved student mastery by 4-9pp, with stronger gains for novice tutors; validates human-AI collaboration in skill development.

— Appen/Harris Poll survey documents declining enterprise AI deployment (47.4% vs 55.5% in 2021) and ROI realization (47.3% vs 56.7%), highlighting implementation challenges that constrain AI skill tool adoption.

— Pilot RCT (n=37) tested 8-week structured deliberate practice course with peer feedback for therapist skill development; demonstrates efficacy of systematized deliberate practice methodology with embedded feedback in professional training context.

— Quasi-experimental study with 40 EFL students compared Duolingo-integrated and control groups, demonstrating AI-powered practice improves language learner engagement and willingness to communicate in classroom settings.

— Khan Academy expanded free Khanmigo access to Arizona, New Hampshire, Ohio, Louisiana, and Oklahoma in partnership with state education departments, demonstrating geographic scaling of AI deliberate practice across U.S. regions.

— Clever survey of 100K+ schools found rising educator optimism about AI in classrooms while identifying critical gaps in inclusive edtech; reflects Q3 2024 shift in teacher sentiment toward deliberate practice tools.

— KQED investigation documented AI hallucinations in mathematics tutoring, including specific Khanmigo failures on algebra problems; provides critical assessment of reliability limitations in AI-based deliberate practice systems.

— Microsoft and Khan Academy released Khanmigo for Teachers free across 49 countries, powered by Azure OpenAI; signals continued institutional backing and intent to scale AI-powered deliberate practice infrastructure globally.

— Production case study detailing Duolingo's four-step course creation combining human curriculum experts with AI content generation, exercise synthesis, and evaluation; demonstrates scaled human-AI collaboration in deployed skill practice platform.

— Research identified algorithm aversion as a critical adoption barrier in AI tutoring despite proven reliability; highlights trust as a bottleneck for broader deployment of deliberate practice systems.

— Peer-reviewed controlled experiment demonstrated that AI-generated evaluative feedback enhances skill acquisition and transfer in sequential decision-making tasks, with empirical validation via NSF-funded research.

— Microsoft partnership scaled Khanmigo for Teachers to all U.S. educators via free pilot access, covering $44 annual cost; signals ecosystem expansion and institutional backing for AI deliberate practice infrastructure.

— RAND survey of 1,020 teachers found 33% overall AI adoption, 18% regular use for skill-building tasks (content adaptation, material generation); adoption higher among higher-poverty schools and middle/high school instructors.

— Temple University faculty integration of AI for formative feedback and deliberate practice showed students shifted from viewing AI as cheating to collaborative learning tool; demonstrated via structured assignments with immediate AI feedback.

— Student survey and interviews found Khan Academy's AI tutor Khanmigo significantly improved comprehension and motivation through immediate feedback and autonomous learning support.

— Controlled study of 77 dental students showed ChatGPT learners achieved higher exam grades (P=.045) than literature-research group, demonstrating effective structured practice with immediate feedback.

— Duolingo deployed GPT-4 for lesson personalization and AI-powered conversation practice, demonstrating production deployment of deliberate practice mechanics at scale.

— IEEE study of ChatGPT, Gemini, and Copilot in mechanical engineering found AI struggles with numerical problem-solving but succeeds in theory-based skill instruction; 172-student survey.

— User evaluation of Duolingo Max's new AI-powered roleplay feature at scale; demonstrates deployed AI practice mechanics including scenario-based conversation with challenge escalation.

— Emerging product addressing core friction in language skill acquisition: finding conversation partners; demonstrates turn-based voice and text chat tutor design pattern for deliberate practice.

— Critical assessment of Duolingo's 21M+ daily active user base questioning whether gamified practice drives durable language proficiency or mainly engagement; identifies core tension between addictiveness and learning efficacy.

— CIPD survey of 1,108 L&D professionals found only 5% currently use AI for learning, with 6% planning adoption; documents very low corporate training adoption despite widespread interest.

— Randomized controlled trial of 1,800 K-12 students from underserved communities showed AI-guided tutoring improved mastery by 4pp (p<0.01), with 9pp gains for students of lower-rated tutors; validates AI-augmented deliberate practice at scale.

— Early pilot in Newark Public Schools documented critical accuracy issues: Khanmigo gave direct answers instead of guidance and fabricated misinformation; provides essential negative signal on deployment reliability risks.

— CEO announced Duolingo Max live product with GPT-4, early rollout stage across 20.3M daily active users; signals real-world deployment of AI-powered skill practice at consumer scale.

— Official product general availability of Duolingo Max subscription with Video Call and Roleplay AI features for conversational practice, powered by GPT-4 across 188 countries; signals consumer-scale deployment of deliberate practice mechanics.

— Randomized controlled trial with 165 students showed AI peer dialogue improved physics post-test scores by 10.5pp despite 40% intentional inaccuracy; demonstrates effectiveness of dialogue-based deliberate practice methodology.

History

2026-Sep: Early September evidence solidified the distinction between capability and outcomes. NBER research (Oreopoulos, Sept 7) reported a new 6,000+ student RCT across 18 Tennessee schools: AI tutoring gains depend entirely on design and guardrails; 96% of students tried the tutor but median engagement fell to 1/3 of practice days (the "just click away" problem persists). EuraStudy's secondary analysis (Sept 8) of 2,899 student sessions found 64–67% of students with high AI-assisted practice scores later scored below 50% on unaided exams, documenting the illusion-of-competence mechanism. Washington University competitive-programming study distinguished AI substitute-mode (harmful) from AI complement-mode (beneficial), with proctored gates separating skill preservation from scaffolding dependence—evidence that deployment context determines outcome. Zoho workplace analysis documented the "deskilling problem": 39% of workers report AI weakened their skill sets, junior engineers show measurably weaker debugging and conceptual ability despite working faster, establishing evidence that substitution for practice can actively degrade rather than support skill. Ecosystem maturity signals persisted: Google announced 30+ integrated Classroom AI tools (Sept 2), and SkillScouter's adoption compilation showed 86% global student AI use, 53% K-12 teacher use, +42% market CAGR coexisting with <10% institutional governance (9× adoption-to-policy gap). El Salvador's national expansion (1,198 students, 171 schools, World Bank backing) announced September 5 demonstrates large-scale institutional deployment feasibility in Global South. Institutional governance failures also surfaced: New Mexico Amira mandate raised data-privacy, scoring-accuracy, and efficacy-validation concerns, reflecting absent standards in procurement and evaluation. The emergent picture: 2026 adoption continues expanding globally at platform scale; but evidence distinguishes well-designed, pedagogically-grounded deployments (showing 0.7–1.3 SD gains, preserved transfer, engagement >50%) from unguarded implementations (showing harm, illusion of competence, <20% engagement), with governance and deployment context determining realized outcome. Later reviews reinforced this: a scoping review of 153 medical-education reports found only 13 tracked retention and none patient outcomes, an evaluation of 20 learning tools found only purpose-built tutors showed promise, and Workera found 56.4% of regular AI users get no work time to upskill.
2026-Aug: Duolingo's Q1 2026 earnings confirmed continued scale (56.5M DAU, +21% YoY) with 10x-accelerated AI course generation (20,500 units/quarter) and new speaking-focused features (spoken tokens, speaking adventures), while a Brookings-aligned synthesis of 20 causal K-12 studies reinforced that tools prompting active reasoning aid learning whereas systems generating complete answers undermine it. New RCT evidence sharpened the performance-learning dissociation: a 454-student RCT found a well-grounded RAG tutor produced zero learning gains despite high satisfaction, a 275-student programming RCT showed AI conditions reduced frustration and improved completion without producing knowledge gains (the "comfort trap"), and a 90-learner RCT found personalized feedback maximized long-term mastery but temporarily hurt next-attempt performance due to cognitive load. Mid-August evidence sharpened the calibration debate: Allen Institute's TutorMoments framework found LLMs instructed to "tutor well" over-help rather than push for reasoning, while a Harvard/MIT RL study (N=1,280) validated adaptive tutoring that calibrates help-level by learner confidence as a mechanism to preserve retention. Google, OpenAI, and Anthropic were reported pivoting learning products from quick answers toward scaffolded Socratic approaches, and medical-education research (Guardian commentary, peer-reviewed proposals) framed "never-skilling" risk with concrete interventions—independent reasoning before AI consultation, periodic no-AI cases—while an American University survey (N=483) showed employer AI-skill demand rising 285% (2024–2026) to 42.6%, cementing AI literacy as baseline rather than differentiator. Late-August evidence reinforced both scale and structural limits: Duolingo's Q2 results showed DAU accelerating to 58.7M (+23%) and paid subscribers to 12.7M (+17%), Google and Khan Academy shipped interactive Gemini-powered diagrams for Khanmigo, and Sindh's UNICEF-backed programme trained 3,503 teachers with measured lesson-planning and assessment gains. Countervailing evidence hardened the performance-learning paradox at population scale: a 26,811-student, 30-month study found autonomous AI homework use raised practice 18% but cut exam performance 20% (1.4 SD), a two-year NBER RCT across 18 Tennessee schools found Khanmigo gains matched Khan Academy alone with only one-third engagement, and Stanford SCALE's synthesis found AI-led tutoring shows 40-47% non-use rates with an insufficient evidence base, while a 13,037-student study showed structured "enacted feedback" dialogue nearly doubled uptake versus directed feedback.
2026-Jul: Multiple RCTs converged to establish the performance-learning paradox with quantified mechanisms: a 26,811-student longitudinal study found autonomous AI homework use raised practice scores 18% but reduced exam performance 20% (1.4 SD) within six months; a 1,222-participant multi-site RCT confirmed unguided AI assistance impairs independent problem-solving within ~10 minutes; and the PNAS guardrail study showed GPT Base achieved +48% on aided exercises but -17% on unassisted exams while a guardrailed version preserved learning—establishing pedagogical design as the critical variable. The policy layer responded: UK DfE Safety Standards now require procurement to track "cognitive offloading rate," while a PE education RCT (N=90) found hybrid human-AI co-design outperformed AI-only and teacher-only conditions on motor-skill gains (d=1.42), validating structured collaboration as the viable deployment pattern. Institutional confidence fractured further later in the month: North Carolina cut its Khan Academy contract 95% ($10M to $500K) following Sal Khan's own admission that Khanmigo has been "a non-event," while VC market analysis reframed the sector's core problem as adoption, not technology—Khanmigo's 700K users convert to only 15% regular use. Countering this, UC San Diego's ASPIRE system cut a precalculus failure rate from 35.3% to 11.4% across 2,200+ students in 12 courses, and a Metro Nashville 911 call-taker training deployment showed 19.5% faster time-to-competence, evidence that well-designed deployments can still produce named-institution gains.
Show earlier history (2023–2026 · 16 more) →

2026

2026-Jun: Multiple converging research streams documented the performance-learning paradox at behavioral and population scale. A World Bank analysis of 26,000 Chinese students found autonomous AI use for homework raised practice rates 18% but reduced exam performance 20% (1.4 SD)—establishing metacognitive laziness as a real population-level effect. A 1,222-participant multi-site RCT (Liu, Christian, Bakker, Dubey) documented that unguided GPT-5 assistance reduces persistence and impairs independent problem-solving within ~10 minutes, identifying the precise behavioral mechanism. A peer-reviewed study of 1,498 undergraduates quantified the illusion of competence gap: AI-assisted output quality 7.62/10 vs. independent mastery 5.55/10. The EFFORT-AI framework (Frontiers in Education) proposed six phases to preserve cognitive effort as a learner-owned process, and a surgical skills meta-analysis (4 RCTs, 268 trainees) found AI tutoring yielded only 0.20 SD gains with significantly higher cognitive load, concluding hybrid human-AI models remain necessary. Founder candor accompanied the research: Sal Khan admitted Khanmigo's real-world impact is "mixed bag / neutral-to-marginal positive." The Brookings synthesis identified three non-negotiable prerequisites for any deployment showing gains—teacher involvement, infrastructure readiness, and rigorous independent evaluation—none of which are consistently met at scale.
2026-May: Production refinement and behavioral evidence emerged alongside new research documenting the performance-learning paradox at scale. Khan Academy released detailed A/B testing results (May 6) from 1.35M+ tutoring sessions showing learning history (+3.4%), prerequisite skill review (+2.7%), and conversation context (+5.09% cognitive engagement) drive measurable skill transfer gains—indicating iterative engineering improves outcomes. However, critical redesign decision (May 5) publicly acknowledged 15% active engagement rate despite 108M total interactions, triggering shift from passive to proactive AI architecture. Duolingo Q1 earnings (May 4) showed 56.5M DAU (+21% YoY), 12.5M paid subscribers (+21%), with accelerated investment in Speaking Adventures, Video Call, and voice-first features. A large-scale longitudinal study (3.2M math problem interactions, arXiv May 2026) found ChatGPT's release caused a 26.9% collapse in college study time on AI-susceptible problems with 25% decline in learning gains on proctored assessment, establishing behavioral cognitive surrender at population scale. OECD analysis documented generic "fast AI" tools raising practice performance +127% while reducing exam scores -17%, naming "metacognitive laziness" as the core mechanism. Enterprise context mirrored: Docebo survey of 2,000+ learners found 85% cannot apply AI training to their roles, with 78% trained in systems disconnected from work context—confirming that skill transfer fails without environmental alignment. Peer-reviewed research (AIED 2026, May 7) analyzing 10,235 code submissions revealed engagement-based behavioral signals predict student learning better than pedagogical quality alone, identifying substantial tutor-to-tutor effectiveness variation in driving student action. Critical assessment intensified: research from MIT, Carnegie Mellon, Oxford, UCLA (May 8) documented 10 minutes of AI assistance paradoxically impairs problem-solving persistence and analytical thinking, with students relying on direct answers underperforming after tool removal. Engineering playbook from vendor confirmed production ITS at 600K–1M+ students (Carnegie Learning MATHia, Khan Academy, Duolingo) requires five-layer architecture (learner model, curriculum, pedagogical strategy, interface, evaluation) to achieve effect sizes d=0.66–0.79; without guardrails, systems regress to unguarded chatbots. The picture stabilizes: efficacy gains remain real and replicable in controlled, pedagogically-aligned contexts; but real-world deployment is bottlenecked by behavioral adoption barriers, cognitive dependency risks (skill atrophy without scaffolding), training burdens, and unresolved equity gaps.
2026-Apr: Deployment reality diverged sharply from efficacy research. Khan Academy founder Sal Khan admitted Khanmigo became "a non-event" for most students, with adoption stalling below initial 2023 promises (Chalkbeat, April 9). Named teachers at flagship schools (Hobart HS) reduced usage, citing student frustration; teacher sentiment skewed toward administrator enthusiasm rather than classroom uptake. Systematic evidence of limitations emerged: (1) RCT (N=1,222) showed AI assistance reduces persistence and impairs unassisted performance after just 10 minutes, directly undermining durable skill acquisition (Liu et al., arXiv April 6); (2) University of Sheffield analysis documented AI tutors delivering errors with identical confidence as correct answers (45% error rate, indistinguishable from correct responses); (3) Stanford preprint found systematic bias in AI tutor feedback by race, gender, and ability—high-achieving/White students receive development-focused critique, Hispanic/ELL students receive grammar-only correction; (4) Umbrella review (102 systematic reviews, Huang et al., Elsevier April 11) identified critical gaps in teacher training, student AI literacy, ethical frameworks, and theoretical grounding. Pedagogical design showed measurable impact: comparative testing found Socratic method (Khanmigo, 23% learning gain) outperformed direct-answer tutors (ChatGPT, 16%), but effectiveness bounded by prerequisite knowledge—Socratic design fails when foundational knowledge absent. ITS meta-analysis (28 studies, 4,597 K-12 students, npj Science of Learning April 10) confirmed medium-to-large effects vs traditional instruction but mixed vs non-intelligent tutoring, with effectiveness contingent on teacher guidance and pedagogical design, not technology. Duolingo's April "AI-first" announcement triggered user backlash (24.5% positive vs 41.1% negative globally; users citing job displacement and quality concerns). Institutional scaling remained stalled: University of Iowa pilot showed <1/week Khanmigo usage despite positive reception; 95% GenAI pilot failure-to-scale rate persisted. Equity gaps intensified: critical assessment noted Socratic design ineffective for underrepresented students lacking pre-existing self-efficacy; neither platform prepared learners for real-world deployment (authentic conversation, transfer to novel contexts).
2026-Mar: Pedagogical design mechanisms came into sharp focus. Khanmigo case study documented Socratic method pedagogy (guiding with questions rather than direct answers) driving 68K→700K user growth across 380+ US school districts with Harvard/Stanford RCT validation. Harvard physics RCT with 194 students showed AI tutoring learning gains >2x versus active learning when designed with active problem-solving-first and cognitive load management. University economics RCT (334 students) found unrestricted GPT-4 access raised exam performance 0.23 SD vs control, contradicting premature-reliance concerns when students adopted scaffolding strategies. Reinforcement learning-enhanced problem sequencing (Taipei high schools, 770 students) improved exam performance 0.15 SD. Systematic review of 2013-2025 literature documented AI adaptive systems yield g=0.50–0.70 effects moderated by study quality and professional development availability. Duolingo reached 135M monthly users (+40% revenue growth YoY, March 2026). Institutional scaling pressures remained: teacher training burden, integration friction, and equity gaps concentrating benefits among high-self-efficacy learners.
2026-Feb: Regional adoption accelerated with LATAM higher education reaching 92% student and 79% faculty AI engagement, while financial pressures mounted on core platforms: Duolingo experienced 23% stock correction due to growth deceleration (18-20% target) and margin compression, signaling economic challenges in AI-first strategy. Critical analysis documented narrow scope of AI tutoring (16% of learning dimensions) and systemic adoption barriers: industry synthesis found 95% GenAI pilot failure-to-scale rate and 80.3% overall AI project failure. SIAI research balanced efficacy gains (Nature-published AI tutor outperformance) against persistent inequality and error risks, indicating maturation beyond early-stage hype toward recognition of real-world adoption constraints.
2026-Jan: Duolingo's DET practice test demonstrated efficacy across 25,000 test-takers with improved confidence and performance through adaptive AI practice generation; Khanmigo continued geographic expansion into New Mexico districts with new pedagogical features. However, critical research from Anthropic revealed that heavy AI assistance impairs skill acquisition (17% decrease in library-specific mastery despite productivity gains), and economic surveys documented persistent adoption failures—56% of CEOs reported no effect from AI investments on costs or revenue. Infrastructure scaling continued but real-world benefit realization remained misaligned with deployment growth.

2025

2025-Q4: Rigorous efficacy research strengthened the case for AI tutoring: Harvard physics RCT (published Scientific Reports, Nov 2025) demonstrated pedagogically-designed AI tutors outperforming human instruction with effect sizes d=0.73-1.3 and 70% efficiency gain; systematic literature review of 48+ studies confirmed comparative effectiveness across global contexts; multi-university qualitative studies provided adoption insights. However, critical limitations emerged: analysis documented Duolingo failing to prepare learners for real-world conversational deployment—grammar drills do not transfer to authentic speech with natural slang and context-dependent patterns. Khanmigo reached 1M+ users with sustained Canvas LMS expansion. Adoption barriers persisted (training burden, institutional friction, equity gaps), and enterprise ROI remained flat, indicating efficacy validation has not yet translated into scaled institutional breakthrough.
2025-Q3: Duolingo achieved 51% year-over-year DAU growth to 47M daily active users with $1B revenue forecast, demonstrating production-scale success of AI-first strategy. Khanmigo student user base expanded from 40K (one year prior) to 700K with 1M+ expected by 2025-26 school year. However, adoption barriers intensified: University of Iowa pilot revealed <1/week faculty usage despite positive concept reception due to integration friction; teacher anxiety about AI errors persisted (42% of untrained users); 51% of professionals reported AI training as excessive burden. Equity concerns surfaced—critical assessment documented Khanmigo failing underrepresented students lacking self-efficacy, raising questions about whose skill development is actually supported. Survey data showed 60% of US K-12 teachers using AI weekly with documented time savings, but deployment economics and training gaps remained structural bottlenecks to scaled adoption.
2025-Q2: Khanmigo expanded Canvas LMS integration across universities (UNMC, Rutgers). Duolingo's AI course generation accelerated to 150 new courses in 12 months but triggered labor displacement concerns about human curriculum designer roles; Lily feature engineering documented guardrailed design patterns for conversational practice. Research validated AI dialogue's effectiveness at correcting misconceptions in learning contexts. Intensifying pedagogical critique questioned whether gamification and habit-formation mechanics in Duolingo actually drive language fluency or primarily psychological dependence on streaks. Labor displacement emerged alongside continued enterprise ROI challenges (47.3%) as dual headwinds to broader institutional acceleration.
2025-Q1: Khanmigo engineering improvements accelerated (math calculator integration, GPT-4 Omni upgrade for numerical reasoning, improved error detection) alongside named school deployments (Enid High School in Oklahoma with documented ELL/special education benefits). Duolingo launched Lily, AI video call feature for real-time conversational speaking practice with personality and adaptive difficulty. Empirical research strengthened guardrails argument: controlled study showed properly designed AI tutors achieved 127% mastery improvement vs unrestricted ChatGPT's 17% loss (Penn, 1,000 students). News coverage documented Tutor CoPilot and Class Companion deployments showing 4pp mastery gains in 1,800-student studies. Critical assessment persisted: educator Dan Meyer documented specific Khanmigo interaction failures and argued AI tutors fail to approximate human teaching, highlighting limitations in current design patterns.

2024

2024-Q4: Khan Academy's large-scale efficacy study of ~350K students confirmed 30+ minutes weekly usage associated with ~20% greater learning gains; 266 U.S. school districts piloting Khanmigo with documented teacher/student success. Philippines Department of Education partnership enabled nationwide free AI access with government data infrastructure. Stanford RCT validated human-AI collaboration model (Tutor CoPilot), with AI-augmented tutoring improving mastery by 4-9pp. However, enterprise AI scaling challenges persisted (deployment declining from 55.5% to 47.4%; ROI realization declining from 56.7% to 47.3%), and educators highlighted core limitation of curriculum-aligned problem-solving in addressing learning motivation and meaningful context necessary for durable skill development.
2024-Q3: Khanmigo expanded to five additional U.S. states with state education department partnerships; Microsoft partnership extended free access to 49 countries via Azure OpenAI (Aug 2024). Empirical validation continued with quasi-experimental studies showing Duolingo's impact on learner engagement; randomized trials validated deliberate practice methodology with structured peer feedback in professional training (Sep 2024). Teacher sentiment toward AI in classrooms shifted more positive (Clever survey, Sep 2024), though critical gaps in inclusive implementation persisted. AI hallucinations in mathematics tutoring emerged as concrete reliability concern, with documented Khanmigo failures on algebra problems (Aug 2024), threatening adoption credibility among institutions.
2024-Q2: Duolingo published four-step human-AI course creation methodology in production (Jun 2024); Khan Academy secured Microsoft partnership to scale Khanmigo for Teachers free to all U.S. educators (May 2024). Empirical research confirmed AI-generated evaluative feedback enhances skill acquisition in sequential tasks (PLOS ONE, May 2024). Teacher adoption survey (RAND, n=1,020) showed 33% overall AI adoption, 18% regular use for skill-building; Temple University faculty integrated AI for formative feedback with positive student outcomes. Algorithm aversion emerged as critical adoption barrier despite proven AI reliability (UCL, Jun 2024); pedagogical critiques raised concerns about AI tutoring mechanics.
2024-Q1: Duolingo deployed GPT-4 for lesson personalization and conversation practice, cutting contractor costs while scaling AI capabilities (Jan 2024); controlled studies showed ChatGPT-supported dental education outperforming traditional research-based learning on exams; Khan Academy's Khanmigo reported improved student comprehension and engagement in mathematics and science; engineering education studies revealed AI strength in conceptual teaching but limitations in numerical problem-solving.

2023

2023-H2: Duolingo reached 21M+ daily active users with new AI-powered roleplay and explanation features (Nov 2023); LearnLingo launched conversational AI language tutor on HN (Jul 2023); critical assessment raised questions about whether gamification-heavy practice produces durable learning outcomes.
2023-H1: Duolingo launched Max subscription (GPT-4 powered Video Call and Roleplay features, Mar 2023) and reported 20.3M daily active users in Q1. Khanmigo early pilot in Newark (Jun 2023) showed accuracy issues including fabricated misinformation, highlighting deployment reliability risks. Tutor CoPilot RCT (1,800 underserved K-12 students) validated AI-augmented tutoring with 4pp mastery gains (9pp for lower-rated tutors). Research demonstrated AI dialogue improves physics misconception correction by 10.5pp. Corporate L&D adoption remained minimal: CIPD survey found only 5% of training professionals using AI for learning, signaling enterprise adoption barriers despite consumer platform growth.

Tools