🎓 Education & Learning
AI for teaching, tutoring, assessing, and managing learning experiences. Mostly leading-edge: adaptive tutoring and automated grading are approaching good practice, but institutional adoption is slow due to academic integrity concerns and uneven infrastructure. Three practices are bleeding-edge, including AI-generated curricula and autonomous classroom agents. Most trajectories are stalled — policy and pedagogy lag behind the technology.
The Headline
In education, AI improves the work students hand in but not what they learn. This fortnight, leading institutions stopped trusting AI detectors and started changing how they test.
The Picture
AI is now standard in education. 68% of US public school districts formally contract a generative AI platform, and more than nine in ten UK undergraduates use AI, mostly on assessed work. So far the payoff sits around the classroom rather than inside it: in admissions processing, administration and simulated practice such as sales role-play, where the output really is the goal. Inside teaching, a small group of institutions is pulling ahead. They choose tutors built to coach rather than answer, and they redesign how they test. Most others are paying for licenses students barely use and relying on grades and detectors that no longer show what a student can actually do.
This Fortnight
Harvard's dean told faculty to get out of the AI-detection business, and the New South Wales education authority told schools not to rely on detectors. UC Berkeley now discourages relying on them because they are too often wrong, and an independent benchmark found Turnitin and Originality missed almost all text that mixes human and AI writing. Any disciplinary process that treats a detector score as proof is now a legal and reputational liability.
A six-month observation of 20 AI products in 16 school systems found none ready to do the teaching job on its own. Purpose-built tutors that coach rather than answer helped most. General chatbots let students skip the hard thinking, and students routinely ignored detailed AI writing feedback. When buying, ask what a tool holds back from students, not just what it can generate.
In a Tennessee school trial, nearly every student tried an AI tutor, but it was used at only 17% of the moments students made errors. Licenses and log-ins are not learning. The same gap appears in corporate sales training, where AI role-play gets a fraction of the practice needed to build skills. Before renewing, measure whether the tool gets used at the moment of struggle.
A randomized trial of 405 pupils found that AI-only summaries produced the worst understanding and recall, even though pupils preferred them. A separate study in Uganda and South Africa found AI feedback raised essay scores while making students more overconfident. Student satisfaction is a misleading measure of these tools. Ask for evidence of what learners can do without the AI.
In one study, sending only the AI's low-confidence grading decisions to people cut manual review by about 80%. It is the most practical answer so far to a 52-study review that found AI and human markers agree only moderately. It also fits a broader review that found no support for fully automated high-stakes scoring. If you use AI marking, build in human-in-the-loop review (a person checks the output before it counts) for the uncertain cases.
Coming Up
Under the EU AI Act, AI used in grading and admissions counts as high-risk, with compliance due by December 2027. The rules require documented human oversight and transparency. Start listing now every tool that touches a grade, a credential or an admission decision, and ask each vendor for its oversight documentation.
Large US public institutions face the ADA Title II web accessibility deadline in April 2027, and smaller ones in 2028. AI is not closing that gap: most accessibility teams now use it, yet 95.9% of home pages still fail standard audits. Budget for human review of captions and documents rather than assuming the tools make you compliant.
Congress has asked the Government Accountability Office to assess AI in college admissions, as automated direct-admissions programs spread to at least 17 states. A public-records survey of major public universities found none had a written policy on AI in admissions. In July the Education Department also dropped its tool for investigating unintentional discrimination, which pushes that risk toward the courts. Write the policy and keep a record of bias testing before a regulator or plaintiff asks for them.
What's Hard About This
AI makes the work better and the learner weaker. A 30-month study of about 27,000 students found that AI homework help lifted homework scores 18% while exam scores fell 20%. Only tools built to withhold answers reliably avoid that penalty, and they are a minority of what gets deployed, so a more capable model will not fix it.
Every credible deployment relies on human review, and that safeguard is fraying. In one UK survey, 56% of teachers said checking AI outputs cancels the time it saves, and other research shows teachers correct harsh AI grades less often than identical human ones. Oversight has to be designed and staffed, not assumed.
When AI can produce a credible answer to almost any take-home task, grades stop measuring the student. The fix is to watch how work gets produced, through oral exams, staged submissions and draft histories, and that costs more: MIT spent $150,000 rebuilding a single course. Budget for assessment redesign as an ongoing cost, including for internal certifications and training exams.
Practices in this Domain (15)
| PRACTICE | TIER | TREND |
|---|---|---|
| Accessibility & accommodation support in education | LEADING EDGE | — Steady |
| AI tutoring — conversational & guided discovery | LEADING EDGE | ↘ Slowing |
| AI tutoring — personalised pacing & adaptive difficulty | GOOD PRACTICE | — Steady |
| Automated grading & assessment | LEADING EDGE | ↘ Slowing |
| Coding education & interactive exercises | LEADING EDGE | — Steady |
| Curriculum design & content generation | LEADING EDGE | — Steady |
| Education administration & admissions automation | LEADING EDGE | — Steady |
| Educational content adaptation & summarisation | BLEEDING EDGE | — Steady |
| Formative feedback generation | LEADING EDGE | — Steady |
| Language learning with conversational AI | LEADING EDGE | — Steady |
| Learning analytics & student risk identification | LEADING EDGE | — Steady |
| Plagiarism & AI-content detection | BLEEDING EDGE | ↓ Declining |
| Question & exam generation | LEADING EDGE | — Steady |
| Simulated practice environments | LEADING EDGE | — Steady |
| Skills assessment & competency mapping | GOOD PRACTICE | — Steady |
Read the full technical briefing (1,777 words) →
Where AI Stands in Education & Learning
Education is the domain where AI adoption and AI results have drifted furthest apart. The tools are everywhere. 68% of US public school districts now formally contract a generative AI platform, up from 42% two years ago. Google's Gemini holds about a third of that market, Khanmigo 22% and Microsoft Copilot 19%, and OpenAI's ChatGPT for Teachers now reaches more than 300,000 educators. Students got there first: more than nine in ten UK undergraduates use generative AI, most of them on assessed work. Yet the most consistent finding from two years of rigorous evidence is that AI improves the work students hand in, not what they learn. A 30-month study of roughly 27,000 Chinese secondary students found that AI homework help lifted homework scores by 18% while exam scores fell 20%. The Turkish high-school trial behind the "48% better in practice, 17% worse in the exam" figure has become the field's reference point. Systems built to withhold answers and scaffold reasoning do produce gains: Google's Gemini tutor in Sierra Leone (+0.26 SD at 69% engagement), Dartmouth's physics tutor and Tübingen's R tutor. General-purpose chatbots handed to students tend to erase them. Design decides the outcome far more than model capability, and most deployments default to the wrong design.
Momentum is real, but mostly around the classroom rather than inside it. Administration and admissions automation is moving fastest, because there the output really is the goal. Virginia Tech's essay reader handles 250,000 essays an hour, with humans settling disagreements. At least 17 US states run automated direct-admissions programmes, and Alabama's alone produced $5.1bn in scholarship offers to more than 12,000 students. Workday's AI products are at roughly $600m in annual recurring revenue, and the UAE's Edu Hub runs admissions across 74 institutions. Simulated practice is the other growth area, because there repetition is the product. Pavilion's benchmark of 268 sales-enablement teams credits AI role-play with 41% faster skill acquisition, and medical, legal and therapy-training deployments are spreading, from Vanderbilt Law's deposition simulator to CUHK's Cantonese-language virtual client. Older adaptive systems such as McGraw-Hill's ALEKS (7m+ users), and objective and code grading through Gradescope (2,600+ universities), are now infrastructure rather than experiments.
The teaching core is stuck, and September's story is institutions pulling back. New York City barred generative AI for students in K-8 (around 600,000 pupils), Los Angeles Unified blocked student access on district devices (around 378,000), and Chicago dropped a district-wide Gemini rollout. MIT's audit of its 6.036 course found 73% of submissions AI-generated, and it spent $150,000 rebuilding the course around oral exams and process portfolios. Six Singapore universities moved assessment away from essays. AI-content detection, the sector's first defensive reflex, is being abandoned by Harvard, Berkeley, Yale, Johns Hopkins and others as unreliable and biased against non-native writers. What sets education apart from neighbouring domains is that the measurable output is not the thing being bought. A tool that makes homework faster can do its job and harm the student in the same moment. The binding constraints are therefore pedagogical design, assessment validity and the capacity for human oversight, and none of these moves on a vendor's release cycle.
What's New, 2026-09-11 to 2026-09-25
The fortnight's most useful evidence came from watching tools in real classrooms rather than labs. Instruction Partners observed 20 AI products across 16 school systems for six months. Purpose-built tutors such as Khanmigo and Quill helped most, while general chatbots let students skip effortful thinking. Detailed AI writing feedback was routinely ignored: students rewrote their work without reading it. Automatic text re-levelling risked becoming a permanent lower setting. The verdict was that nothing observed was ready to do the pedagogical job on its own. In a Tennessee trial across 18 schools, 96% of students tried an AI tutor, but it was used at only 17% of the moments when they made errors. A randomised trial of 405 pupils aged 14–15 found that LLM-only summarisation gave the worst comprehension and memory, even though pupils preferred it. A controlled study of 180 undergraduates in Uganda and South Africa found that LLM feedback raised essay scores but significantly increased overconfidence. The PersonaPath benchmark found that the best of ten open models passed only 29.5% of learning-path planning tasks, despite producing structurally valid plans 90.9% of the time. The counter-evidence again favoured design: Tübingen's preregistered study found 12-point transfer gains from a Socratic R tutor, and a 72-study review found that graduated access to AI reduces passive reliance.
Assessment and integrity moved most. Harvard's dean told faculty to get out of the AI-detection business, Berkeley discourages reliance on detectors, and the NSW education authority told schools not to treat them as a primary safeguard. An independent benchmark put Turnitin and Originality below 0.55 macro F1, a measure of accuracy that weighs missed and false flags equally. Vanderbilt estimates about 750 false flags per 75,000 papers. Pangram has become the contested exception: it has admirers at Bloomberg and Wiki Education, and a public dispute after it rated the Dartmouth provost's recent writing a median 96% AI. On AI grading, a 52-study meta-analysis put human–LLM agreement at only r=0.66. A preregistered study of 1,426 dissertations found LLM graders scored student work below AI-generated text. A 43-study scoping review found no support for autonomous high-stakes scoring. One study showed that routing only low-confidence cases to humans cut manual review by about 80%. Elsewhere, Wisconsin pulled its Dropout Early Warning System after a reported 75% error rate and racial bias, and a review found only 11 of 689 learning-analytics papers had evaluated an intervention. A 100-student trial of a multi-agent clinical-interview simulator raised exam and communication scores but left diagnostic accuracy unchanged (84% against 86%). A 153-report review of generative AI in medical education found only 13 reports with retention data and none measuring patient outcomes. Level Access found that 80% of professionals use AI for accessibility work while 95.9% of home pages still fail WCAG audits. Cornell's $2m teaching pilots explicitly fund avoiding AI where appropriate. Admissions automation was quiet this fortnight.
Key Tensions
Better work, weaker learners. AI raises the quality of what students submit while lowering what they can do unaided. The 27,000-student Chinese study (homework up 18%, exams down 20%) and a 1,066-undergraduate natural experiment point the same way: in the latter, failure rates rose from 2–6% to 18.4% once AI was taken away. This fortnight's overconfidence finding points there too. Only tutors built to withhold answers reliably avoid the penalty, and they are a minority of what schools actually deploy.
Access is not use. Districts buy licences; students do not turn up. Stanford SCALE trials found that 40–47% of elementary pupils never logged in, and active users averaged 2–5 minutes a week against a 30-minute threshold for measurable gains. Khan Academy says only 15% of students with access use Khanmigo regularly. The pattern extends to corporate training, where sales teams average 2–6 AI role-play sessions a year against the 20–30 repetitions skill-building requires.
Assessment is losing its evidential footing. When AI can produce a credible answer to almost any take-home task, grades stop measuring the student, and detectors cannot restore the signal. Institutions are shifting from policing finished work to observing how it was produced: oral exams at MIT, Kellogg and Singapore's universities, staged submissions, draft histories. That costs more per student, which is why MIT's redesign of one course cost $150,000.
Human oversight is the safeguard, and it is fraying. Nearly every credible deployment rests on a human reviewing AI output: Singapore's Markly feedback tool, Virginia Tech's admissions reader, and the EU AI Act's high-risk rules for grading and admissions due by December 2027. Yet teachers correct harsh AI grades 22% less often than identical human ones, 80% of AI feedback in one study went out unedited, and 56% of UK teachers surveyed say checking outputs cancels the time saved. Routing only low-confidence cases to people is the most promising fix so far.
Equity harms are outrunning governance. Evidence of bias keeps arriving while the checks weaken. Identical essays draw different AI feedback depending on students' race and language status, detectors misflag 61% of non-native English essays, and Wisconsin's early-warning system was withdrawn after disproportionately flagging Black and Latino students. Meanwhile the US Department of Education dropped its Title VI disparate-impact tool in July, and a public-records survey found none of 24 public universities had a written policy on AI in admissions.
Top 10 Evidence Items
- MIT's 26,000-Student Learning Outcome Study: AI Homework Gains, Exam Losses, and Institutional Assessment Redesign (news-coverage) — Anchors the central tension: AI lifts homework by 18% while secure exam scores fall 20%, and institutions respond by redesigning assessment. https://www.inquirer.com/education/artificial-intelligence-college-students-universities-approaches-mit-harvard-ohio-chicago-20260918.html
- To Get Past AI Hype, Researchers Watch Students Use Actual Tools in Class (news-coverage) — Direct source for the fortnight's headline: purpose-built tutors help, general chatbots let students skip effortful thinking, and no tool can do the pedagogical job alone. https://www.the74million.org/article/to-get-past-ai-hype-researchers-watch-students-use-actual-tools-in-class/
- Student Adoption of an AI Tutor for Statistical Programming: A Longitudinal Study (Tübingen, R-Tutor) (research-paper) — The design-decides-outcome counter-evidence: a preregistered Socratic tutor delivering 12-point transfer gains. https://edtechdev.github.io/aied/articles/ai-tutor-statistical-programming-adoption-2026/
- Algorithmic Self-Deception: Overconfidence from LLM-Generated Feedback (research-paper) — Shows the better-work, weaker-learner pattern in feedback: essay scores rise while learners become measurably overconfident. https://lanfrica.com/fr/record/algorithmic-self-deception-how-ai-generated-feedback-skews-learners-self-reflection
- AI is now part of the edtech stack — and schools are repeating the mistake they made in every wave before (opinion) — Quantifies access versus use: 96% of students tried the tutor, yet it was used in only 17% of error moments. https://fortune.com/2026/09/11/ai-edtech-schools-repeating-tech-mistake/
- Evaluating the accuracy and reliability of AI content detectors in academic contexts (research-paper) — Independent benchmark showing the two most-deployed detectors below 0.55 macro F1, which is why the sector is abandoning detection. https://edtechdev.github.io/aied/articles/hadra-ai-detector-accuracy-efl-2026/
- Unis and schools are moving away from AI-detection software. They should stop using it altogether (opinion) — Reports the NSW education authority's move this month to stop schools relying on detectors, and argues for dropping them altogether. https://theconversation.com/unis-and-schools-are-moving-away-from-ai-detection-software-they-should-stop-using-it-altogether-292593
- Wisconsin Dropout Early Warning System removed after documented failures and racial bias (news-coverage) — A concrete equity and governance failure: a system withdrawn after a 75% error rate and racial bias. https://coffee-web.ru/blog/what-do-students-think-of-wisconsins-dropout-algorithm/
- Confidence-routing reduces manual assessment review by ~80% while preserving reliability (research-paper) — The most promising fix for fraying oversight: routing only low-confidence cases to humans cuts manual review by about 80%. https://edtechdev.github.io/aied/articles/know-when-to-trust-ai-scoring-reliability-2026/
- New York City restricts student-facing AI for grades K-8 (~600K students); Los Angeles Unified bars student device access (~378K students) (news-coverage) — Captures the September institutional pullback, with NYC and LA restricting student AI use despite rising district contracts. https://www.bloomberg.com/news/articles/2026-09-11/nyc-la-montessori-schools-limit-ai-as-families-flock-to-low-tech-classrooms