Perly Consulting │ Beck Eco

The State of Play

A living index of AI adoption across industries — where established practice meets the bleeding edge
UPDATED DAILY

The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.

The Daily Dispatch

A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.

Pick a role above to explore practices

BLEEDING EDGE

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🔬 RESEARCH & KNOWLEDGE
⚖️ LEGAL, COMPLIANCE & RISK
🎧 CUSTOMER OPERATIONS
🏛️ AI GOVERNANCE & SAFETY
📊 DATA & ANALYTICS
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💼 SALES & REVENUE
🎬 CREATIVE & GENERATIVE MEDIA
👁️ COMPUTER VISION & SENSING
💹 FINANCE & ACCOUNTING
🔄 OPERATIONS & PROCESS AUTOMATION
🚗 AUTONOMOUS SYSTEMS & VEHICLES
🦾 PHYSICAL AI & ROBOTICS
🎓 EDUCATION & LEARNING
PERSONAL EFFECTIVENESS

LEADING EDGE

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🔬 RESEARCH & KNOWLEDGE
⚖️ LEGAL, COMPLIANCE & RISK
🎧 CUSTOMER OPERATIONS
🏛️ AI GOVERNANCE & SAFETY
📊 DATA & ANALYTICS
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💼 SALES & REVENUE
🎬 CREATIVE & GENERATIVE MEDIA
👁️ COMPUTER VISION & SENSING
💹 FINANCE & ACCOUNTING
🔄 OPERATIONS & PROCESS AUTOMATION
👥 PEOPLE & TALENT
🚗 AUTONOMOUS SYSTEMS & VEHICLES
🦾 PHYSICAL AI & ROBOTICS
🎓 EDUCATION & LEARNING
PERSONAL EFFECTIVENESS

GOOD PRACTICE

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🔬 RESEARCH & KNOWLEDGE
⚖️ LEGAL, COMPLIANCE & RISK
🎧 CUSTOMER OPERATIONS
🏛️ AI GOVERNANCE & SAFETY
📊 DATA & ANALYTICS
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💼 SALES & REVENUE
🎬 CREATIVE & GENERATIVE MEDIA
👁️ COMPUTER VISION & SENSING
💹 FINANCE & ACCOUNTING
🔄 OPERATIONS & PROCESS AUTOMATION
👥 PEOPLE & TALENT
🚗 AUTONOMOUS SYSTEMS & VEHICLES
🦾 PHYSICAL AI & ROBOTICS
🎓 EDUCATION & LEARNING
PERSONAL EFFECTIVENESS

ESTABLISHED

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💹 FINANCE & ACCOUNTING
👥 PEOPLE & TALENT

🎓 Education & Learning

AI for teaching, tutoring, assessing, and managing learning experiences. Mostly leading-edge: adaptive tutoring and automated grading are approaching good practice, but institutional adoption is slow due to academic integrity concerns and uneven infrastructure. Three practices are bleeding-edge, including AI-generated curricula and autonomous classroom agents. Most trajectories are stalled — policy and pedagogy lag behind the technology.

15 practices: 2 good practice, 11 leading edge, 2 bleeding edge

Where AI Stands in Education & Learning

Education is the domain where AI adoption is closest to complete and the evidence of benefit is thinnest. A Becker Friedman Institute survey of more than 1,200 K-12 principals published this month found 90% of schools now use generative AI for teaching — and 30% of principals believe it has improved student learning. Gallup's survey of 2,069 US teachers puts weekly use at 30% and self-reported time savings at 5.9 hours a week, roughly six working weeks a year, while only 18% have formal district guidance. In UK higher education, undergraduate use runs above 90%. No other sector we track has reached this level of penetration with so little agreement about whether it works. The gap is not a lag in measurement; it is the finding. In August the US Department of Education's own research arm, the Institute of Education Sciences, published a rapid synthesis of twenty rigorous causal studies and reached three conclusions: teacher-mediated AI tutoring that gives hints rather than answers performs about as well as conventional instruction; student-facing tools produce mixed results; and general-purpose AI use is associated with worse learning outcomes.

The mechanism behind that last finding is now well characterized, and it is the single most important fact in the domain. Learning gains are real where the system is architecturally constrained to withhold answers. A pre-registered randomized trial of Guided Learning in Gemini across 1,763 students in twelve Sierra Leonean schools produced +0.258 standard deviations in math over eight weeks with 69% engagement, and interaction analysis confirmed the design was doing the work: 76% of AI turns were scaffolding questions, 2% were direct solutions. A Wharton trial across ten Taipei schools found 0.15 SD gains — six to nine months of additional learning — when Socratic guidance was paired with adaptive problem sequencing. Where the constraint is absent, the direction reverses. Turkish high-schoolers with unrestricted chatbot access scored 17% worse than controls on exams taken without the tool. A longitudinal study following 26,000-plus Chinese secondary students through a large deployment of automated grading recorded a 20% fall in exam scores within six months, even as teachers cut grading time from forty minutes to ten. A controlled study published this month found ChatGPT lifted essay scores while producing no measurable knowledge gain or transfer at all. Research connected to Anthropic found that professional developers learning an unfamiliar Python library with AI assistance scored 17% lower on a knowledge quiz afterwards, with no speed advantage to show for it.

What is actually advancing in this domain is not teaching. Two practices carry an advancing trajectory: admissions and administration automation, where 68% of major universities now run AI in the enrollment pipeline against 29% in 2024, and simulated practice environments, where corporate sales and contact-center roleplay has reached 58% Fortune 500 penetration and is now expanding into surgical, emergency-medicine and legal training. Both sit outside the learning-outcome argument entirely — one automates back-office paperwork, the other trains professional behavior in domains with measurable performance endpoints. Everything pointed directly at instruction and assessment is stalled: tutoring, grading, feedback, curriculum generation, language practice, learning analytics, question generation, content adaptation, and detection. That distribution is the clearest available statement of where the domain has and has not found product-market fit.

What's New, 2026-07-31 to 2026-08-14

The fortnight's defining development was the arrival of a government verdict alongside a formal institutional retreat. The IES synthesis gives US policymakers a citable finding that unstructured AI use hinders learning, published within days of the 90%-adoption, 30%-benefit principal survey. Two independent benchmarks then localized the problem in the models themselves rather than in deployment. The Allen Institute for AI's TutorMoments benchmark tested seven frontier models against 462 authentic math tutoring transcripts containing more than 1,500 teacher-flagged decision moments and found every model defaults to over-helping when given a plain instruction to tutor well; prompts that explicitly encode the trade-off improve behavior but do not reach human judgment. A separate empirical test across 779 simulated tutoring conversations and six frontier models found that without tutoring-specific prompting, models gave direct answers 97% of the time — effectively zero successful tutoring — and that no model could simultaneously diagnose misconceptions and withhold solutions. Sal Khan conceded in an August interview that the original Khanmigo "was a non-event" for most students, with the redesign now built around teacher-driven engagement rather than tool availability. A research brief on two Stanford-linked trials confirmed why: close to half of elementary students never logged in at all, and those who did averaged two to five minutes a week against the thirty needed for measurable reading gains.

The assessment-integrity architecture broke in public. New South Wales initiated a moratorium on unsupervised take-home assessment for the roughly 95,000 students in the 2027 HSC cohort, explicitly on the grounds that detection tools have failed — the largest formal admission yet that a national qualification cannot verify work done outside a supervised room. The legal exposure hardened in parallel: the Newby v. Adelphi decision annulled an academic integrity finding based solely on a 100% Turnitin AI score after independent detectors scored the same essay at 0%, and a Yale MBA student has filed a thirteen-count federal suit over a false positive tied to $208,500 in tuition. Independent testing of GPTZero on 861 verified human samples found a 13.8% false positive rate against a vendor claim of 1% or less, rising to 16% on writing by non-native English speakers. A systematic review of 26 studies concluded detection is unsuitable as a stand-alone basis for high-stakes judgment. Meanwhile a Higher School of Economics study found 90% of Russian high-schoolers use AI and 79% rewrite output specifically to look human-written.

Distribution kept expanding regardless. Google began pushing Gemini to 150 million K-12 Classroom users and made Gemini Notebook generally available inside Schoology; OpenAI shipped dedicated K-12 educator, college educator and student plugins. Against that, Chalkbeat documented AI-generated curriculum with factual errors, missing alphabet letters and nonsensical content selling at scale on Teachers Pay Teachers, a marketplace used by 85% of pre-K–12 educators. On the regulatory side, the EU AI Act's high-risk obligations for automated grading slipped to 2 December 2027 while transparency duties remain live from August 2026, and Ofqual issued guidance naming AI item generation as a priority use case for awarding bodies subject to mandatory human review. Nothing shifted position this cycle. Thirteen of fifteen practices remain stalled; the two advancing ones remain the two furthest from the classroom.

Key Tensions

  • Adoption has outrun outcome evidence, and the gap is now officially documented. Ninety percent of schools use generative AI; 30% of principals think it helped. The US Department of Education's own evidence review found general-purpose use associated with worse results, and a meta-analysis of 936 learning analytics papers found 70% contained no learning outcome measure at all. The field has spent three years measuring engagement and time savings while the outcome question went largely unasked.

  • The pedagogy that works is not the behavior models ship with. Scaffolded, answer-withholding systems produce genuine gains — +0.258 SD in Sierra Leone, 0.15 SD in Taipei. But the Allen Institute benchmark shows all seven tested frontier models default to over-helping, and without explicit tutoring prompts six models gave direct answers 97% of the time. Stanford research finds models affirm user actions 49% more often than humans do. Good tutoring is an architectural constraint that has to be imposed against the model's grain, and students bypass it when they can.

  • Detection has failed as a control, and institutions are absorbing the cost in courts and in curriculum. False positives run 13.8% on verified human text and 16% on non-native English writing against vendor claims near 1%; a July mathematical analysis argues the floor is information-theoretic rather than an engineering deficit. Courts have now labeled detector scores probabilistic guesses, litigation is naming disability and language discrimination, and NSW has responded by removing half of a national qualification's assessment from home. The replacement is supervised assessment, oral defense and process evidence — all of which cost staff time.

  • Efficiency accrues to institutions; the measured costs land on students, disproportionately on some. Teachers bank 5.9 hours a week; the Shanghai grading deployment saved 30 minutes per teacher per assignment while exam scores in the tracked cohort fell 20% over six months. A University of Pennsylvania trial with 193 teachers and 2,816 students found AI teaching assistants reduced student motivation and achievement, with the largest harm where teachers used AI output without revising it. Speech recognition still misrecognizes Black speakers at 35% against 19% for white speakers; AI assessment systems score identical resumes carrying disability-related honors lower 75% of the time. The federal remedy narrowed on 24 July 2026, when the Department of Education eliminated the disparate-impact investigation tool under Title VI, requiring proof of intentional discrimination.

  • Where deployment is genuinely accelerating, governance has not followed. Admissions automation went from 29% to 68% of major universities in two years, with Virginia Tech processing 250,000 essays an hour and named deployments reporting multimillion-dollar revenue effects. A public records survey found zero of 24 major public universities had a written policy governing AI in undergraduate admissions and zero provided staff training. Independent technical review of leading enrollment platforms found FERPA audit-logging and data-residency gaps, concluding that deployment pace outstrips compliance frameworks. UNAM's AI-proctored entrance exam for 160,000 applicants ended in roughly 58,000 mandated retakes — a preview of what failure looks like at that scale.

Top 10 Evidence Items

  1. AI Diffusion Gaps: Unequal Integration of AI Across K-12 Schools (adoption-metric) — The Becker Friedman Institute's 1,200-principal survey behind the domain's defining statistic: 90% adoption, 30% belief it helped, and the "adoption paradox" that opens this fortnight's summary. https://bfi.uchicago.edu/insights/ai-diffusion-gaps-unequal-integration-of-ai-across-k-12-schools/

  2. AI in K-12 Education: The Good, the Bad, and the Guardrails to Consider (industry-report) — The US Department of Education's own research arm publishing a rapid synthesis of 20 causal studies and concluding general-purpose AI use is associated with worse learning outcomes; the government verdict that turns adoption penetration into an officially documented problem. https://ies.ed.gov/learn/blog/ai-k-12-education-good-bad-and-guardrails-consider

  3. TutorMoments: Do AI tutors know when to help and when to hold back? (industry-report) — Allen Institute for AI's benchmark of seven frontier models against 462 real tutoring transcripts, showing every model defaults to over-helping; localizes the tutoring-outcome problem in model architecture rather than deployment discipline. https://huggingface.co/blog/allenai/tutormoments

  4. Measuring the impact of learning with AI in Sierra Leone and beyond (case-study) — The pre-registered RCT (1,763 students, 12 schools) that supplies the domain's clearest positive result, +0.258 SD in math, and the mechanism — 76% scaffolding questions, 2% direct solutions — that explains why most other deployments don't replicate it. https://www.toolai.io/zh/info/3288

  5. Half of Every HSC Grade Is Unenforceable: NSW Ends Take-Home Work Over AI Detection Failure (adoption-metric) — Australia's largest state qualification body pulling roughly 95,000 students' assessments out of the home entirely, the starkest institutional admission yet that detection has failed as a control. https://www.techtimes.com/articles/323700/20260810/half-every-hsc-grade-unenforceable-nsw-ends-take-home-work-over-ai-detection-failure.htm

  6. GPTZero Bypass: What That Search Means and What Actually Helps (research-paper) — Independent testing finding a 13.8% false-positive rate on verified human writing (16% for non-native English) against a vendor claim near 1%, the empirical basis for the detection-litigation wave now working through US courts. https://tohuman.io/blog/gptzero-bypass

  7. An AI-supervised remote exam went so badly that 58,000 students must retake it (news-coverage) — UNAM's AI-proctored entrance exam for 160,000 applicants, the concrete failure case for what happens when administrative AI scales faster than its validation — a preview of the risk in the 68%-of-universities admissions-automation boom. https://arstechnica.com/culture/2026/08/an-ai-supervised-remote-exam-went-so-badly-that-58000-students-must-retake-it/

  8. Student Defense Report Highlights Lack of Policies Governing AI Use By College Admissions Offices (industry-report) — A public-records survey finding zero of 24 major public universities have a written AI admissions policy or staff training, documenting the governance vacuum beneath the domain's fastest-growing practice. https://defendstudents.org/all/student-defense-report-highlights-lack-of-policies-governing-ai-use-by-college-admissions-offices

  9. Teachers report low-quality AI content on popular marketplace for classroom materials (news-coverage) — Chalkbeat's investigation of Teachers Pay Teachers, used by 85% of pre-K-12 educators, finding AI-generated curriculum with factual errors and nonsensical content selling at scale — distribution outrunning quality control in exactly the practices the IES synthesis flags as stalled. https://www.chalkbeat.org/2026/08/03/ai-slop-on-teachers-curriculum-marketplace/

  10. Teachers save time with AI. Their students may pay the price (news-coverage) — The University of Pennsylvania RCT (193 teachers, 2,816 students) showing AI teaching assistants reduced student motivation and achievement, worst where teachers used AI output unrevised — direct evidence for the tension between institutional time savings and student-level cost. https://abc17news.com/stacker-news/2026/07/30/teachers-save-time-with-ai-their-students-may-pay-the-price/