The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
AI for individual productivity, communication, organisation, and self-directed learning. The most polarised domain: writing assistance and meeting summarisation are good practice, but nearly half the practices are bleeding-edge — personal AI agents, life planning, and autonomous scheduling lack reliable implementations. Most trajectories are stalled, reflecting a gap between consumer hype and sustained daily utility.
This is the domain where AI is most widely used and least able to prove it pays. The everyday substance of knowledge work sits here — writing, email, brainstorming, research, scheduling, presentations, spreadsheets — and collectively it is the single largest category of workplace generative AI use, with writing alone accounting for roughly 80% of it. Commercially, the vendors have just posted their strongest quarter yet. Microsoft's FY26 Q4 earnings confirmed more than 30 million paid Copilot seats with net additions doubling quarter on quarter and 40 million agents registered; Adobe's Firefly reached roughly $300M in annual recurring revenue, up 50% in a single quarter; Gamma runs $100M ARR on a fifty-person team; Duolingo holds 56.5 million daily active users. On the other side of the ledger, a multi-source synthesis drawing on MIT, BCG, KPMG and McKinsey data puts 95% of AI pilots at zero measurable profit impact, with only 5–8% of organisations achieving returns at scale. Of the thirteen practices tracked here, twelve remain stalled; only translation is advancing. Vendor revenue and customer-side proof have decoupled, and nothing in this scan closed the distance.
What makes this domain structurally different from, say, fraud detection is that the user is an individual and the output is often that person's own thinking, voice or judgment. That produces two costs no amount of model improvement removes. The first is the verification tax, which now has a defensible number attached to it: a Work AI Index survey of 6,000 workers found AI saving eleven hours a week while claiming 6.4 of them back in checking and fixing — a net gain of four to five hours, with 69% admitting they ship unreviewed work anyway and only 13% reporting improved organisational performance. A G2 analysis of 2,771 verified reviews explains the distribution: 88% of buyers are casual users for whom AI produces genuine time savings, while the 12% who write professionally report that review overhead cancels the stated gain outright. The reliability floor beneath that tax is now measured rather than asserted. Independent benchmarking of six legal AI platforms — Harvey, CoCounsel, Lexis+, Westlaw, Spellbook and Ironclad — documented hallucination rates of 8–23% across research, contract review and regulatory lookup, with vendor accuracy disclosures inconsistent and self-referential. Fabricated citations then surfaced in PwC's own AI-generated client reports, making it the second Big Four firm caught after KPMG's June retraction.
The second cost is cognitive, and it is the domain's most uncomfortable finding. Peer-reviewed work by Capraro and colleagues found AI advice cutting human accuracy from 27% to 9% while nearly tripling confidence and collapsing the recognition of uncertainty from 44% to 3%. Three new randomised controlled trials in skill acquisition converged on the same shape: a 454-student trial found a well-grounded AI tutor produced zero learning gains despite high satisfaction; a 275-student programming trial found AI reduced frustration and improved task completion while producing no knowledge gain — practitioners are now calling it the comfort trap. The through-line is that the felt experience of using these tools is a poor and sometimes inverted signal of what they deliver. That, more than any capability ceiling, is why individual value refuses to aggregate: a WRITER-commissioned survey of 2,400 respondents recorded individual task gains as high as 5x against just 29% of organisations reporting significant return. Translation remains the exception that clarifies the rule — it advances precisely because its output is meant to be invisible rather than personal, its returns are documented, and its residual constraints are narrow and honest.
Position held everywhere. No practice changed tier or trend, and the fortnight's evidence deepened existing structural reads rather than redirecting them. The clearest movement was commercial. Beyond Microsoft's seat count came named agentic deployments at scale — Atos running 19,000 agents, Chow Tai Fook 400-plus agents with a 57% sales conversion lift, EY across 150,000 employees, and Accenture's 743,000-seat rollout, the largest publicly announced. Microsoft claimed weekly Copilot engagement now at parity with Outlook and Teams, though independent analysts flagged the methodology behind that claim as opaque. Google shipped its full Gmail AI stack to general availability on 15 July and put formula-error diagnosis into Sheets across 28 languages. Against all of it: 35.8% of licensed employees use Copilot regularly versus 83.1% for ChatGPT; a fresh survey found 95% of organisations reporting zero measurable return despite 80% adoption; and strategic analysis noted GitHub Copilot has lost competitive ground to Cursor and Claude Code despite Microsoft's distribution dominance and $13B OpenAI investment. Adobe supplied the sharpest version of the same paradox from its own worker survey — 90% of workers want creative AI, 9% use it, with roughly 400 hours a year going unclaimed.
The most consequential research this fortnight concerns confidence. Alongside the Capraro finding and the education trials, a BCG field experiment with 758 consultants drew the capability boundary precisely: +12.2% productivity and +40% quality on tasks inside the model's competence frontier, against a 19-point correctness drop on tasks outside it — with no reliable way for the user to tell which side they are on. Related work showed chain-of-thought explanations systematically misrepresent the computation the model actually performed, meaning the reasoning trace cannot serve as a check. An HBR field study completed the picture with frontline workers defending AI decisions they neither made nor understood. Read together, these findings shift the risk from "the model is wrong" to "the human cannot tell", which is a governance problem rather than an accuracy one.
Two structural signals closed the fortnight. First, the ROI narrative itself is being repriced. A Futurum survey of 830 IT decision-makers found time savings collapsing 5.8 points as the leading purchase justification (23.8% to 18.0%) as CFOs demand direct profit-and-loss evidence; CoreView found 66% of enterprises have delayed or cancelled Copilot deployments over concerns it surfaces confidential SharePoint data, with 75% of C-level respondents hesitant. Where deployments do stick, organisational design is the differentiator: named case studies at Ropes & Gray, Citigroup and Mars showed internal champion networks driving roughly twice the sustained usage of top-down mandates, with Harvey scaling to 282,000 prompts a month. Second, personal productivity tools are quietly becoming substrates for AI assistants rather than destinations in their own right. Superhuman shipped an MCP server letting Claude and ChatGPT read, draft and send mail directly; Obsidian's local REST API MCP server passed 2,700 GitHub stars as Claude Code plus Obsidian consolidated into the reference pattern for agent-native personal knowledge management; Akiflow and Reclaim have moved the same way in scheduling. Meanwhile translation added named production scale — NVIDIA's internal Nemotron Speech cut translation time 70% and cost 25%, Grab piloted Gemini 3.5 Live Translate across more than 10 million monthly Southeast Asia voice calls, and five named enterprises (CATS, Rosenbauer, Schrack Technik, Fandom, ALN Africa) reported 12.5–75% cost and time reductions — against a persistent equity gap in which accented speakers face 12% error rates versus 1% for General American English.
Vendor revenue and buyer-side proof have decoupled. Microsoft's 30 million-plus paid seats, Adobe's $300M Firefly line and Gamma's $100M ARR are real businesses built on this domain. The buyers funding them cannot yet demonstrate the return: 95% of pilots show no measurable profit impact, only 5–8% achieve returns at scale, and 95% of organisations in one survey reported zero measurable ROI despite 80% adoption. Procurement is responding before the vendors do — the "saves time" justification fell 5.8 points among 830 IT decision-makers as CFOs began demanding direct financial evidence.
The verification tax is now measurable, and it falls hardest on the people who most need the tool. Eleven hours saved a week against 6.4 hours returned to checking is a net gain, but a modest one — and it inverts for professionals, with G2's 2,771-review analysis showing the 12% of buyers who write for a living report review overhead cancelling the gain entirely. The floor under this is a measured 8–23% hallucination rate across the six leading legal AI platforms and fabricated citations inside PwC's own client reports, which is why 69% of workers shipping unreviewed output is the number that should worry boards most.
Confidence rises as competence falls, and users cannot detect the crossover. AI advice cutting accuracy from 27% to 9% while nearly tripling confidence is the fortnight's headline mechanism, but it recurs everywhere: education trials producing high satisfaction and zero learning, a BCG experiment showing a 19-point correctness drop outside the model's competence frontier, chain-of-thought traces that misrepresent the actual computation. The practical consequence is that self-reported productivity and satisfaction are unreliable adoption metrics across this entire domain.
Data governance, not model quality, now decides who deploys. Two-thirds of enterprises have delayed or cancelled Copilot over fears it surfaces confidential SharePoint files — a permissions problem, not a capability one. The pattern repeats in accessibility, where 45% of workers report accessibility is absent or unclear in their organisation's AI governance despite 60% believing AI improves accessibility knowledge. Where deployment does succeed, the differentiator is organisational: champion networks outperform mandates roughly two to one.
Shared channels are absorbing the cost of everyone's adoption. Cold email reply rates fell from 0.50% to 0.35% across 2025 as AI-drafted outreach saturated inboxes, and only 1.1% of professionals trust AI to send mail autonomously. Style research quantified the same dynamic elsewhere: homogenisation now measured at 31% trust loss and 52% audience disengagement, with models reverting to generic phrasing within roughly 300 words. Translation escapes this trap for a structural reason — nobody wants their interpreter to have a personal voice — which is exactly why it is the one practice here crossing from adoption into infrastructure.