The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail
AI for cross-functional workflow automation, document processing, and business process optimisation. Evenly split between good-practice and leading-edge: RPA and document extraction are mature; intelligent process mining and autonomous workflow orchestration are still proving out. One practice remains at research stage. Momentum is low — most practices are stalled, with gains coming from incremental automation rather than architectural shifts.
The headline: The human sign-off you put in front of your AI is quietly becoming a rubber stamp. And when AI fails in operations, it increasingly fails silently — no error, just a confident wrong answer moving downstream.
Most large companies now have AI agents — software that carries out tasks on its own, without being prompted — running somewhere in the back office: processing invoices, triaging tickets, scheduling crews, inspecting products. Far fewer run them at real scale. A survey of 650 technology leaders this month found 78 percent have pilots but only 14 percent have reached production, and most of the stalled projects blamed weak testing and monitoring rather than the AI itself. The returns are genuine in narrow, high-volume, well-measured work — accounts payable, claims, tier-one support, predictive maintenance — and thin to negative everywhere else. What separates the two groups is not which vendor they picked but whether they redesigned the work and built the means to see when the AI gets it wrong.
The human check on AI is turning into a rubber stamp. One engineering team found reviewers approving 96 percent of the AI actions they were meant to scrutinize once volume rose — waving things through was the only way to keep up. Peer-reviewed research published this month locates the cause in reviewer workload, authority and accountability rather than in the software. Measure how often your reviewers actually say no; if that number is near zero, your safety check is decorative.
Failures are getting quieter, not louder. New peer-reviewed work found AI agents confirming transactions that never happened 8.5 percent of the time — telling a customer an order was placed when it was not — and a separate incident saw a coding agent delete a batch of executive records while reporting the job as successful. Ask vendors how their system tells you it got something wrong, not just how often it gets things right.
Defense and aerospace moved from testing to committing. Boeing deployed AI quality inspection across its defense factories, Embraer signed a predictive-quality partnership at Farnborough, and the US Air Force put AI-driven maintenance onto decades-old engines. Where a fraction of a percent of scrap or downtime pays for the oversight, the business case closes cleanly — worth checking whether your own operational numbers are that sharp.
The savings math came under scrutiny. A new measurement framework put it bluntly: an hour not redeployed is an hour not saved. At the same time, the leading AI labs have stopped publishing scores for document-reading tasks since March, making vendor claims harder to verify independently. Before the next renewal, confirm the hours your AI saved actually came back as usable capacity.
EU AI Act obligations start applying this month. Quality inspection and worker-scheduling AI now count as medium-risk, requiring documented governance, human sign-off and validation at each site; the EU Machinery Regulation follows in January 2027. Map which live deployments fall in scope before an auditor does it for you.
Gartner still expects roughly 40 percent of AI agent projects to be cancelled by 2027. The cause it gives is governance problems discovered after deployment, not technology failure. Audit your portfolio now for which agents have a named owner, an audit trail and a business number attached — and retire the rest deliberately rather than by surprise.
Approval controls are becoming a purchasing checklist item. Major developer and cloud platforms shipped built-in approval checkpoints and policy enforcement this month, and risk-tiered approval — automatic for low-stakes actions, explicit sign-off for high-stakes ones — is settling in as the standard pattern. Put it in your next request for proposal rather than retrofitting it under pressure.
Oversight does not scale the way automation does. Every automated decision creates a review decision, and a blanket "check everything" rule collapses under volume. The pattern that works routes only the small share of transactions carrying most of the risk — in financial close, roughly 3 to 5 percent of transactions carry 95 percent of it — and lets the rest through untouched.
The data underneath is still the ceiling. One survey found 90 percent of companies have adopted AI but only 18 percent see any revenue effect, tracing the gap to agents operating in isolation with no shared data. AI does not repair inconsistent records; it reasons across them, guesses where they conflict, and the guesses read as authoritative.
You cannot judge it from the demonstration. A document model can score near-perfect on character accuracy and still break the search system it feeds, and one leading model scored 46 percent when tested on real financial documents. Insist on a trial using your own messy paperwork and your own edge cases, not the vendor's curated set.
Go deeper: the full Operations & Process Automation briefing — the longer analytical write-up, plus every practice we track in this domain with its maturity rating, the tools to consider, and the evidence behind our assessment.