Perly Consulting │ Beck Eco

The State of Play

A living index of AI adoption across industries — where established practice meets the bleeding edge
UPDATED DAILY

The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.

The Daily Dispatch

A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.

Pick a role above to explore practices

BLEEDING EDGE

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🔬 RESEARCH & KNOWLEDGE
⚖️ LEGAL, COMPLIANCE & RISK
🎧 CUSTOMER OPERATIONS
🏛️ AI GOVERNANCE & SAFETY
📊 DATA & ANALYTICS
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💼 SALES & REVENUE
🎬 CREATIVE & GENERATIVE MEDIA
👁️ COMPUTER VISION & SENSING
💹 FINANCE & ACCOUNTING
🔄 OPERATIONS & PROCESS AUTOMATION
🚗 AUTONOMOUS SYSTEMS & VEHICLES
🦾 PHYSICAL AI & ROBOTICS
🎓 EDUCATION & LEARNING
PERSONAL EFFECTIVENESS

LEADING EDGE

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🔬 RESEARCH & KNOWLEDGE
⚖️ LEGAL, COMPLIANCE & RISK
🎧 CUSTOMER OPERATIONS
🏛️ AI GOVERNANCE & SAFETY
📊 DATA & ANALYTICS
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💼 SALES & REVENUE
🎬 CREATIVE & GENERATIVE MEDIA
👁️ COMPUTER VISION & SENSING
💹 FINANCE & ACCOUNTING
🔄 OPERATIONS & PROCESS AUTOMATION
👥 PEOPLE & TALENT
🚗 AUTONOMOUS SYSTEMS & VEHICLES
🦾 PHYSICAL AI & ROBOTICS
🎓 EDUCATION & LEARNING
PERSONAL EFFECTIVENESS

GOOD PRACTICE

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🔬 RESEARCH & KNOWLEDGE
⚖️ LEGAL, COMPLIANCE & RISK
🎧 CUSTOMER OPERATIONS
🏛️ AI GOVERNANCE & SAFETY
📊 DATA & ANALYTICS
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💼 SALES & REVENUE
🎬 CREATIVE & GENERATIVE MEDIA
👁️ COMPUTER VISION & SENSING
💹 FINANCE & ACCOUNTING
🔄 OPERATIONS & PROCESS AUTOMATION
👥 PEOPLE & TALENT
🚗 AUTONOMOUS SYSTEMS & VEHICLES
🦾 PHYSICAL AI & ROBOTICS
🎓 EDUCATION & LEARNING
PERSONAL EFFECTIVENESS

ESTABLISHED

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💹 FINANCE & ACCOUNTING
👥 PEOPLE & TALENT

🔄 Operations & Process Automation

AI for cross-functional workflow automation, document processing, and business process optimisation. Evenly split between good-practice and leading-edge: RPA and document extraction are mature; intelligent process mining and autonomous workflow orchestration are still proving out. One practice remains at research stage. Momentum is low — most practices are stalled, with gains coming from incremental automation rather than architectural shifts.

14 practices: 6 good practice, 6 leading edge, 2 bleeding edge

Where AI Stands in Operations & Process Automation

Operations is where enterprises have pushed AI agents deepest into the back office, and where the resulting reliability reckoning is now most heavily instrumented. The vendor argument has been settled for a year: agentic automation is table-stakes across UiPath, Automation Anywhere, Microsoft, ServiceNow, SAP, Celonis and Coupa. UiPath's most recent quarter shows 950 companies building agents on its platform and 365,000 processes orchestrated through its Maestro layer, with net new annual recurring revenue up 81% year on year and AI-inclusive deals averaging six times the size of non-AI deals. Automation Anywhere reports more than a billion IT service requests auto-resolved at better than 80% success. SAP now runs over fifty Claude-powered agents making live supply-chain decisions, with 20-30% procurement efficiency gains and 55% scrap reduction attached. The unsettled question is operational rather than technical: whether these systems can be run, watched and corrected inside a real enterprise. Only one area in this domain — multimodal document understanding, where models read tables, charts and diagrams together — is still advancing chiefly on model capability. Everywhere else the binding constraint has moved off the model and onto the organization, and organizations do not upgrade on a release cadence.

The outcome distribution stayed stubbornly bimodal, and this fortnight sharpened the numbers on both tails. Kyndryl's survey of 1,100 global leaders found 57% have deployed AI but only 11% are meeting their objectives; 81% expect agents to make autonomous decisions within a year, yet only 25% would trust them to do so without oversight. A survey of 650 technology leaders put the production cliff at its starkest yet: 78% have pilots running, 14% have scaled to production — a 64-point gap, with 64% of stalled projects naming evaluation and observability, not model quality, as the blocker. In procurement, Ardent Partners and Ivalua found only 23% of organizations have the data foundations to qualify as "AI-first" and just 11% operate a unified procurement data model, despite 90% exploring or piloting. Against that, the value in bounded, well-instrumented processes keeps compounding: Accenture's accounts-payable deployment across 990,000 invoices a year raised its auto-clear rate 300%; Boeing's inspection tooling cuts 17-plus hours per aircraft with 90% first-attempt accuracy against a 50% manual baseline; Denmark's largest wholesaler runs supplier confirmations at over 90% touchless with 98% ERP-matching accuracy; and at semiconductor scale, a three-point yield improvement from AI-driven process control is worth roughly $1.3B a month to a fab the size of TSMC's.

Underneath the headline, the domain continues to stratify into three layers maturing at different speeds. The extraction layer has commoditized to the point of irrelevance as a differentiator: Alibaba open-sourced an 0.8-billion-parameter end-to-end document model that became the first to beat traditional multi-stage pipelines on the standard benchmark, Salesforce folded the open-source Docling parser into its customer data platform, and email classification tops out at 98-100% accuracy across every major vendor at fractions of a cent per message. The orchestration layer has hardened into genuine infrastructure — Temporal's serverless workers reached general availability on AWS Lambda at 150,000 actions per second with Stripe, Netflix, Datadog and Snap named in production, and Microsoft shipped declarative workflows in its Agent Framework, separating orchestration logic from application code. The governance layer is where the domain is stuck, and this cycle it acquired something it previously lacked: a precise, measurable failure taxonomy of its own.

What's New, 2026-07-19 to 2026-08-02

The fortnight's most consequential finding is that human oversight — the control every enterprise has been leaning on — is itself failing at scale, and now has numbers attached. One engineering team processing thirty pull requests a day documented 96% rubber-stamping once approval gates were triggered by action category rather than risk: reviewers approved almost everything, because approving everything was the only way to keep up. A production-AI failure taxonomy published this window formally names "human-in-the-loop collapse" as a measurable failure family, observable before launch through override and rejection-rate instrumentation. Peer-reviewed work from the National Research Council of Canada goes further, showing that reviewer workload, conformity and automation complacency determine oversight effectiveness independently of how good the surrounding infrastructure is — and that most organizations have given their reviewers neither the training, the authority nor the accountability to dissent. A separate industry report found 62% of organizations have no defined process for human review of AI outputs at all, and 18% have rolled back an AI initiative on quality grounds. Distributed-systems analysis added a mundane but damaging mechanism: approval state held in the orchestration runtime is silently discarded when a process redeploys or an autoscaler terminates a worker mid-wait, so the approval that was supposed to happen simply does not. The industry response is already visible and converging: risk-tiered, signal-based approval replacing blanket category gates; BlackLine's financial-close architecture routing the 3-5% of transactions that carry 95% of the risk; Google shipping policy enforcement at the infrastructure layer through an agent gateway; GitLab shipping native approval checkpoints with multi-approver logic; and enterprise governance frameworks drawing on ACSC, CISA and NSA guidance to formalize "approval debt" — the shared service accounts, broad OAuth grants and sandbox drift that pilots accumulate and production inherits.

The second thread is that failures in this domain are increasingly silent rather than loud. Peer-reviewed work presented at ACM UMAP 2026 documents "journey hallucinations": agents confirming transactional state that never occurred — telling customers an order is placed when it is not — at an 8.5% false-positive rate, which then triggers the wrong downstream workflow. Complementary arXiv research names operational hallucination (repetitive flawed tool calls) and safety drift (alignment eroding over long multi-turn sessions) as distinct failure modes. A Replit agent deleted 1,200 executive records through a missing approval gate and returned HTTP 200 — success — while doing it. The same pattern recurs across unrelated practices: predictive-maintenance models lose 20-40% of their precision within six months of sensor drift while continuing to report 99% confidence; handwriting recognition exhibits "fluent rewriting," silently substituting plausible but wrong text; and a SIGIR workshop taxonomy catalogues phantom grounding and provenance hallucination in multimodal agentic search. Alongside this, the industry's own measurements came under scrutiny. An ACL 2026 industry-track benchmark showed that high character-level OCR accuracy does not predict downstream retrieval success, because structural and semantic errors survive low error rates; GPT-4o scored 46% on real financial document tasks in the MultiFinBen evaluation; frontier labs have stopped publishing quantified document-understanding scores since March 2026; and an ROI-governance framework for document automation landed the sharpest line of the cycle — "an hour not redeployed is an hour not saved." On maturity, nothing moved. Every practice in the domain held its position, and the one still advancing on capability is still the same one. That stability is the signal: the capability curve keeps climbing while adoption stays pinned against an organizational wall that better models do not touch.

Key Tensions

  • Human oversight is failing as a control mechanism, not as a principle. The fix everyone reached for after last quarter's inspection failures — put a person in front of every output — does not survive volume. One team's approval gates produced 96% rubber-stamping; peer-reviewed research attributes oversight effectiveness to reviewer workload, authority and accountability rather than tooling; and 62% of surveyed organizations have no defined human-review process to degrade in the first place. The emerging answer is not more review but tiered review: route the small fraction of decisions carrying most of the risk, and instrument override rates so you can see when reviewers stop reading.

  • Silent failure, not visible error, is the dominant production risk. Agents falsely confirm transactions at 8.5%; a coding agent deleted 1,200 records and reported success; maintenance models keep reporting 99% confidence as sensor drift strips 20-40% of their precision; document models rewrite text fluently and wrongly. In every case the system does not raise an error — it produces a confident, plausible, incorrect output that flows straight into a downstream workflow. This is why observability and evaluation, not accuracy, is what 64% of stalled deployments name as their blocker, and why detection architecture is now a harder procurement question than capability.

  • The measurements systematically overstate readiness. Benchmark scores do not predict production behavior: high OCR character accuracy does not deliver working retrieval, a leading model scores 46% on real financial documents, and independent testing puts large language models at roughly 12% success on end-to-end scheduling tasks against classical solvers that carry correctness guarantees. ROI accounting is equally soft — recovered hours that are never redeployed are not savings, and frontier labs have quietly stopped publishing the document benchmark numbers they once led with. Organizations buying on demo performance are buying a number that does not transfer.

  • The data substrate is still the wall, and agents make it worse rather than better. A survey pairing 90% AI adoption against 18% revenue impact traced the gap directly to half of deployed agents operating in isolation with no shared data. Three-quarters of UK and Ireland manufacturers still run on legacy systems or spreadsheets; only 11% of procurement teams operate a unified data model. Process mining, the discipline meant to reveal how work actually happens, is itself blind to the 80-90% of enterprise work conducted in email, chat and calls. Agents do not repair fragmented data — they reason across it, guess where it conflicts, and their guesses carry unwarranted authority.

  • Value is concentrating in capital-intensive industry just as the regulatory clock starts. Defense, aerospace and semiconductors moved decisively this cycle: Boeing deployed Palantir Foundry for real-time quality anomaly detection across defense factories, Embraer partnered with Hexagon on predictive quality using 40-micron laser metrology, and the US Air Force put AI reliability-centered maintenance onto legacy propulsion systems. These are sectors where a fraction of a percent of yield or downtime justifies the governance investment outright. Meanwhile the EU AI Act's obligations land this month, classifying quality inspection and worker scheduling as medium-risk and requiring documented governance and facility-specific validation — with the EU Machinery Regulation following in January 2027 and the revised automotive statistical process control standard already tripling its mandatory controls. The organizations that built governance for commercial reasons will absorb this; the ones that did not now face it as a deadline.

Top 10 Evidence Items

  1. 30 PRs Daily: Why HITL Approval Gates Break at Scale (practitioner case analysis) — The single best-documented instance of this fortnight's central finding: category-based approval gates producing 96% rubber-stamping, the exact mechanism behind the summary's lead tension that human oversight is failing as a control, not just as a principle. https://waxell.ai/blog/human-in-the-loop-approval-fatigue-policy-fix

  2. What an AI Agent Production Incident Actually Looks Like (case study) — A Replit agent deleted 1,200 executive records through a missing approval gate and returned HTTP 200, the starkest available illustration of "silent failure, not visible error" as the domain's dominant production risk. https://tessary.ai/blog/ai-agent-production-incident-response

  3. AI Agents Are Confirming Orders That Were Never Placed (peer-reviewed research, ACM UMAP 2026) — Names and quantifies "journey hallucinations" at an 8.5% false-positive rate, giving the silent-failure thesis a formal, measured failure mode rather than an anecdote. https://industrycontents.com/ai-agent-journey-hallucinations/

  4. Realizing the Promise of AI Governance Involving Humans-in-the-Loop (peer-reviewed research, National Research Council of Canada) — Academic grounding for the claim that oversight effectiveness depends on reviewer workload, conformity and accountability rather than tooling quality — the mechanism underneath the rubber-stamping problem. https://nrc-publications.canada.ca/eng/view/object/?id=a4d815b0-eab9-432f-9f43-95b5ad52ff24

  5. The AI Implementation Cliff (survey-backed analysis, 650 technology leaders) — Source of the domain's starkest adoption statistic this cycle — 78% running pilots versus 14% scaled to production, with 64% of stalled projects blaming evaluation and observability rather than model quality — the number anchoring the summary's opening paragraph. https://www.themindfinders.com/2026/07/20/the-ai-implementation-cliff-what-happens-between-success-in-sandbox-and-failure-in-reality/

  6. New Research by Ardent Partners and Ivalua Reveals Gap between AI Ambition and Execution in Procurement (industry survey) — Documents the data-substrate wall inside a single function: only 23% of organizations have AI-first data foundations and just 11% run a unified procurement data model despite 90% exploring or piloting. https://finance.yahoo.com/technology/ai/articles/research-ardent-partners-ivalua-reveals-140000835.html

  7. Accenture: Accentuating Digital Transformation with Cloud-Based and AI-Driven Capabilities (case study) — The counterweight to the failure evidence: a bounded, well-instrumented process (990,000 invoices a year) delivering a real 300% auto-clear improvement, showing the positive tail of the domain's stubbornly bimodal outcome distribution. https://www.sap.com/asset/dynamic/2025/11/541ac8fd-2d7f-0010-bca6-c68f7e60039b.html

  8. U.S. Boeing selects Palantir AI to speed weapons output and strengthen defense programs (case study) — Direct evidence for the closing tension: value concentrating in capital-intensive, high-governance sectors just as the EU AI Act's obligations land this month. https://www.armyrecognition.com/archives/archives-aerospace-defense/defense-news-aerospace-2025/boeing-selects-palantir-ai-to-speed-weapons-output-and-strengthen-sensitive-defense-programs

  9. Alibaba Open-Sources OvisOCR2: End-to-End Model First Exceeds Pipeline Methods (product release) — Makes concrete the extraction layer's commoditization: a 0.8-billion-parameter open-source model beating traditional multi-stage pipelines, reinforcing that multimodal document understanding is the one practice still advancing chiefly on model capability. https://www.53ai.com/news/OpenSourceLLM/2026072413568.html

  10. How AI Improves Semiconductor Yield and Die-per-Wafer Economics (industry analysis) — Puts a hard number on why capital-intensive industry absorbs governance costs the rest of the domain balks at: a three-point yield improvement is worth roughly $1.3B a month at TSMC scale. https://www.softwebsolutions.com/resources/semiconductor-yield-optimization-using-ai/