The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
AI for turning raw data into queryable, analysable, actionable insight. Streaming analytics, MLOps, and feature engineering are good practice with proven deployments at scale. The bulk sits at leading-edge, held back not by tooling but by data quality and governance gaps — 60% of AI projects stall on data readiness. Nearly all practices are stalled in trajectory.
The tooling argument in enterprise analytics is over. Databricks, Snowflake, Microsoft, Google and Amazon all ship broadly comparable AI analytics: autonomous query agents, embedded profiling, managed forecasting, native anomaly detection, catalog-aware metadata services. The routine plumbing of the discipline — data cataloguing and lineage, quality automation, AutoML and feature engineering, experiment tracking and model monitoring, real-time streaming — has settled into established good practice, deployed at scale with defensible economics. What sits above it, the interpretive layer where AI is supposed to earn its keep, has not. Natural-language querying, narrative generation, causal inference, automated exploratory analysis, geospatial AI, graph analytics, synthetic data and time-series forecasting all remain leading-edge: proven at named organisations, unavailable to most. General-purpose anomaly detection is the one practice still genuinely unsolved at scale, and the retirement of Azure Anomaly Detector in October, following AWS Lookout for Equipment, is the ecosystem admitting it. Across sixteen practices, exactly one — model monitoring — is gaining ground. The rest are stalled or plateaued. That is not vendor failure. It is the signature of a domain whose binding constraint has moved from technology to organisation.
The constraint has a name and a number. Every practice in this domain now converges on the same finding from a different angle: AI analytics performance is a function of how well a company has written down what its own data means. dbt Labs' benchmark puts frontier models at 84.1 percent SQL accuracy against a raw schema and 100 percent against a governed semantic layer. Data agents plateau around 70 percent accuracy without curated metadata and reach the 95 percent production threshold with it. Semantic context adds 17 to 23 accuracy points on enterprise schemas. This is the closest thing to a law that the domain has produced, and it is unhelpful to most buyers, because the input is documentation discipline rather than a product. Snyk's survey of 3,044 enterprise accounts found that 51 percent declare no dataset in their repositories at all — no lineage, no provenance, nothing to govern. Cloudera's survey of 1,500 enterprise architects found 95 percent had delayed or cancelled AI initiatives despite 77 percent actively using AI, with data architecture cited as the primary cause and 66 percent moving workloads back to private cloud. RAND's meta-analysis of more than 2,400 initiatives attributes 77 percent of failures to governance and strategy rather than technology.
Yet the domain is not stuck, and it would be a mistake to read stalled trend lines as absent value. The deployments that landed in the last fortnight are large, specific and unglamorous. Uber detailed Tarot, a production targeting orchestrator combining uplift models with multi-treatment constrained optimisation across Mobility and Delivery at millions-of-users scale, and separately published an exactly-once ad event pipeline for UberEats across Flink, Kafka and Pinot. Taobao ran a fourteen-day test of multi-channel uplift policy learning across 300,000 items for a 3.53 percent lift in paid orders and 3.26 points of margin. Carrefour's retrieval-grounded data assistant took support resolution from hours to two minutes for 700–800 users at 75 percent self-service. Apache Fluss graduated to top-level project status with production use at Alibaba, JD.com, Ant Group, Xiaohongshu, Fresha and iQiyi handling hundreds of billions of events. A reference supply-chain knowledge graph now runs at 460 billion nodes and 1.1 trillion edges. The pattern across all of them is identical: a narrow, heavily instrumented problem inside an organisation that had already done the definitional work. Value in this domain is not distributed evenly across enterprises; it is concentrated in the ones that were disciplined before the models arrived.
The defining development of this fortnight is that the industry's measurement layer failed audit, in public, from six directions at once. A University of Illinois team's VLDB 2026 audit of BIRD — the benchmark most widely cited for text-to-SQL accuracy — found 52.8 percent annotation errors, and determined that 19 percent of what had been recorded as model failures were in fact improvements over flawed gold answers. Michael Stonebraker's BEAVER benchmark, run against real corporate query logs, put a plain language model at 0 percent execution accuracy, a retrieval-augmented agentic system at 10 percent, and an oracle-joins ceiling at 30 percent, against the 70-plus percent those same systems post on public benchmarks. An ICPR 2026 evaluation of seven anomaly detection algorithms across 690 datasets found rankings highly unstable and driven mainly by dataset choice and hyperparameter configuration, destabilising a decade of comparative claims. AutoML benchmarking analysis showed that correcting for test-set leakage and unenforced time budgets cut reported win rates from 59.4 percent to 34.3 percent. In forecasting, a practitioner evaluation of thirteen models across ten datasets and four horizons found foundation models winning 30 of 38 contests but with zero transfer between contexts — and GIFT-Eval's top-ranked model, Toto-2.0, winning none of the 38 real-world cells. GISAgentBench put the best language agent at 32.7 percent strict accuracy on 349 practitioner-sourced GIS tasks; DataClawEval put the best of sixteen frontier agents at 74.9 percent on real production data-engineering work; the KDD Cup DataSpace benchmark capped data agents at 66.34 percent across 410 heterogeneous tasks. Individually these are papers. Together they are a verdict: public leaderboards in this domain do not predict enterprise performance, and in at least two cases the leaderboards are themselves measurably wrong.
The second thread was regulatory and infrastructural, and it moved unusually fast. EU AI Act obligations requiring automatic logging of AI inference and provenance took effect on 2 August. Within nine days the vendor ecosystem shipped the corresponding plumbing: Amazon Quick's agentic catalog experience reached GA on 1 August with automatic upstream catalog discovery and metadata inheritance for AI-generated datasets; Databricks made Unity Catalog AI Gateway lineage GA on 6 August, capturing model-service lineage with upstream foundation models and downstream inference tables; OpenLineage 2.0 reached GA on 10 August with multi-year vendor commitments, Databricks targeting Delta Live Tables lineage emission in Q4 2026, Snowflake a Metadata Connector preview in H2, Google a BigQuery integration by year end. Lineage has stopped being audit paperwork and become the control plane for agent governance. Against that, regulatory audit testing in the same window documented systematic failure: most catalog tools infer lineage from query logs and miss the application-layer transformations regulators actually scrutinise. Elsewhere, a GPTZero investigation documented systematic hallucinations in PwC consulting reports — fabricated products, invented government customers, false citations — a reminder that the failure mode is client-facing, not academic, precisely as Salesforce shipped GA narrative explanations in Tableau Pulse and Oracle GA'd narrative close insights in NetSuite 2026.2. No practice changed tier or trend this cycle.
The evaluation layer is no longer trustworthy, and procurement has not caught up. Six independent results in two weeks — the BIRD audit's 52.8 percent annotation error rate, BEAVER's 10 percent against 70-plus, ICPR's unstable anomaly rankings across 690 datasets, corrected AutoML win rates falling from 59.4 to 34.3 percent, GIFT-Eval's top forecaster winning none of 38 real cells, GISAgentBench's 32.7 percent — converge on the conclusion that published accuracy claims in this domain are not evidence of enterprise performance. Buyers who evaluate on leaderboards are pricing a number that has been shown not to transfer. The only defensible test is a bake-off on the buyer's own schemas, and few procurement processes are built for it.
Metadata became the AI control plane before most enterprises built any metadata. The regulatory clock started on 2 August, the vendors shipped lineage GA within the week, and Snyk found 51 percent of 3,044 enterprise accounts declare no dataset at all. Worse, capability and compliance are not the same thing: audit testing showed catalog tools reconstruct lineage from query logs and miss application-layer transformations, and consultant analysis found that lineage without named ownership goes stale within three to six months — with the governance layer routinely the first line cut from a budget despite the technical capture being complete.
Observability has a measurable price, and nobody has priced it. MLflow tracing at production scale takes baseline latency from 80 milliseconds to 770 milliseconds, forcing an explicit architectural trade-off between monitoring density and response time. The same economics bite elsewhere: a Panther analysis found 42 percent of security teams deploy AI anomaly tooling without environment-specific customisation, burning roughly 395 hours a week on false alerts, about $1.3 million a year. Watching the system is becoming a comparable cost line to running it, and it is the line that gets cut first.
Where accountability is real, simpler methods keep winning. A 41-day live challenge on German transmission-grid load — safety-critical infrastructure under the EU AI Act — saw locally trained statistical models beat Chronos-2 and other 100-million-parameter-plus foundation models. Fine-tuning Kronos on Turkish equities produced no advantage over a naive baseline and made hourly error worse, from 1.79 to 2.01 percent MAPE. A demand-forecasting autopsy traced a $240,000 winter-coat overstock to an automated neural network later replaced by a boosted tree. Architectural ambition is not a defence when someone has to explain the number.
The pilot-to-production gap is now a governance metric, measured identically across unrelated practices. Between 65 and 78 percent of large enterprises pilot knowledge graphs and fewer than 15 percent reach full production; 88 percent of organisations report AI adoption while only 10 percent scale AutoML systems; just 9 percent of geospatial practitioners are building with AI agents, with data quality and provenance named as the top blocker; 88 percent of enterprise AI analytics pilots reportedly never reach production. Different tools, different vendors, different verticals, the same failure rate — which is a strong signal that the variable being measured is the organisation, not the technology.
Half of BIRD Text-to-SQL Benchmark Gold Answers Are Wrong, UIUC Study Finds (news-coverage) — The lead exhibit in the fortnight's measurement-layer audit: the benchmark most cited for text-to-SQL accuracy is itself 52.8 percent wrong, meaning a fifth of recorded "model failures" were actually improvements over flawed gold answers. https://gyaansetu.com/ai/half-of-bird-text-to-sql-benchmark-gold-answers-are-wrong-uiuc-study-finds
Azure AI services retiring Oct 2026: migration playbook (news-coverage) — Microsoft's retirement of Azure Anomaly Detector, following AWS Lookout for Equipment, is the vendor ecosystem's own admission that general-purpose anomaly detection remains the one practice in this domain genuinely unsolved at scale. https://ecorpit.com/azure-anomaly-metrics-personalizer-retirement-migration-2026/
Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons (research-paper) — Correcting for test-set leakage and unenforced time budgets cut reported AutoML win rates from 59.4 to 34.3 percent, one of six independent results this cycle showing published accuracy claims don't transfer to enterprise conditions. https://arxiv.org/abs/2608.07303
GISAgentBench: A Practitioner-Sourced Benchmark for Evaluating LLM Agents on GIS Tasks (research-paper) — The best language agent hit only 32.7 percent strict accuracy on 349 real GIS tasks, extending the measurement-layer crisis into geospatial AI and reinforcing why just 9 percent of practitioners are building agents into production workflows. https://arxiv.org/abs/2608.01645
OpenLineage 2.0 standard wins major vendor backing (product-ga) — Shipped nine days after EU AI Act inference-logging obligations took effect on 2 August, with multi-year commitments from Databricks, Snowflake and Google — lineage has stopped being audit paperwork and become the control plane for agent governance. https://contentwave.net/article/openlineage-20-wins-major-vendor-backing-raises-the-bar-for-trainingdata-lineage
Enterprises Are Blind to Two-Thirds of Their Own AI Attack Surface (adoption-metric) — Snyk's survey of 3,044 enterprise accounts found 51 percent declare no dataset in their repositories at all, the governance-gap number that explains why lineage tooling arrived faster than the metadata it's meant to govern. https://finance.yahoo.com/technology/ai/articles/enterprises-blind-two-thirds-own-120000640.html
Solving the Multiple Knapsack Problem at Scale (case-study) — Uber's Tarot orchestrator, combining uplift models with multi-treatment constrained optimisation across Mobility and Delivery at millions-of-users scale, is the domain's clearest example of value concentrating in organisations that did the definitional work before the models arrived. https://www.uber.com/dk/da/blog/solving-multiple-knapsack/
From repetitive queries to instant SQL: Building Carrefour's internal data assistant in a weekend (case-study) — A retrieval-grounded assistant took support resolution from hours to two minutes for 700-800 users at 75 percent self-service, a narrow, well-instrumented deployment that stands in sharp contrast to the sector's public benchmark failures. https://discuss.google.dev/t/from-repetitive-queries-to-instant-sql-building-carrefour-s-internal-data-assistant-in-a-weekend/387879
Chasing the Hallucinations: PwC report hallucinates product and government customers (opinion) — A GPTZero investigation into fabricated products, invented government customers and false citations in PwC consulting reports shows the narrative-generation failure mode is client-facing, not academic, arriving in the same window vendors GA'd automated narrative insights in Tableau and NetSuite. https://gptzero.me/news/investigations-pwc/
Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments (research-paper) — In a 41-day live challenge on German transmission-grid load, locally trained statistical models beat Chronos-2 and other 100-million-parameter-plus foundation models, the domain's sharpest instance of the rule that where accountability is real, simpler methods keep winning. https://arxiv.org/abs/2608.05018