The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail
AI for turning raw data into queryable, analysable, actionable insight. Streaming analytics, MLOps, and feature engineering are good practice with proven deployments at scale. The bulk sits at leading-edge, held back not by tooling but by data quality and governance gaps — 60% of AI projects stall on data readiness. Nearly all practices are stalled in trajectory.
The headline: The scoreboards vendors cite to prove their AI analytics tools work have been audited, and they are wrong. Judge tools on your own data, not on a leaderboard.
Every major platform — Databricks, Snowflake, Microsoft, Google, Amazon — now sells broadly the same AI analytics, so the difference between companies is no longer which product they bought. It is whether they have written down what their own data means. Where a firm has a governed dictionary of terms like "revenue" and "active customer", AI tools answer business questions correctly around 95 percent of the time; pointed at a raw database, the same tools sit near 70 percent. A small group did that documentation work years ago and is now scaling. Most have not: a survey of 3,044 enterprise accounts found 51 percent do not record what data they hold at all. None of the sixteen practices we track moved position this fortnight, which is itself the finding.
The industry's main accuracy benchmark failed an audit. University of Illinois researchers re-checked BIRD, the test vendors cite when quoting how often their AI turns a plain-English question into a correct database query, and found 52.8 percent of its answers annotated wrong. Five further benchmark studies landed the same fortnight with the same message, including one on real corporate data where systems scoring above 70 percent publicly managed 10 percent. Stop accepting vendor accuracy figures and insist on a trial against your own data before you sign.
A Big Four firm shipped hallucinated client reports. Detection firm GPTZero documented systematic hallucinations — where an AI tool confidently makes things up — in PwC consulting reports: fabricated products, invented government customers, false citations. This landed as Salesforce and Oracle switched on automatic plain-language explanations in Tableau and NetSuite. Require that any AI-drafted number can be traced to a named source before it reaches a client, board or regulator.
EU AI logging rules went live on 2 August, and vendors shipped the plumbing within days. Firms must now automatically log what an AI system inferred and where the data came from; Amazon, Databricks and the open OpenLineage standard all released data-lineage features — automatic records of where each number came from — between 1 and 10 August. Audit testing the same week found these tools miss changes made inside applications, exactly where regulators look, so buying the tool is not the same as passing the audit.
Real deployments kept landing, and none needed a bigger model. Carrefour's data assistant cut support resolution from hours to two minutes for 700–800 staff, with 75 percent handled without a human; a nonprofit reached 94–95 percent query accuracy through a curated data dictionary rather than a more powerful model. Curation is the consistent ingredient, so scope your first project to one well-documented question rather than a broad rollout.
Simpler models beat the expensive ones where it counted. In a 41-day live test on German electricity grid load — safety-critical infrastructure under the EU AI Act — small local statistical models beat foundation models (the largest, most capable current models). A separate post-mortem traced a $240,000 winter-coat overstock to an automated neural network, later replaced with a simpler method. Where you must explain a number, complexity is a liability rather than a credential.
Vendor lineage roadmaps run past your compliance date. Databricks is targeting the fourth quarter of 2026, Snowflake the second half of this year, Google an integration by year end. Your obligation started in August — get those dates written into the contract.
What counts as "anonymous" data has narrowed. European regulators now ask whether a record can be singled out, linked to another dataset or inferred from, and fresh analysis confirms synthetic data alone does not qualify without formal mathematical privacy guarantees. Have counsel re-check which datasets you treat as outside privacy law.
Synthetic data is scaling faster than its safety checks. Thirty-five percent of Fortune 500 firms in regulated sectors report production use, and Gartner projects 75 percent adoption by end-2026, up from under 5 percent in 2023. A clinical study this fortnight found datasets passing standard quality checks can still be 79 percent empty, so set your own acceptance test.
Watching the system costs almost as much as running it. Turning on full monitoring in the most widely used model-tracking tool takes response time from 80 to 770 milliseconds. Separately, 42 percent of security teams run AI alerting untailored to their environment, burning about $1.3 million a year chasing false alarms.
Documentation rots without an owner. Consultants tracking these projects find lineage records go stale within three to six months when nobody is accountable, and the governance role is usually the first budget line cut once the technical work looks finished.
The failure rate is identical everywhere, which points at you, not the tools. Between 65 and 78 percent of large enterprises pilot knowledge graphs; fewer than 15 percent reach full production. Eighty-eight percent report AI adoption while 10 percent actually scale automated modelling.
Go deeper: the full Data & Analytics briefing — the longer analytical write-up, plus every practice we track in this domain with its maturity rating, the tools to consider, and the evidence behind our assessment.