📊 Data & Analytics
AI for turning raw data into queryable, analysable, actionable insight. Streaming analytics, MLOps, and feature engineering are good practice with proven deployments at scale. The bulk sits at leading-edge, held back not by tooling but by data quality and governance gaps — 60% of AI projects stall on data readiness. Nearly all practices are stalled in trajectory.
The Headline
The tools all work now. Almost nobody's data does. The companies getting returns spent their money writing down what their business terms mean, not on better AI.
The Picture
Every major data platform now ships conversational querying, automated quality checks and pipeline generation as standard. The buying decision is effectively over, and it is no longer what separates companies. What separates them is whether anyone has written down, in a form software can execute, what "revenue" or "active customer" actually means in your business. Accenture puts the share of enterprises with genuinely AI-ready data foundations at 7 percent; Teradata's survey of a thousand senior leaders found 77 percent say their data is not ready for AI that acts on its own without being prompted. The organizations posting hard numbers this fortnight — HP, Unilever, Comcast, CVS Health — all did that definitional work first. Everyone else is running pilots on data nobody has defined, and getting confident wrong answers that look exactly like right ones.
This Fortnight
HP scrapped a natural-language query system it had built in-house and deployed a governed off-the-shelf tool instead, reporting 40 to 50 percent efficiency gains. It had spent considerable time and money on the build. The lesson is not that the vendor was smarter — it is that HP curated its data estate first, and that is the part no vendor sells. If you have a build-versus-buy paper in flight, the governance work sits on both sides of it.
A peer-reviewed benchmark recorded zero hallucinations — invented answers stated with confidence — only when queries ran through governed business definitions. Without that layer, the same systems failed across six separate categories. Miro separately lifted query accuracy from under 40 percent to over 90 percent through governance and context work alone, without changing the AI model. Your reliability lever is your data dictionary, not your model contract.
dbt released incremental pipeline builds as a generally available product, with one customer cutting its warehouse compute bill 59 percent. RxBenefits posted that 59 percent reduction, Virgin Media O2 25 percent, with 15 to 30 percent typical across early adopters. This moves data-engineering hygiene out of the engineering preference column and into a line item your finance team can see and audit.
Independent studies found general-purpose AI forecasting models losing to older, simpler statistical methods. In one, a benchmark on 20,330 real orders found the models ranked best on forecast accuracy performed worst on the thing that actually matters — whether the order got served. Treat vendor accuracy claims in forecasting with the skepticism you would apply to a backtested trading strategy.
More standalone data tools shut down or went quiet, including the archiving of a widely used open-source data catalog. Great Expectations closed its commercial cloud product in June with about thirty days' notice; Microsoft's knowledge-graph project is now bug-fixes-only; Confluent stripped a streaming component out of its own console. The middle tier of point tools is being absorbed into the big platforms or narrowed to regulated niches, so it is worth auditing which of yours still has a viable vendor behind it.
Coming Up
Microsoft retires Azure Anomaly Detector on 1 October, one of several first-generation AI services being switched off this year. AWS Lookout for Equipment, the synthetic-data vendor MOSTLY AI and Gretel's self-serve tier have all gone the same way. Ask your data team for a list of every AI service running in production with a published end-of-life date, and what replaces each one.
European regulators have proposed replacing the yes-or-no test for anonymized data with a context-dependent one, with consultation closing 30 October. The draft also narrows the processor exemption that most software-as-a-service contracts lean on when they promise your data is only used "aggregated or anonymized." Your general counsel should be reading your ten largest vendor agreements against this before the year turns.
Boards are about to start asking what the AI budget bought, and most organizations cannot answer. Tempo.io found 91 percent of technology leaders unable to link AI work to business outcomes while AI consumes 20 to 30 percent of research and development budgets, and supply-chain benchmarks show 70 percent of deployments contributing nothing to profit. Agree now, in writing, on the two or three metrics each deployment is meant to move.
What's Hard About This
The asset that makes these tools work cannot be bought from anyone. Anthropic's published internal figures show its own analytics reaching 95 percent accuracy with a curated layer of business definitions and verification checks in place, and 21 percent without — maintained by roughly thirty analysts, about 2 percent of headcount. That is a permanent staffing line, not a project with an end date.
The governance bill arrives after the pilot succeeds, and it is usually the first thing cut. Tagging a 40,000-table estate for column-level access control implies around 10,000 hours of stewardship, and enterprise knowledge-graph projects carry a $10 million to $20 million modeling cost with fewer than 15 percent reaching production. Documentation of where data came from goes stale within three to six months without a named owner, so budget the maintenance or expect quiet decay.
What works keeps being narrow and hand-built, which is precisely what does not scale. In independent benchmarking, no general-purpose model consistently beat carefully tuned conventional methods on anomaly detection, and plain ridge regression with good preprocessing matched far more complex models on six of eight standard forecasting benchmarks. Plan to win one governed domain at a time rather than buying a platform that promises to work everywhere at once.
Practices in this Domain (16)
| PRACTICE | TIER | TREND |
|---|---|---|
| Anomaly & outlier detection | BLEEDING EDGE | — Steady |
| Automated exploratory data analysis | LEADING EDGE | — Steady |
| Causal inference & uplift modelling | LEADING EDGE | — Steady |
| Data catalogue, metadata & lineage management | GOOD PRACTICE | — Steady |
| Data pipeline, dashboard & report generation | LEADING EDGE | — Steady |
| Data privacy & anonymisation automation | LEADING EDGE | — Steady |
| Data quality, cleaning & transformation automation | GOOD PRACTICE | — Steady |
| Feature engineering, AutoML & predictive modelling | GOOD PRACTICE | — Steady |
| Geospatial data analysis & visualisation | LEADING EDGE | — Steady |
| Graph analytics & relationship discovery | LEADING EDGE | — Steady |
| MLOps — experiment tracking & model monitoring | GOOD PRACTICE | — Steady |
| Narrative generation from data | LEADING EDGE | — Steady |
| Natural language data querying & semantic search | LEADING EDGE | — Steady |
| Real-time streaming analytics | GOOD PRACTICE | — Steady |
| Synthetic data generation | LEADING EDGE | — Steady |
| Time series forecasting | LEADING EDGE | — Steady |
Read the full technical briefing (1,691 words) →
Where AI Stands in Data & Analytics
Data and analytics is the domain where the rest of the AI stack's dependencies become visible, and the picture it presents is uncomfortable. The tooling question is settled almost everywhere. Every major platform ships conversational analytics, automated profiling, lineage capture, drift monitoring and pipeline generation as standard. Databricks, Snowflake, Google and AWS have absorbed what were once separate product categories. And yet the binding constraint has not moved in two years: organisations cannot get their data into a state where any of it works reliably. Accenture puts the share of enterprises with genuinely AI-ready data foundations at 7 per cent. Teradata's survey of a thousand senior leaders found 77 per cent reporting their data is not ready for agentic AI and only 37 per cent seeing measurable business impact. dbt Labs reports just 16 per cent of companies running agentic AI at enterprise scale, with 80 per cent blocked by data access and 70 per cent lacking governance trust.
What has changed is that the fix is now precisely specified, and it is not a model. It is the semantic layer — a governed, deterministic definition of what business terms mean and how they compile to queries. dbt Labs' 2026 benchmark puts governed semantic-layer accuracy at 98 to 100 per cent against 84 to 90 per cent for raw text-to-SQL using identical models. Anthropic's published internal architecture reaches 95 per cent accuracy on business queries with semantic, skills and verification scaffolding in place; the same model without that scaffolding manages 21 per cent, and the scaffolding is maintained by a data team of roughly thirty analysts, about 2 per cent of headcount. Miro lifted query accuracy from under 40 per cent to over 90 per cent through governance and context engineering alone, without changing models. HP abandoned a home-built text-to-SQL system it had spent considerable time and money on, deployed Databricks Genie against a governed estate, and reported 40 to 50 per cent efficiency gains. The lesson repeats across Comcast, Thales, Atlassian, AngelList and CVS Health: the deployments that work are the ones where somebody did the curation work first.
The consequence is a domain moving in two directions at once. Vertical, governed, narrowly-scoped deployments are producing hard numbers — Unilever cutting pipeline costs 25 per cent and accelerating them two to five times on Spark Declarative Pipelines; RxBenefits taking 59 per cent off its Snowflake bill with dbt State; Zepto reporting 52x ROI on MLflow-traced LLM evaluation across 80,000 daily tickets; JPMorgan's OmniAI saving $2bn on anti-money-laundering false positives. Meanwhile the horizontal, domain-agnostic tier is closing. Great Expectations shut its commercial cloud product in June with about thirty days' notice. Amundsen was archived this month. Microsoft's GraphRAG repository sits in maintenance mode. Azure Anomaly Detector retires on 1 October. MOSTLY AI ceased operations before being acquired by Syntho; Gretel's self-serve tier is gone. Even the streaming incumbents are pulling back: Confluent removed Kafka Streams from its own Control Center, cutting startup from fifty minutes to one and lifting partition scaling from 120,000 to 400,000, and Kestra 2.0 stripped it from core entirely. Experiment tracking and model monitoring is the one practice here with genuine forward momentum — MLflow at 60 million monthly downloads, Kubeflow through CNCF graduation, managed monitoring reaching general availability at both AWS and Google — and that is not a coincidence. Monitoring is where the value of everything upstream gets measured, and organisations have finally noticed that 91 per cent of deployed models degrade without it.
What's New, 2026-09-09 to 2026-09-23
This was a two-week window, and the sharpest movement was in evidence quality rather than capability. Three independent results tightened the case that governance, not model choice, is the reliability lever: a peer-reviewed benchmark (GROUND) recorded zero hallucinations only under governed semantic layers, failing across six categories without them; Miro's context-engineering result showed a doubling of accuracy with no model change; and HP's abandonment of its in-house text-to-SQL build in favour of a governed Genie deployment gave the argument a named casualty. dbt State reached general availability with quantified compute savings — 59 per cent at RxBenefits, 25 per cent at Virgin Media O2, 15 to 30 per cent on average — making incremental, metadata-aware builds a cost story rather than an engineering nicety. Five further named Genie deployments landed in a fortnight (Grupo Panvel, Webmotors, FinThrive, Transferz, The AA), all reporting the same pattern: metadata quality and phased curation drive adoption, not model selection.
Running against that, a cluster of results pushed back hard on foundation-model universality. inovex benchmarked tabular foundation models (TabPFN, FoMo-0D, AnoLLM) for anomaly detection and found none consistently beats grid-search-tuned classical detectors. A contamination-free temporal hold-out study traced time-series foundation-model wins to pretraining corpus overlap rather than generalisation — the 28 per cent MASE gain isolated to Wikipedia, where TimesFM trained. A decision-aware benchmark on 20,330 real orders found forecast accuracy rank correlates negatively (−0.555) with order-service outcomes, which is to say the field has been optimising the wrong quantity. Zero-shot foundation models failed to beat elastic net on continuous glucose monitoring across eight datasets. Elsewhere, cost and risk got priced: Unity Catalog's column-level access tagging implies roughly 10,000 steward-hours for a 40,000-table estate with no inheritance; the EDPB's draft Guidelines 02/2026 replace binary anonymisation with context-dependent relative identifiability and narrow the processor exemption most SaaS contracts rely on; a Dgraph production incident lost data on 6,500 companies after a failed backup restore. New general availability arrived at pace — Google Spanner Graph, Neo4j Virtual Graph (zero-copy knowledge graphs over warehouses, no ETL), OpenAI's Data agent in ChatGPT Work, Oracle Select AI natural-language actions, Confluent's AI inference functions in Flink SQL, Databricks AI Functions and Unity Catalog trace redaction — but none of it addresses the curation bill.
Key Tensions
The semantic layer has become the product. Deterministic compilation of governed metric definitions now outperforms probabilistic query generation by 8 to 16 percentage points on identical models, and the gap widens on multi-hop joins. Anthropic's internal analytics reach 95 per cent accuracy with full scaffolding against 21 per cent without, maintained by roughly thirty analysts. The uncomfortable implication is that the differentiating asset is curated business context, which no vendor can sell you and no model can infer.
Nobody can price the returns. Tempo.io found 91 per cent of technology leaders unable to tie AI work to business outcomes despite AI consuming 20 to 30 per cent of R&D budgets. Supply-chain benchmarks show 70 per cent of deployments contributing nothing to EBIT, with median realised ROI of 10 per cent against a 20 per cent target. Protiviti recorded financial-sector adoption rising from 58 to 76 per cent year-on-year while only 35 per cent measure returns effectively.
Pretrained generality keeps losing to tuned specifics. Independent benchmarking found no tabular foundation model beats grid-search-tuned classical detectors on unsupervised anomaly detection, and a contamination-free hold-out attributed time-series foundation-model gains to pretraining corpus overlap rather than generalisation. Ridge regression with careful preprocessing matches Transformers on six of eight standard forecasting benchmarks. Domain-specific engineering keeps winning, which is precisely what does not scale.
The governance bill arrives after the pilot. Attribute-based access control at Unity Catalog requires column-level tagging with no inheritance: a 40,000-table estate implies a million decisions and roughly 10,000 steward-hours. Atlan puts the ontology tax on enterprise knowledge graphs at $10m to $20m, with fewer than 15 per cent of pilots reaching production. Lineage without ownership governance goes stale within three to six months, and the governance layer is commonly the first budget cut.
Standalone tooling is being squeezed out of the middle. Great Expectations shut GX Cloud with thirty days' notice, Amundsen was archived, MOSTLY AI ceased trading, Microsoft's GraphRAG went to maintenance mode and Azure Anomaly Detector retires next month. Confluent and Kestra both removed Kafka Streams from their own products. What survives is either absorbed into a hyperscaler platform or narrowed to a regulated vertical where the customisation is the value.
Top 10 Evidence Items
- GROUND: Semantic layers eliminate hallucinations in enterprise analytics—zero measured failures vs six categories without governance (research-paper) — The GROUND benchmark is the strongest peer-reviewed evidence that governance, not model choice, eliminates hallucinations — the briefing's central mechanism. https://techpulse.ro/en/news/items/ground-reducing-hallucinations-in-llm-based-enterprise-analytics-through-governed-semantic-definitions-23xnid
- HP: Genie production deployment achieving 40–50% efficiency gains across multiple analytics teams; replaced abandoned custom text-to-SQL solution (case-study) — HP's abandonment of its home-built text-to-SQL system for a governed Genie deployment is the named casualty that gives the semantic-layer argument teeth. https://www.databricks.com/customers/hp/genie
- Context Engineering for AI Agents at Scale (opinion) — Miro's doubling of accuracy through context engineering alone, with no model change, is the clearest demonstration that curation beats capability. https://datahub.com/blog/context-engineering-for-ai-agents/
- dbt State: Incremental pipeline builds with 59% cost reduction (RxBenefits) and 25% savings (Virgin Media O2) (product-ga) — dbt State's quantified compute savings turn incremental, metadata-aware builds into a cost story, not just an engineering nicety. https://www.getdbt.com/blog/dbt-state-is-ga
- Teradata survey: 77% of enterprises report data unready for agentic AI (adoption-metric) — The Teradata survey supplies the uncomfortable headline stat that most enterprises still lack agentic-AI-ready data despite the tooling being settled. https://kesq.com/stacker-ai/2026/09/09/companies-keep-spending-on-ai-despite-roadblocks-on-returns/
- ABAC scaling barrier: 40,000-table estate requires ~10,000 steward labour-hours for PII tagging (opinion) — The Unity Catalog ABAC tagging bill quantifies the governance cost that arrives after the pilot and rarely gets budgeted. https://lovelytics.com/post/databricks-abac-policies-someone-still-has-to-tag-40000-tables/
- DataKitchen 2026: GX Cloud shutdown signals vendor consolidation and open-source fragmentation (opinion) — Great Expectations' abrupt shutdown is the sharpest evidence that the horizontal, domain-agnostic tooling tier is being squeezed out. https://datakitchen.io/blog/the-2026-open-source-data-quality-and-data-observability-landscape/
- Anomaly detection: Do tabular foundation models help? (opinion) — Independent benchmarking showing tabular foundation models can't beat tuned classical detectors undercuts the foundation-model-universality narrative. https://www.inovex.de/en/blog/anomaly-detection-do-tabular-foundation-models-help/
- A Later Test Set Is Not a New Domain: Pretraining Familiarity Survives Contamination-Free Hold-Out (research-paper) — The contamination-free hold-out study shows time-series foundation-model gains trace to pretraining overlap rather than genuine generalisation. https://papers.cool/arxiv/2609.10357
- Accuracy Is Not Service: A Decision-Aware Benchmark for Intermittent-Demand Forecasting (research-paper) — This decision-aware benchmark's negative correlation between forecast accuracy and service outcomes shows the field has been optimising the wrong metric entirely. https://arxiv.org/abs/2609.13840