The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail
AI that generates realistic but artificial datasets for testing, training, and privacy-preserving data sharing. Includes tabular, text, and image synthetic data; distinct from data augmentation which modifies real data rather than generating from scratch.
Synthetic data generation promises to break the deadlock between data access and data privacy, but after seven years of development the practice remains experimental, with production use confined to a handful of high-governance verticals. The core idea -- generating artificial datasets that preserve the statistical properties of real data across tabular, text, and image domains -- has attracted substantial vendor investment and regulatory attention. NVIDIA's acquisition of Gretel and Microsoft's integration of synthetic data into Phi-4 training signal genuine commercial confidence. Yet independent research from EPFL and Max Planck has formalised hard limits: for many use cases the trade-off between fidelity and privacy cannot be overcome algorithmically. Vendor consolidation reinforces the caution; multiple funded startups have shut down or been acqui-hired, and surviving companies are pivoting toward platform embedding rather than standalone tools. Recent June 2026 ICML research has refined understanding of model-collapse mechanisms, proving that quality-assurance verifiers optimized for local datasets (healthcare consortia, financial institutions) paradoxically accelerate collapse when distribution coverage is incomplete—turning safeguards into systemic risks. Where synthetic data works -- fraud detection in banking, clinical trial augmentation in pharma, QA in regulated software, high-fidelity simulation training in aviation -- it works within tightly bounded conditions with careful real-synthetic mixing. Broader enterprise scaling remains blocked not by a lack of tooling but by unresolved privacy-validation standards, relational-data quality gaps, and refined understanding of model-collapse risks in siloed operational environments.
— Case study of Ideally deployment across 23 Omnicom agency brands with 60+ tests at 1/10 traditional budget cost, validated limitations (regression to mean, need for human oversight), reflecting mature understanding of synthetic data in production workflows.
— Peer-reviewed study quantifying fundamental gap between utility-passing synthetic clinical benchmarks and structural realism (79% missingness, 12% actionable rows), proving utility checks insufficient for production readiness.
— Peer-reviewed research documenting silent bias amplification before standard collapse metrics trigger alarms, a critical production risk for recursive synthetic training in language models requiring separate fairness monitoring.
— Market adoption data: 35% of Fortune 500 firms in regulated sectors have deployed synthetic data in production; Gartner projects 75% adoption by end-2026 (up from <5% in 2023), indicating mainstream acceleration in enterprise.
— Practitioner analysis identifying verification and anchoring as prerequisites for synthetic data success, with concrete examples (AlphaGeometry proofs, Phi-4 manufactured data, robotics simulators) showing how quality-assurance and real-data retention prevent collapse.
— Peer-reviewed deployment demonstrating synthetic agricultural imagery solving real annotation burden while maintaining accuracy through domain-gap-aware optimization, with practical impact on real-world phenotyping generalization.
— AWS Transform signals synthetic data generation as standard service in enterprise database migration workflows, demonstrating major cloud vendor embedding practice at scale for schema validation without privacy/compliance risk.
— Production defect detection case study: synthetic data achieved 80.9% MAP on real rotogravure manufacturing defects from zero real-data starting point, solving long-tail scarcity in constrained industrial environments.