The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail
AI models that forecast future values from historical time series data across demand, revenue, usage, and other metrics. Includes deep learning forecasting and automated model selection; distinct from financial forecasting which applies time series to a specific finance context.
AI-driven time series forecasting has reached the point where forward-leaning organisations extract real value from it -- but most have not yet started, and the field's central question remains unresolved. Neural and foundation model approaches (Transformers, TimeGPT, TimesFM) promise zero-shot generality across demand, revenue, and operational metrics, yet empirical evidence stubbornly shows that simpler methods -- gradient boosting, ARIMA, exponential smoothing, even optimized linear regression -- match or beat them on most production workloads. The M4 Competition, repeated benchmarking studies, July 2026 peer-reviewed comparisons across 30+ datasets and 50 financial assets, and practitioner case studies all converge on the same finding: model performance is task-dependent, not architecture-dependent. What makes this a leading-edge practice is not proof that deep learning wins, but that a mature vendor ecosystem, cloud-managed services, confirmed multi-sector deployments, and operationalized break-even analysis have made automated forecasting accessible at scale. The tension that defines this tier is method selection: organisations can deploy forecasting today with documented ROI (20-50% error reduction, 15-30% inventory improvements, $1.7B+ enterprise value creation confirmed), but choosing when neural/foundation-model complexity justifies its cost over classical alternatives still requires domain expertise and empirical validation rather than default architectural commitment.
The vendor ecosystem is consolidating around foundation models even as evidence mounts against their universal superiority. AWS completed its deprecation of Amazon Forecast, retreating from specialised forecasting-as-a-service -- a significant signal from the category's largest cloud provider. Foundation model vendors filled the gap: Google released TimesFM 2.5 (March 2026) with 200M parameters and 16k context length (8x expansion), integrated into BigQuery ML and Google Sheets for consumer-grade accessibility; Amazon Chronos-2 achieved 600M+ HuggingFace downloads and added multivariate/covariate support; Salesforce released Moirai-MoE with sparse mixture-of-experts outperforming larger rivals at 28x parameter efficiency. Datadog released Toto 2.0 (May 2026), an open-weights TSFM scaling from 4M to 2.5B parameters with continuous improvement and no saturation, signaling an ecosystem pivot toward scaling-driven architectures.
Real-world deployments confirm adoption breadth: retail (The Very Group: 9.9% SKU management improvement across 8M+ forecasts, o9 Solutions with AB InBev and Kraft Heinz achieving 60% stock-out reduction and 87% forecast accuracy at 99.5% service levels), manufacturing (Foxconn: 8% accuracy gain $553K annual savings; statworx case: 10% accuracy on 20K products), energy (renewable forecasting 14% balancing cost reduction, Belgium grid operators validating Chronos-2 and TimesFM 2.5 on volatile electricity pricing), and healthcare (ICLR 2026 confirms TSFM calibration superiority for risk-sensitive deployment; GlucoFM-Bench validates zero-shot transfer on diabetes prediction yet documents domain-specific challenges in T1D cohorts). Yet May 2026 production benchmarks reveal critical limitations: ARFBench on 63 real Datadog production incidents shows current TSFMs, LLMs, and VLMs achieve only 62.7% accuracy versus 87.2% oracle performance, documenting substantial gaps in multi-step reasoning capability over production data. Infrastructure-scale deployment evidence also surfaces barriers: data architecture gaps (nShift analysis of missing returns/cancellation data), misaligned optimization metrics (Expectations vs. Realities paper: MSE-optimal point forecasts systematically produce under-dispersed distributions failing in production), and empirical electricity market studies finding foundation models underperform on volatile real-world pricing signals.
June 2026 ecosystem maturity signals: production observability tooling emerging (ForecastOps open-source for TSFM monitoring), multi-model routing frameworks (TimeRouter achieving SOTA on GIFT-Eval without LLM overhead), domain-specific TSFM validation (APEX on 4,500 wireless networks demonstrating 18% MAE improvement over generic Toto), and next-generation benchmarking (TIME benchmark with 50 fresh datasets and zero-shot data-integrity validation). July 2026 research advances operationalized deployment decisions: break-even analysis across 30 datasets establishes data-volume thresholds (classical methods beat zero-shot on 6 datasets with <2,700 samples), while systematic volatility benchmarking on 50 financial assets confirms only small models (TTM) narrowly beat econometric methods—demonstrating that TSFM superiority is conditional, not universal. These advances indicate operational readiness—but also reveal fragmentation: specialist models (APEX for networks, domain-tuned GlucoFM instances) outperform generic TSFMs within their domains, yet zero-shot universality remains unproven; Amazon's SCOT (proprietary decade-refined supply chain optimizer) outperforms Chronos on domain data but is non-transferable, suggesting that domain-specific excellence and generic deployability are still misaligned.
Peer-reviewed research continues converging on fundamental findings that challenge vendor enthusiasm. Ridge regression with carefully tuned preprocessing matches or exceeds Transformer/MLP/CNN baselines on 6 of 8 standard benchmarks—demonstrating that model capacity does not automatically unlock forecasting accuracy and supporting pragmatic method selection over architectural commitment. Critical ICML 2026 research on traffic speed forecasting reveals that aggregate benchmark metrics mask regime-dependent calibration failures: Chronos and TimesFM collapse to 54.9% coverage in transition regimes versus 90% in stable conditions, with root causes being bimodal distribution misalignment that requires post-hoc Bayesian correction for production deployment. Healthcare domain evaluation shows mixture-of-experts TSFMs effective for epidemiological forecasting, yet LLM-based methods underperform relative to numerical forecasters—reinforcing that TSFM applicability is domain-specific and not universally superior. On high-frequency trading (5-minute Bitcoin), Kronos foundation models show no statistically significant advantage over classical Brownian motion baselines, documenting TSFM limitations without retraining.
Production deployments and adoption metrics confirm real-world uptake but with caveats: AWS Connect Decisions GA achieved 40% forecast accuracy improvement at Wells Vehicle Electronics with 90% automation; AWS-Kearney demand sensing platform delivered 10-20% accuracy improvement and 2% revenue lift at scale with multi-signal integration. Financial sector adoption metrics show 82% of CFOs plan increased AI/ML investment, with 20-50% error reduction versus traditional models and 71% reporting improved accuracy. Manufacturing practitioners caution that anomaly detection works reliably but demand sensing is overstated for B2B sparse signals; core barriers are data quality and explainability rather than model architecture. Method selection complexity persists as the primary adoption barrier—not which model architecture to choose, but whether forecasting teams optimize for business value and whether zero-shot generic foundation models offer genuine ROI over domain-specific fine-tuning. Portfolio approaches (Amazon Science: specialist models outperforming single monolithic TSFMs) and hybrid routing strategies with adaptive domain selection are emerging as pragmatic production patterns, reducing inference cost while maintaining accuracy. Emerging research on domain adaptation (Guard framework) demonstrates that selective distillation and contextual routing can address distributional misalignment in specialized domains without retraining, pointing toward pathways for broadening TSFM applicability beyond zero-shot generality.
July 2026 deployment evidence reinforces balanced realism on TSFM applicability: production benchmarks confirm forecasting accuracy gains (14 APAC retailers averaged 34% MAPE reduction; global food manufacturer achieved 20-30% improvement; grid infrastructure on 200 real feeders validated Chronos-2 superiority for peak prediction; Amazon SCOT forecasts 400M+ products daily with 20-50% enterprise error reduction) alongside empirical limitations. Extreme-event forecasting (California wildfire PM2.5) shows BiLSTM outperforming zero-shot TSFMs across all thresholds; volatility forecasting on 50 financial assets reveals only small models narrowly beat econometric baselines; optimized linear methods match Transformers on 6 of 8 benchmarks. Adoption reaches mainstream CFO sign-off (94% of organizations plan TSFM deployment within 2 years; Unilever delivered $1.7B value through integrated AI forecasting) yet hidden costs surface: 88% PoC-to-production gap, decision bullwhip risk in multi-agent scenarios, token cost explosion, and data architecture gaps remain binding constraints. Leading-edge classification holds because deployment breadth is confirmed across retail, manufacturing, finance, energy, and healthcare with quantified ROI; however, the central tensions persist: zero-shot universality remains unproven, method selection still requires domain expertise, and simpler approaches remain competitive on most production data—indicating that leadership in this domain flows from pragmatic method selection and business-value alignment rather than from architectural commitment.
— 70% of large organizations will adopt AI-based forecasting by 2030 (Gartner). But 55% unclear on ROI, 56% struggle with legacy integration, 50% lack expertise. Signals mainstream adoption with persistent implementation barriers.
— Production failure autopsy: $240k winter-coat overstock from automated purchasing. Replaced neural net with boosted tree. Critical findings on seasonality encoding, data leakage, loss function misalignment, and infrastructure cost optimization.
— Training-free alignment method for frozen TSFMs: 3.75% MSE improvement on Chronos-Bolt; 2.5-13.7% zero-shot gains on TimesFM, Moirai, Toto across seven benchmarks without per-backbone tuning.
— Fine-tuned Kronos on Turkish stock market: zero predictive advantage over naive baseline. Hourly fine-tuning worsened MAPE (1.79% → 2.01%). Critical negative signal on domain transfer and fine-tuning reliability.
— EU-AI-compliant local statistical models beat 100M+ parameter foundation models on critical infrastructure. Critical negative signal: simplicity outperforms foundation model complexity in safety-critical settings.
— Comprehensive practitioner evaluation: 13 models × 10 datasets × 4 horizons. Foundation models won 30/38 contests but no model transfers across contexts. Critical finding: leaderboard rank doesn't predict real-world performance.
— Federated parameter-efficient fine-tuning of Chronos-T5 with LoRA across distributed clients. Achieved 31% MAPE improvement with differential privacy on Indian agricultural markets.
— Enterprise integration pattern: Chronos-2 on SageMaker + Lambda + Snowflake external functions with CloudFormation infrastructure-as-code. Demonstrates TSFM adoption in data warehouse workflows.