The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail
AI generation of natural-sounding speech from text for audiobooks, accessibility, navigation, and content delivery. Includes multi-language synthesis and emotional expression; distinct from voice cloning which replicates specific voices rather than generating generic natural speech.
The vendor landscape has stratified by performance envelope and use-case. Sub-300ms P95 latency tier (real-time conversational): Cartesia 40–90ms (SSM architecture, Sonic-Turbo), OpenAI Realtime-2 (May 7 launch with speech-to-speech), LiveKit+Gemini (295ms P95), Gradium Sonic (155ms P50), Inworld TTS-2 (<130ms P90 on Mini tier). Batch/content tier: ElevenLabs (41% Fortune 500, 2M agents, $500M+ ARR post-Series D, $11B valuation), Amazon Polly (31 generative voices, AWS SageMaker JumpStart integration June 2026), Azure (400+ voices, 140+ languages, MAI-Voice-2 zero-shot cloning June 2). Efficiency tier: Inworld TTS-1 Max ($10/M chars, ELO 1,162), Minimax Speech 2.8 HD (86.2% approval, ELO 1,107), PlayHT (85.6%), WellSaid Labs (82%). Open-source maturity: Higgs Audio v3 (Boson AI, 4B-parameter, 100+ languages at 3.61% WER, inline emotion/style control), Chatterbox-Turbo (65.3% preference over ElevenLabs Turbo v2.5 in blind tests), Kokoro ($0.70/M chars), Fish Speech, CosyVoice 2.0. Independent evaluation platforms (Coval, LMSYS, Artificial Analysis, Voice AI Leaderboard) standardize benchmarking; Coval emphasizes vendor benchmarks unreliable on P95 latency and domain-specific pronunciation. Vendor-published P50 metrics systematically outperform real production P95/P99 under load.
Production deployment spans content (5.2x monthly growth in independent audiobook production via Inkfluence, ACX policy enabling AI narration June 2026), accessibility (7.5M K-12 students under US IDEA; $4B TTS market 2024 → $7.6B 2029, compliance-driven adoption but implementation quality bottlenecked), healthcare (30M clinician minutes via voice agents), customer service (Klarna 10x resolution, Mahindra 8% conversion, Revolut 31 languages, UK motor insurer 4× handle time reduction, Berlin-Brandenburg Airport zero-wait), and public sector (UK government MoU with ElevenLabs for national-scale AI voice requiring 300+ languages). Real production telemetry from 10+ live deployments: median 680ms P50 latency (5–8× slower than human turn-taking at 239ms); all-in costs $0.07–$0.21/min with containment 62–88%; latency > 1,500ms triggers 40%+ call abandonment. Vendor-neutral empirical evaluation (Podonos, 1,060 human evaluators): ElevenLabs ranked highest naturalness; AWS Polly shows 28.6% unnatural intonation and 8.8% sudden emotion changes. Open-source production maturity: Hugging Face + Cerebras speech-to-speech stack (Alibaba Qwen3-TTS + Gemini 4 LLM) deployed on 9,000+ Reachy Mini robots, establishing vendor-agnostic infrastructure maturity. However, production-reality gap emerges: vendor demos achieve MOS 4.5–4.8 on benchmark audio; real Twilio deployments (8 kHz mu-law codec compression, 5,000-character prompts) show robotic prosody and quality degradation measurable via automated MOS predictors (NISQA/UTMOS). Expressive-domain limitations documented: Interspeech 2026 peer-reviewed research shows naturalness and appropriateness vary independently—systems excel at reading (newscast) but fail on expressive domains (acting, animation, spontaneous speech). Multilingual adoption barriers: phonology-informed evaluation shows high-naturalness systems fail to preserve language-specific phonological structure; Meta's MMS TTS realized [-ATR] vowels as [+ATR] in 1/3 of tokens despite high quality scores, impairing intelligibility in tone-sensitive and vowel-harmony languages. Emotional expressiveness and zero-shot cloning now standard across major vendors (Microsoft MAI-Voice-2, ElevenLabs v3, Google Gemini 3.1 Flash); architecture innovation (flow-matching, diffusion, SSM) continues sub-100ms synthesis. Cost-quality-latency cluster converging: infrastructure-maturity shift reflected in voice cloning becoming standard API parameter across platforms (Fish Audio S2 $15/M chars, 70–100ms TTFA; Qwen3-TTS Apache 2.0 self-hostable; ElevenLabs Flash $0.05–$0.10/1K chars). Adoption barriers persist: consumer preference (Audio Publishers: 16% tried AI audiobooks, AI revenue 0.03% of market, willingness dropped 70%→61% YoY—Audible 63% audiobook market shows TTS quality/authenticity barriers dominate despite availability), expressive-domain limitations (prosody, emotion, code-switching, multispeaker coherence remain unsolved), multilingual phonological gaps, regulatory risk (BIPA litigation, $100M–$2B+ precedent damages), operational instability (recurring platform incidents), regional constraints (emerging markets 800ms–1.4s latency penalties, language quality gaps, compliance misalignment), vendor lock-in (94% enterprise concern, 2.3–5.7x switching costs). Evaluation methodology standardized via independent platforms (Coval, LMSYS, Artificial Analysis); quality commoditization at MOS 4.2–4.3 shifted competitive differentiation to latency consistency (P50 vs P95 variance), pronunciation accuracy (1–3% WER variance), domain-specific appropriateness, and production reliability under load.
— Market consolidation evidence: LOVO bankruptcy (May 2026), Play.ht shutdown (Dec 2025); vendor rankings show Speechify Simba 3.2 #1 on Artificial Analysis leaderboard, documenting competitive displacement and market maturity.
— Official incident tracking within evaluation window shows 7 incidents in 30 days with 2h 1m avg recovery; documents operational challenges and single-point-of-failure risks despite broad Fortune 500 deployment.
— NVIDIA containers TTS as production microservice with July 2026 benchmarks: 46-52ms first-chunk latency on H100, sub-100ms inter-chunk on A100/L40; signals infrastructure-layer commoditization across GPU platforms.
— Peer-reviewed TensorRT optimization achieves 5.0x speedup on autoregressive GPT component and 3.6x end-to-end with minimal quality loss; streaming support enables production deployment at scale.
— Independent streaming-latency benchmark: Palabra v1 achieves 103ms TTFA (sub-100ms viable); reveals production hierarchy and distinguishes batch throughput from streaming metrics critical for voice agents.
— Comprehensive landscape analysis signals H1 2026 inflection point: on-device quality parity with cloud APIs, quality gap narrowed 223 to 81 Elo in 3 years, 54-voice Kokoro ranks top-5 globally at 82M parameters.
— Open-source speech-to-speech stack (Alibaba Qwen3-TTS + Deepmind Gemma 4 LLM) deployed on 9,000+ Reachy Mini robots in active production, demonstrating vendor-agnostic infrastructure maturity for voice AI.
— Interspeech 2026 peer-reviewed research reveals multilingual TTS adoption barrier: Meta's MMS TTS fails to preserve phonological structure (realized [-ATR] vowels as [+ATR] in 1/3 of tokens) despite high naturalness scores.