The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail
AI that provides real-time spoken translation during customer support calls, enabling cross-language support. Includes live interpreter replacement and bidirectional voice translation; distinct from content localisation which translates pre-written materials rather than live speech.
Real-time voice translation in support calls promises to decouple agent hiring from language requirements, letting monolingual contact centers serve global customers. The premise is compelling: translate speech live, eliminate interpreter costs, and staff for skill rather than fluency. Vendor consolidation accelerated through June 2026 with simultaneous GA launches from Google (Gemini 3.5 Live Translate, June 9), Krisp (Voice Translation API), and competitive offerings from Gradium, DeepL, OpenAI, and Microsoft solidifying real-time translation as standard enterprise platform feature. Yet the practice remains firmly experimental. Independent benchmarks consistently show AI translation accuracy between 60-85% against a 95%+ human baseline, with hallucination rates of 33-60% and cultural mistranslation rates around 40%—gaps that confine deployments to lower-stakes, cost-sensitive scenarios. Production deployment barriers have hardened: code-switching (70% of Indian contact centers naturally code-switch to Hinglish) fails in traditional ASR pipelines; acoustic artifacts from poor microphones and overlapping speech cause 63% more failures than noise alone; and entity accuracy gaps (16.7-25.5% miss rates) expose risks in regulated industries. Cloud infrastructure compounds the problem: latency varies by architecture—cascaded pipelines (STT→translate→TTS) spike to 2-6 seconds, while new end-to-end models drop to 2.9-3.6 seconds, narrowing but not eliminating the gap. The defining tension is not whether the technology works in demos—it does—but whether it can sustain the accuracy, compliance, and responsiveness that live customer conversations demand. Production deployments at scale (Alorica hospitality and healthcare, Krisp healthcare 90% resolution without interpreters, OpenAI early customers including Zillow 69%→95% call success, Instadesk manufacturing 58%→79% FCR lift, Google Grab pilots with 10M calls/month) demonstrate real-world viability and ROI, but adoption remains minority-level (17% of enterprises) with 83% still relying on manual or traditional workflows. The pilot-to-production conversion gap is severe: 88% of voice AI pilots fail to reach production due to accent/dialect bias, latency spikes, and compliance barriers. For most contact center operators, this is a space to watch and pilot, not to bet operations on—though successful pilots signal material ROI for operators that clear the engineering and governance barriers.
Vendor ecosystem expanded through July-August 2026 with continued platform consolidation and new entrants. xAI entered the market (July 29) with Grok Voice Think Fast 2.0, achieving 82.9% on speech-to-speech benchmarks (vs GPT-Realtime-2.1 at 79.1%, Gemini 3.1 Flash at 69.5%), 0.70s first-audio latency, and verified Starlink phone service deployment improvements. Krisp released Voice Translation v3 (July 13) with automated quality scoring per translated call, six-domain accuracy benchmarking, and supervisory monitoring in 60+ languages. Zoom achieved GA for live voice translation (July 24) with voice cloning and intonation preservation, reaching 55%+ videoconferencing market share (300M daily active users). These maturity signals combine with June 2026 launches: Google's Gemini 3.5 Live Translate (June 9) with audio-to-audio architecture and 70+ languages; Krisp's Voice Translation API with 1M+ production minutes and 96% healthcare accuracy; Gradium's two-pass architecture achieving 3.0s latency; OpenAI gpt-realtime-translate at 4.53/5 fidelity. Established vendors—Microsoft (Dynamics 365 May GA), AWS, DeepL (Feb GA), and TTEC—solidified real-time translation as standard platform feature. Independent benchmarking confirms measurable quality gaps: Slator March 2026 showed DeepL Voice at 79% fully-correct segments (vs 42% for competitors), while latency comparisons document cascaded systems at 800ms-2s vs end-to-end models at 2.9-3.6s. Evaluation methodology challenges emerged: IWSLT 2026 peer-reviewed research revealed that audio-infused metrics fail to reliably surpass text-only baselines due to noise pollution and transcript mismatches, signaling fundamental measurement barriers. The vendor ecosystem now includes 21+ competing platforms with diverse cost structures ($0.04–$1.25/minute, Fora Soft 2026 production data), though vendor claims on language coverage frequently exceed production-quality support—marketing counts mask real-world mid-conversation switching failures and quality degradation outside English.
Production-scale real-world deployments demonstrate viability in specific use cases with material ROI. Krisp's healthcare provider achieves 90% end-to-end multilingual call handling without interpreter intervention, zero patient safety incidents across 8 languages (June 2026). Automated Health Systems achieved documented operational improvement: multilingual call handle time reduced from 40+ minutes to 9 minutes via Voice Translation (CCW 2026 case study, July 2026). OpenAI's early customers show measurable improvements: Vienna boutique hotel (direct-booking conversion +38%), DACH e-commerce (time-to-market compression 9 months→3 weeks), Zillow (26-point call-success-rate improvement: 69%→95%). Google is piloting Gemini Live Translate with Grab (10M voice calls monthly) in high-noise Southeast Asia (Thai, Vietnamese, Bahasa Indonesia, Tagalog). Alorica ReVoLT serves Fortune 25 healthcare organizations with 75+ languages (Everest Group validation). eesel.ai documented three production deployments: German jewelry (1K tickets/month), Spanish insurance (564 conversations/48 hours), German lending (100K+ tickets/month), fully automated at scale. Adoption signals accelerating: survey of 815 CX leaders (August 2026) shows 60% plan to adopt AI Voice Translation within 6-12 months, up from 36% in 2025; 28% currently deployed. Yet ecosystem adoption remains constrained by engineering barriers: 17% enterprise adoption as of March 2026, with customer support the 23% primary driver—indicating market remains early-stage despite vendor product maturity. Mordor Intelligence sizes the market at $2.74B (2026) growing to $5.58B (2030) at 19.5% CAGR.
Critical deployment barriers remain hardened despite vendor architectural progress. Code-switching is quantified as systemic blocker: ServiceNow's June 2026 SWER benchmark documents hallucination and phonetic errors across frontier ASR models on Hinglish and Spanglish, with 30-50% WER degradation on mixed-language utterances (70% of Indian contact center users naturally code-switch). India-specific ASR shows 94-96% accuracy on Indian English but drops to 88-92% on Hinglish, a 15-20 point delta that determines deployment viability in South Asian markets. Acoustic failures in real-world support calls create cascading accuracy loss: background voices cause 63% more failures than noise alone; elderly care deployments show 78% failure from TV audio interference. Entity accuracy gaps remain substantial (16.7-25.5% miss rates per AssemblyAI and Deepgram), particularly problematic in healthcare and financial support where terminology precision is non-negotiable. Enterprise compliance remains adoption barrier: 84% of organizations failed pre-deployment AI compliance audits, citing SOC 2, GDPR, BAA certification gaps. Platform-specific failures documented: Azure Speech Service systematically filters code-switched English from Cantonese transcription; Retell AI voice agents fail mid-call language switching with documented revenue impact. Infrastructure barriers persist: AudioCodes reports most enterprise voice AI projects stall not from AI failure but from telephony integration complexity (SIP/VoIP inconsistencies), scale bottlenecks at hundreds of concurrent sessions, vendor lock-in risk, and transcription accuracy as the binding constraint rather than translation quality. Latency trade-offs remain: cascaded pipelines (STT→translate→TTS) achieve 2-6 second round-trip latency; end-to-end architectures (Gemini 3.5, Gradium S2S) improve to 2.9-3.6 seconds but still exceed conversational naturalness thresholds (<800ms). These constraints mean that while the technology works in bounded use cases (monolingual to single-target-language high-volume support in non-critical contexts), real-time translation remains confined to cost-sensitive, non-regulated support routing until code-switching, compliance, latency, and acoustic robustness barriers narrow substantially.
— Survey of 815 CX leaders shows 60% plan to adopt AI Voice Translation in 6-12 months (up from 36% in 2025), 28% currently deployed; adoption intent nearly doubled year-over-year despite tech integration barriers remaining low on priority list.
— Fora Soft's production playbook details four-stage multilingual AI pipeline (ASR→MT→TTS→transport) with sub-1-second latency now achievable for high-resource pairs, cost $0.06–$0.18/min, and enterprise deployment patterns from 250+ projects.
— Peer-reviewed IWSLT 2026 research reveals evaluation challenges: audio-infused metric models fail to reliably surpass text-only baselines due to noise pollution and audio-transcript mismatches, signaling fundamental measurement barriers in speech translation.
— xAI launches Grok Voice Think Fast 2.0 with 82.9% speech-to-speech benchmark (vs GPT-Realtime-2.1 79.1%, Gemini 69.5%), 0.70s first-audio latency, $0.08/min pricing, and verified Starlink phone service deployment improvements.
— Zoom GA live voice translation (55%+ videoconferencing market share) with voice cloning and intonation preservation in 5 languages (Arabic +3 more Q4 2026), 300M daily active users, privacy-first architecture.
— Production case from CCW 2026: Automated Health Systems reduced multilingual call handle time from 40+ minutes to 9 minutes via Voice Translation, demonstrating material operational ROI in healthcare support context.
— AWS tutorial for Amazon Connect real-time translation in support: 75+ languages, bidirectional customer-agent message translation with PII redaction, claimed 60–80% staffing cost savings, 40–50% agent productivity gains.
— Krisp Voice Translation v3 GA: automatic quality scoring on translated calls, six-domain accuracy benchmarking, 60+ languages, live multi-language supervisory monitoring, session-start speed 50% reduction—signals productization maturity.