The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail
AI that monitors agent interactions for quality and compliance while providing real-time sentiment and tone coaching. Includes automated QA scoring and in-call coaching prompts; distinct from agent assist which drafts responses rather than evaluating agent performance.
AI-driven quality monitoring and coaching is a proven capability with mature GA tooling, measurable deployment ROI at scale, and ecosystem-wide adoption—yet implementation and operationalization remain the binding constraints, not technology. The technology works: auto-scoring accuracy reaches 99%+, real-time coaching reduces attrition by 20-35% in operationalized deployments, and organisations scaling 100% coverage report 25-40% faster agent ramp-up and 20-30% quality gains. Yet new August 2026 evidence reveals two critical failure modes. First: 74% of enterprises roll back deployed AI agents due to governance failures; even organisations with mature governance roll back at 81%, suggesting governance surfaces rather than prevents failures. Second: 15% increase in AI adoption (2023-2025) accompanied by 0.5-point loss in CX/EX scores, showing that adoption without proper scorecard design and coaching integration actively harms outcomes. Only 29% of organisations with deployed QA tools use them effectively, and deployment-to-value gap remains unresolved. Coaching delivery also lags: 81% of agents report conversations never reviewed despite full-coverage tooling, and only 15% use automated feedback methods, indicating that technology readiness outpaces operational capability. Fairness in scoring remains limited: peer-reviewed ACL 2026 research documents systematic bias (5.4-13.0% judgment reversals, worse with contextual priming), and 75-80% of deployed systems carry documented accent, sentiment, and gender bias. The practice remains in good-practice tier because deployment evidence is credible and ROI is real where operationalized; advancement requires solving the operationalization, governance, and fairness gaps that characterise current deployments.
Calabrio, Observe.AI, NICE, Verint, and Gryphon all ship GA products offering 100% interaction coverage, automated scoring, and real-time agent coaching. August 2026 ecosystem maturity reached new milestone: Cisco's Webex (reaching millions of enterprise CCaaS users) launched AI Quality Management as GA, confirming that automated QA and real-time coaching are now table-stakes across major contact center platforms. Named deployments continue to validate ROI where operationalized: Crash Champions deployed an AI agent handling 400-500 calls/day with mature monitoring workflows including A/B/C performance ratings, escalation-based intervention, and continuous feedback loops. Cresta's Snap Finance deployment showed 40% AHT reduction, containment improvement from 6% to 33%, and 23% CSAT increase. Balto's multi-customer deployments (InteLogix, Credit Control, Bordner) demonstrated 24% enrollment lift, compliance score improvement from 70s to 90-100%, and appointment-rate gains. Verint and Observe.AI continue documenting scale ROI: €79M benefit with 30-second AHT reduction at a telco, $12.5M savings from 3%→97% QA coverage scaling. All evidence points to genuine, measurable business value where operationalization is mature.
Yet August 2026 evidence also surfaces three critical barriers now blocking tier advancement. First, governance failures cascade: Sinch's 2,527-respondent AI Production Paradox study found 74% of enterprises rolled back customer-facing AI agents; crucially, organisations with mature governance roll back at 81%, indicating governance surfaces failures rather than preventing them. Second, adoption paradox: Deloitte Canada's analysis of AI adoption (2023-2025) revealed 15% adoption increase accompanied by 0.5-point loss in CX/EX scores, showing that adding coverage without addressing scorecard alignment or coaching integration actively harms outcomes. Third, coaching delivery gap persists: Solidroad survey (500 agents) found 81% report conversations never reviewed despite full-coverage tooling, and only 15% use automated feedback methods, indicating that technology exceeds operational capability to deploy. Fairness in LLM-based QA systems remains limited despite high accuracy: ACL 2026 peer-reviewed research documents systematic bias with 5.4-13.0% judgment reversals, worse with contextual priming (up to 16.4%), calling for standardized fairness auditing before deployment. The practice maturity paradox: major vendors (Webex, Cresta, Balto, Verint, Calabrio/Verint) deploy 100% coverage with proven ROI; yet 74% rollback, adoption harming outcomes, and coaching delivery gaps reveal that technology maturity exceeds implementation maturity. Operationalization remains the binding constraint.
— Cisco Webex AI Quality Management reached GA with 100% automated interaction evaluation, real-time coaching insights, and Agent Performance Dashboard, confirming platform maturity across major CCaaS ecosystem.
— State of CX survey (500 agents): 81% say most conversations never reviewed; 79% find feedback helpful but only 15% use automated feedback method; highlights coaching delivery gap and coverage expectations shifting to 100% as table stakes.
— Three named customer deployments from independent review: InteLogix (24% enrollment lift, ACW halved), Credit Control (compliance 70s→90-100%), Bordner (improved appointment rates); demonstrates enterprise-scale QA + coaching deployment outcomes.
— Peer-reviewed ACL research documents systematic fairness gaps in LLM-based QA systems: CFR (judgment reversals) 5.4-13.0%, worse with contextual priming (16.4%); calls for standardized fairness auditing before deployment.
— 12-month banking case study shows AI deflection creates four-tier call distribution; traditional metrics (AHT, FCR) break post-automation; argues quality monitoring must shift to per-conversation contribution, not operational metrics.
— Sinch AI Production Paradox: 74% of enterprises rolled back customer-facing AI agents; governance failures primary cause; 81% rollback even among organizations with mature governance, indicating governance surfaces rather than prevents failures.
— Identifies insight-to-action gap: 18-30% AHT reductions when QA connects to training/coaching, but most platforms stop at scoring; lack of integrated training layer prevents value realization; links measurement to business outcomes.
— Crash Champions deployed AI agent handling 400-500 calls/day with mature monitoring workflow: A/B/C performance ratings, escalation-based intervention, outcome-focused metrics, continuous feedback loops.