Perly Consulting │ Beck Eco

The State of Play

A living index of AI adoption across industries — where established practice meets the bleeding edge
UPDATED DAILY

The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.

The Daily Dispatch

A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.

AI Maturity by Domain

Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail

DOMAIN
BLEEDING EDGEESTABLISHED

Agent quality monitoring & coaching

GOOD PRACTICE

TRAJECTORY

Stalled

AI that monitors agent interactions for quality and compliance while providing real-time sentiment and tone coaching. Includes automated QA scoring and in-call coaching prompts; distinct from agent assist which drafts responses rather than evaluating agent performance.

OVERVIEW

AI-driven quality monitoring and coaching is a proven capability with mature GA tooling, measurable deployment ROI at scale, and ecosystem-wide adoption—yet implementation and operationalization remain the binding constraints, not technology. The technology works: auto-scoring accuracy reaches 99%+, real-time coaching reduces attrition by 20-35% in operationalized deployments, and organisations scaling 100% coverage report 25-40% faster agent ramp-up and 20-30% quality gains. Yet new August 2026 evidence reveals two critical failure modes. First: 74% of enterprises roll back deployed AI agents due to governance failures; even organisations with mature governance roll back at 81%, suggesting governance surfaces rather than prevents failures. Second: 15% increase in AI adoption (2023-2025) accompanied by 0.5-point loss in CX/EX scores, showing that adoption without proper scorecard design and coaching integration actively harms outcomes. Only 29% of organisations with deployed QA tools use them effectively, and deployment-to-value gap remains unresolved. Coaching delivery also lags: 81% of agents report conversations never reviewed despite full-coverage tooling, and only 15% use automated feedback methods, indicating that technology readiness outpaces operational capability. Fairness in scoring remains limited: peer-reviewed ACL 2026 research documents systematic bias (5.4-13.0% judgment reversals, worse with contextual priming), and 75-80% of deployed systems carry documented accent, sentiment, and gender bias. The practice remains in good-practice tier because deployment evidence is credible and ROI is real where operationalized; advancement requires solving the operationalization, governance, and fairness gaps that characterise current deployments.

CURRENT LANDSCAPE

Calabrio, Observe.AI, NICE, Verint, and Gryphon all ship GA products offering 100% interaction coverage, automated scoring, and real-time agent coaching. August 2026 ecosystem maturity reached new milestone: Cisco's Webex (reaching millions of enterprise CCaaS users) launched AI Quality Management as GA, confirming that automated QA and real-time coaching are now table-stakes across major contact center platforms. Named deployments continue to validate ROI where operationalized: Crash Champions deployed an AI agent handling 400-500 calls/day with mature monitoring workflows including A/B/C performance ratings, escalation-based intervention, and continuous feedback loops. Cresta's Snap Finance deployment showed 40% AHT reduction, containment improvement from 6% to 33%, and 23% CSAT increase. Balto's multi-customer deployments (InteLogix, Credit Control, Bordner) demonstrated 24% enrollment lift, compliance score improvement from 70s to 90-100%, and appointment-rate gains. Verint and Observe.AI continue documenting scale ROI: €79M benefit with 30-second AHT reduction at a telco, $12.5M savings from 3%→97% QA coverage scaling. All evidence points to genuine, measurable business value where operationalization is mature.

Yet August 2026 evidence also surfaces three critical barriers now blocking tier advancement. First, governance failures cascade: Sinch's 2,527-respondent AI Production Paradox study found 74% of enterprises rolled back customer-facing AI agents; crucially, organisations with mature governance roll back at 81%, indicating governance surfaces failures rather than preventing them. Second, adoption paradox: Deloitte Canada's analysis of AI adoption (2023-2025) revealed 15% adoption increase accompanied by 0.5-point loss in CX/EX scores, showing that adding coverage without addressing scorecard alignment or coaching integration actively harms outcomes. Third, coaching delivery gap persists: Solidroad survey (500 agents) found 81% report conversations never reviewed despite full-coverage tooling, and only 15% use automated feedback methods, indicating that technology exceeds operational capability to deploy. Fairness in LLM-based QA systems remains limited despite high accuracy: ACL 2026 peer-reviewed research documents systematic bias with 5.4-13.0% judgment reversals, worse with contextual priming (up to 16.4%), calling for standardized fairness auditing before deployment. The practice maturity paradox: major vendors (Webex, Cresta, Balto, Verint, Calabrio/Verint) deploy 100% coverage with proven ROI; yet 74% rollback, adoption harming outcomes, and coaching delivery gaps reveal that technology maturity exceeds implementation maturity. Operationalization remains the binding constraint.

TIER HISTORY

ResearchJan-2020 → Jan-2020
Bleeding EdgeJan-2020 → Jan-2022
Leading EdgeJan-2022 → Jan-2025
Good PracticeJan-2025 → present

EVIDENCE (155)

— Cisco Webex AI Quality Management reached GA with 100% automated interaction evaluation, real-time coaching insights, and Agent Performance Dashboard, confirming platform maturity across major CCaaS ecosystem.

— State of CX survey (500 agents): 81% say most conversations never reviewed; 79% find feedback helpful but only 15% use automated feedback method; highlights coaching delivery gap and coverage expectations shifting to 100% as table stakes.

— Three named customer deployments from independent review: InteLogix (24% enrollment lift, ACW halved), Credit Control (compliance 70s→90-100%), Bordner (improved appointment rates); demonstrates enterprise-scale QA + coaching deployment outcomes.

— Peer-reviewed ACL research documents systematic fairness gaps in LLM-based QA systems: CFR (judgment reversals) 5.4-13.0%, worse with contextual priming (16.4%); calls for standardized fairness auditing before deployment.

— 12-month banking case study shows AI deflection creates four-tier call distribution; traditional metrics (AHT, FCR) break post-automation; argues quality monitoring must shift to per-conversation contribution, not operational metrics.

— Sinch AI Production Paradox: 74% of enterprises rolled back customer-facing AI agents; governance failures primary cause; 81% rollback even among organizations with mature governance, indicating governance surfaces rather than prevents failures.

— Identifies insight-to-action gap: 18-30% AHT reductions when QA connects to training/coaching, but most platforms stop at scoring; lack of integrated training layer prevents value realization; links measurement to business outcomes.

— Crash Champions deployed AI agent handling 400-500 calls/day with mature monitoring workflow: A/B/C performance ratings, escalation-based intervention, outcome-focused metrics, continuous feedback loops.

HISTORY

  • 2020: Observe.AI and Calabrio establish AI-powered agent quality monitoring as a distinct capability; Observe.AI secures $80M in funding and lands partnerships with HCL and 3CLogic, enabling 100% call coverage for sentiment and compliance scoring.
  • 2021: Observe.AI reaches 160 customers with 20,000+ agent licenses; launches AI-powered coaching product suite (4X coaching session increase); integrates with Microsoft Azure. Calabrio expands to 100% omnichannel interaction analytics. Cloud adoption accelerates (68% of contact centers), validating infrastructure readiness. Implementation challenges emerge: nonlinear effort to achieve quality thresholds and risk of agent demotivation from automated scoring.
  • 2022-H1: Observe.AI reports 150% ARR growth and 40% enterprise customer increase; launches Auto QA for adaptive automation (up to 1,000x coaching insights increase). Calabrio earns G2 Leader recognition and integrates with Talkdesk. Independent survey shows 67% of contact centers still manual but CI adopters 10x more confident; ecosystem maturity advances via platform integrations and product innovation.
  • 2022-H2: Vendor ecosystem accelerates real-time coaching capabilities. Calabrio and Verint present advanced coaching and quality management features at industry conferences. Market adoption survey shows 78% of contact centers plan AI deployment within 3 years, with quality management as a top priority use case. Implementation focus shifts from manual sampling to 100% automated interaction evaluation.
  • 2023-H1: Observe.AI launches Real-Time AI product suite adding live guidance and supervisor coaching; Calabrio maintains leadership with mature Auto QM evaluation forms. Third-party adoption metrics show significant gaps: 62% of contact center managers cannot analyze enough calls for accurate performance evaluation (Invoca). Vendor innovation focuses on real-time coaching ROI with measurable behavioral impacts (5-6% win rate lifts). Two-tier market emerges between cloud-native leaders and traditional centers.
  • 2023-H2: Critical research from SQM Group reveals persistent adoption barriers: only 19% of managers believe QA programs improve CSAT, and 83% of agents don't believe QA helps their performance. Despite mature product capabilities and vendor innovation in ROI tooling, fundamental user skepticism remains a deployment barrier. Quality monitoring reaches mainstream commercial stage with 100% interaction evaluation becoming standard, but adoption unevenness persists between cloud-native and traditional contact centers.
  • 2024-Q1: Calabrio acquires Wysdom.AI to expand bot QA analytics; deployment evidence shows Awaken achieving 56% reduction in difficult calls and 10% sales uplift. ISG analyst research projects two-thirds of contact centers will increase training/coaching budgets by 2026. Level AI survey finds 100% of leaders considering AI adoption with 23% higher satisfaction among agents using real-time AI tools. Community skepticism persists about AI agent reliability in live customer interactions, highlighting deployment risks alongside vendor momentum.
  • 2024-Q2: NICE releases Real-Time Interaction Guidance for AI-driven agent coaching with contextual compliance prompts, confirming category-wide focus on real-time coaching as table-stakes capability. Calabrio continues product innovation with Bot Analytics tools. Vendor ecosystem shows steady product maturation in real-time guidance and evaluation, though no major new deployment case studies emerge in this quarter.
  • 2024-Q3: Market adoption metrics show 39% of CX leaders using AI-driven scoring for employee and customer evaluation (CallMiner survey, 700 leaders). AWS ecosystem integration advances with real-time sentiment analysis templates for contact center deployment. Research from Salesforce reveals fundamental vulnerabilities in AI evaluation systems: LLMs vulnerable to deceptive feedback with 50%+ performance degradation. Analysis of 47 failed enterprise AI deployments ($127M sunk) identifies testing gaps and insufficient human oversight as key adoption barriers. Industry data shows 74% of contact centers still rely on random sampling, with AI achieving 100% coverage—adoption bifurcation persists between cloud-native and traditional centers.
  • 2024-Q4: NICE releases AI for Agents with 100% conversation evaluation and real-time coaching, confirming vendor commitment to quality monitoring as table-stakes. Adoption momentum continues: 33% of contact centers actively using emotion recognition for sentiment analysis; Frost & Sullivan projects two-thirds of CX operations plan AI-driven coaching within 3-5 years. However, critical legal barriers emerge: class-action lawsuits under privacy statutes (CIPA) challenge automated quality management when customer consent absent. Market bifurcation persists—cloud-native leaders deploying 100% AI-automated evaluation while traditional centers remain largely manual; execution gaps widen between vendor innovation and customer deployment capability.
  • 2025-Q1: Named deployments confirm measurable ROI: AAA Northeast reduced AHT by 14 seconds via AI analytics (equivalent to 1 FTE); Australian energy provider cut QA scoring inconsistencies by 35%, improving FCR/CSAT; GE Appliances/Delta Dental show cost reduction and attrition/defect improvement. AWS Marketplace integration of Observe.AI signals cloud ecosystem maturity. McKinsey/Gartner data shows 30% CSAT improvement and 25% productivity gains from AI monitoring. Practitioner analysis emphasizes hybrid human+AI model necessity—pure automation risks employee pushback, judgment gaps, and legal compliance issues. Market remains bifurcated between cloud-native leaders and traditional centers.
  • 2025-Q2: Calabrio's survey reveals near-universal AI adoption (98%) but persistent implementation challenges: 61% of centers report more difficult conversations since AI deployment, 32% cite agent distrust as critical barrier. Calabrio releases 70+ new features (Auto QM, Trending Topics, Interaction Summary), confirming ecosystem maturity. Observe.AI documents 350+ enterprise deployments with 60% efficiency gains and 75% QA time reduction. Critical shift in evidence landscape: practitioner analysis exposes measurement gaps (e.g., telecom provider showed 12% AHT improvement but 7% revenue decline), revealing tension between operational metrics and business outcomes. Legal/compliance barriers persist; market bifurcation between cloud-native leaders and traditional centers widens.
  • 2025-Q3: Vendor ecosystem innovation continues: Omind launches AI QMS platform with 100% automation, 30% cost reduction, 95% compliance accuracy, 20% CSAT gains, and up to 59-second AHT improvement. Observe.AI case studies document RealDefense (103% quota lift, 13% revenue boost) and Nations Info Corp (50% save rate improvement, 43% AHT reduction). However, critical deployment risks crystallize: 75-80% of enterprises deploying AI QA grading; documented bias manifestations in scoring (accent, sentiment, gender, script-adherence bias) affecting agent reviews and coaching. Organizational failure rates spike: Gartner forecasts 85% of AI projects fail; S&P Global 2025 data shows 42% of companies abandoned most AI initiatives. Fundamental gaps persist—legacy system integration, change management, causation modeling between metrics and business outcomes, and compliance under CIPA privacy constraints. Market bifurcation widens: cloud-native leaders achieve strong ROI, traditional centers struggle with execution. Technology maturity exceeds implementation maturity.
  • 2025-Q4: Technical standardization emerges: Deepgram and UseScore publish production guidelines for speech-to-text sentiment analysis (5-10% WER) and scaling from 3-5% manual sampling to 100% coverage (70-90% workload reduction achieved 8-12 weeks post-deployment). Industry benefit data solidifies: 20-40% CSAT, 15-25% repeat-call reduction, 30-50% escalation gains consistently reported. Observe.AI validated as IDC MarketScape Leader in Workforce Engagement Management. However, deployment fundamentals remain unchanged: 61% of centers report conversation quality degradation post-deployment; 32% cite agent distrust; organizational failure rates sustained at 42% abandonment; bias risks (accent, sentiment, gender, script-adherence) persist across 75-80% deployed systems. No tier-advancement signals emerge; market bifurcation between cloud-native leaders (with strong ROI) and traditional centers (struggling with execution) persists unchanged.
  • 2026-Jan: Calabrio launches Omni Agent Intelligence for unified human+AI agent monitoring, confirming market evolution toward hybrid deployment frameworks. New product capabilities extend to quality measurement across autonomous and human agents. However, Gartner forecasts 40% of agentic AI projects will be canceled by 2027 due to cost, unclear ROI, inadequate risk controls, and integration friction (70% of developers report integration problems). Technical maturity advances (95-98% call-scoring accuracy achieved) but organizational adoption barriers persist: unoptimized QA processes, inadequate change management, and measurement gaps between operational metrics and business outcomes remain tier-limiting factors.
  • 2026-Feb: Calabrio QM deployment demonstrates advanced technical capability: 99%+ auto-scoring accuracy, 90% manual QM time reduction, 25% agent attrition improvement, 41% ACW reduction; healthcare deployment (CareAI) manages 53% of inquiries via automated quality evaluation. However, strategic optimization analysis reveals critical disconnect: 98% AI adoption across contact centers but only 12% claim fully optimized value; 86% remain in "pilot purgatory." Leadership and cultural barriers intensify—only 35% of agents understand AI tool usage, >50% fear automation, and 64% of leaders neglect empathy training. Deployment maturity remains unchanged with persistent challenges: bias risks in 75-80% of systems, 42% organizational abandonment rates, measurement gaps between operational and revenue metrics. Technology capability advances but implementation/organizational maturity static, preventing tier advancement.
  • 2026-Apr: New deployment evidence confirms scale ROI where adoption is mature: Verint Coaching Bot delivers €67.8M benefit and 20-second AHT reduction at a telco, $70M savings at an insurer, and +39 NPS at a mortgage lender; platform comparison data shows BPO customers scaling QA coverage from 3% to 100% with 5-point CSAT gains. The operationalization gap sharpens as the defining constraint: CMSwire data shows 88% of contact centers deployed AI but only 25% operationalized it into daily workflows, and only 52% allow shared visibility between agents and AI systems. A real failure case (insurance company's knowledge base error affecting 660 calls over 11 days before a customer complaint surfaced it) illustrates why 100% monitoring coverage has practical value beyond efficiency — it catches systematic errors that sampling misses.
  • 2026-May: Major vendor commitments formalize next-generation capabilities. Microsoft launches Quality Assurance Agent within Dynamics 365 Contact Center (GA April 2026), emphasizing shift from sampling to real-time evaluation across both AI and human agents. Palomarr analyst framework ranks 94 vendors on transcription accuracy, real-time analytics, and coaching automation, identifying LevelAI (9.8), Cresta (9.7), and Observe.AI (9.6) as leaders. Independent survey of 815 enterprise executives (Liveops/Peter Ryan Strategic Advisory) identifies continued maturity gap: 65% remain in Walk/Run stages (hybrid human-AI workflows) requiring quality management infrastructure, while only 14% reach Fly stage with continuous real-time optimization. McKinsey data confirms 90%+ AI accuracy vs 70% manual; $286K annual savings per 1% FCR improvement. Verint Quality Bot delivers enterprise-scale outcomes: €8.6M digital channel savings at 80% containment and 30% attrition reduction, with Fiserv achieving 1%→96% QA coverage and $12.5M annual savings by eliminating 1,200 manual QA roles. Salesforce State of Service study (3,075 respondents) confirms 66% adoption rate with 70% observing measurable value within 60 days—strongest rapid-ROI signal in the category this cycle. Metrigy's CX Assurance research study finds companies with advanced assurance practices are 2.2× more likely to succeed, and explicitly extends the monitoring mandate from human agents to autonomous AI systems. Real-time coaching architecture consolidates as a three-layer pattern (pre-interaction compliance cues, during-call guidance, post-interaction micro-coaching), with independent technical assessment of Observe.AI confirming platform production-readiness. Vendor comparison ranking 7 automated coaching platforms documents 25-40% faster new-rep ramp time as a repeatable outcome. Tele Access BPO hybrid QA model combines 100% AI coverage with adult-learning-grounded coaching methodologies (Good Call/Bad Call peer sessions, morning briefings), demonstrating mature operationalization. 88% centers deployed AI but only 25% operationalized into daily workflows; operationalization gap and real-time coaching as attrition driver when QA feels punitive remain the binding constraints.
  • 2026-Jun: Outcome-based measurement, scale evidence, and M&A consolidation define the June picture. COPC's third-party case study of a B2B e-commerce deployment reveals that only 7 of 60 percentage points of unresolved issues were agent-controllable—53 pp were policy, process, or tool-driven—fundamentally reshaping how QA findings should drive operational change rather than individual coaching. ETS Labs' QEval documents 2.5 billion interactions processed in 2025 with sub-4-minute latency and zero backlog, confirming 100% coverage is now operationally achievable at industry scale. Level AI's Vista deployment illustrates the step-change from 1-2% manual sampling to 100% AI scoring with improved coaching effectiveness. Research (Journal of Applied Psychology via TechBullion) reinforces that real-time feedback produces 3x larger behavioral effect than delayed review, strengthening the case for in-call coaching over post-call evaluation. Regulatory pressure sharpens: an RBI fine of Rs 1.31 cr against a BFSI firm exposed that 3% sampling carries a 97% miss-probability per violation with a 4-8 week detection lag, making 100% coverage a compliance necessity rather than a performance aspiration. Verint's $2B acquisition of Calabrio (confirmed June 2026) merges two of the three largest QA/WFM vendors, signaling category consolidation and enterprise strategic validation. Verint Engage 2026 conference produces three named quantified deployments: FNB South Africa (14x automated evaluations, 15% compliance improvement), NOS Portugal (40% productivity gain, 61-point NPS lift), and a major unnamed bank ($10M agent capacity savings). Market adoption metrics firm up: 80% of contact centers now use AI-based QA technologies (up from 30% in 2021), with primary research on 109 directors/VPs finding 85% deployed AI QA tools but only 29% use them effectively—the deployment-to-value gap, not technology maturity, remains the binding constraint.
  • 2026-Jul: Named deployments continue to quantify coaching ROI at scale: a Verint telco case shows €79M benefit with 30-second AHT reduction, an equipment distributor scaled Quality Bot coverage from 3% to 97% for $12.5M in savings, and Insignia's real-time coaching cut first-90-day attrition by 34%. Execution gaps persist alongside the shift to 100% AI-powered evaluation (from 1-5% manual sampling): 65% of leaders call their AI deployment successful, but 43% of projects are delayed or stalled and 28% report lost revenue from poor complexity handling—reinforcing that operationalization, not detection accuracy, remains the binding constraint.
  • 2026-Aug: Cisco Webex AI Quality Management reaches GA with 100% automated evaluation and real-time coaching, extending the category-wide shift to full coverage across major CCaaS platforms; Crash Champions' mature deployment (A/B/C ratings, escalation-based intervention, continuous feedback) and Cresta's Snap Finance case (40% AHT reduction, 23% CSAT gain) show what operationalized coaching looks like in production. Peer-reviewed ACL research documents systematic fairness gaps in LLM-based QA scoring (5.4-16.4% judgment reversals under contextual priming), and a Solidroad survey finds 81% of conversations still go unreviewed despite full-coverage capability existing—reinforcing that the insight-to-action and governance gaps, not scoring technology, remain the binding constraints.