The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail
AI that automatically summarises support calls and generates disposition codes and structured notes for CRM entry. Includes after-call work automation and key moment extraction; distinct from call transcription in sales which focuses on sales conversations rather than support calls.
Call summarisation and disposition has reached mainstream platform maturity with ecosystem-wide GA adoption and documented ROI from early production deployments. Every major vendor (Microsoft, AWS, Zendesk, Genesys, Talkdesk, ServiceNow, Oracle, Webex, Dialpad, Five9) ships AI-generated post-call summaries and automated disposition as core GA features, and proven deployments report 25-50% AHT reductions with sustained agent productivity gains (40-90 seconds saved per call in field implementations). The practice replaces manual after-call work—typing summaries, selecting disposition codes, updating CRM—with LLM-based automation that extracts issue, resolution, action items, and classification codes on-call or immediately post-call. The technology clearly works at scale: Genesys, Vapi, and AWS implementations confirm sub-100-second automation cycles, and third-party platforms (Lindy.ai, Lindy) are shipping plug-and-play solutions. However, production deployments universally require deliberate architectural choices, quality validation protocols, and acceptance of persistent limitations. Hallucination, fabricated customer statements, and diarisation failures on hybrid calls remain unsolved technical barriers that separate capability availability from mid-market adoption. Critically, detection methods fail to reliably identify hallucinated summaries (even ensemble approaches collapse to near-random performance on summarisation tasks); hallucination detection cannot be the primary control, making human-in-loop review mandatory before CRM entry. Additionally, legal liability attaches to the deploying organisation, not the vendor—Air Canada precedent and emerging regulatory frameworks (Australia Privacy Act, ASIC Corporations Act) establish post-deployment accuracy monitoring as compliance obligation. Organisations willing to invest in RAG-grounded architectures, fine-tuning, and human-in-the-loop validation achieve genuine efficiency gains; those deploying out-of-the-box models continue to experience agent distrust, quality failures, and regulatory exposure.
Platform ecosystem standardisation is complete: all major contact centre vendors (AWS, Microsoft, Genesys, Zendesk, ServiceNow, Oracle, Webex, Talkdesk, Dialpad, Five9, CloudTalk) offer GA call summarisation bundled into base platform pricing. Geographic expansion continues (AWS Contact Lens June 2026 added Portuguese, French, Italian, German, Spanish, Chinese, Japanese, Korean support), confirming sustained investment in global production readiness. Third-party vendors (Vapi, Lindy.ai, Aircall, Nextiva) are shipping independent AI summarisation and disposition products, demonstrating ecosystem depth beyond platform incumbents.
Real-world deployments confirm ROI at scale: Genesys implementations document 45-90 seconds saved per call; Lindy.ai reports 40-60 second reductions with +40% contact centre capacity; Five9 TruConnect (healthcare) achieved 40% ACW reduction; Verint baseline research (1,000 agents) establishes 54% of calls require after-call work including summarisation, confirming scale of demand. Field implementations validate 25-50% AHT reduction from combined front-of-call and back-of-call AI automation, establishing credible mid-range ROI.
However, production deployments reveal persistent technical limitations constraining mid-market adoption. Hallucination remains endemic: independent comparative testing of five AI medical scribes documents systematic fabrication of medications/dosages, phantom exam findings, and confabulated patient statements—error types that directly map to call summarisation risks (invented customer statements, misrepresented agreements, fabricated action items). AI Evals production framework establishes threshold requirements (faithfulness >95%, coverage >85%) that most raw summaries fail; organisations deploying out-of-the-box models face 63-89% raw accuracy, rising to 94-96% only with structured human-in-the-loop validation. Critically, hallucination detection methods fail catastrophically on summarisation tasks specifically (detection ensembles collapse to near-random AUC 0.47-0.57, vs. 0.79 on QA and 0.71 on dialogue), meaning detection-only pipelines cannot flag hallucinated summaries reliably—human-in-loop review before CRM entry remains the only effective control. Speaker diarisation accuracy drops ~30 percentage points on hybrid calls; domain jargon blindness requires custom vocabulary tuning; context reconstruction on escalations costs $200-500 per incident. Genesys implementations explicitly document mandatory human review of AI outputs before finalisation (agents cannot skip validation step), and UJET research confirms 93% of agents feel need to double-check AI outputs pre-deployment despite crediting summarisation with ACW reductions. Regulatory exposure has emerged as a new constraint: Air Canada chatbot precedent establishes deploying organisations (not vendors) as liable for AI output accuracy; Australia's Privacy Act (2026-12-10) and ASIC Corporations Act require post-deployment accuracy monitoring and transparency as compliance obligation. The practice tier has stabilised at good-practice: mainstream feature availability coexists with explicit technical barriers (hallucination, quality validation costs, tuning complexity, regulatory compliance burden, detection failure) that separate capability from confident autonomous deployment.
— ISI Analytics GA release of Call Summary template for Standard call data, confirming automated call summarization as baseline feature in contact center analytics platforms.
— Zoom official documentation confirms GA post-call AI summary feature with real-time question answering and administrative controls—evidence of mainstream vendor adoption of call summarization.
— Microsoft Dynamics 365 Contact Center admin documentation for enabling Copilot case summaries with token thresholds and exclusion controls—production-grade GA feature in major CCaaS platform.
— Openlayer (Gartner-named evaluator) documents hallucination taxonomy and rates (3–27% of queries, 35–40% fabrication in legal research); explains why detection fails—critical for understanding call summarization quality risks in production systems.
— Veteran BPO analyst positions auto-summarization as 'cheapest, cleanest win' among contact center AI, documenting 10–20 second AI wrap-up versus 45–90 seconds manual, yielding 18 FTE capacity recovery per 10k daily contacts.
— ERAA-2026 benchmark tests 15 RAG systems on 10K multi-hop queries; reveals 51% of responses omit material contradictions, 38% citations unsupported—directly applicable to call summarization synthesis quality risks.
— Regulatory and liability framework citing Air Canada chatbot precedent; establishes deploying organisation (not vendor) is legally liable for AI output accuracy; Privacy Act (Australia) + ASIC Corporations Act require post-deployment accuracy monitoring—new governance barrier for disposition automation.
— Peer-reviewed TrustNLP workshop paper demonstrating span-level unlikelihood training reduces hallucinated summaries from 31% to 13% on CNN dataset (58% reduction) and 33% to 20% on SAMSum (39% reduction)—validating fine-tuning approaches to core hallucination barrier.