🎧 Customer Operations
AI for supporting, retaining, and understanding customers after the sale. The highest concentration of good-practice tiers: chatbots, ticket routing, sentiment analysis, and voice-of-customer are deployed at scale in most industries. Bleeding-edge frontiers include autonomous resolution without human escalation and real-time emotion detection. Momentum is steady but churn prediction and proactive outreach remain stalled.
The Headline
AI that helps your service agents is proven and cheap. AI that replaces them is being rolled back by three companies in four. Fund the first, gate the second.
The Picture
Most companies now get AI assistance for service agents bundled into the platform they already pay for: drafting replies, summarizing calls, routing tickets, scoring quality. Two-thirds of service organizations run some form of AI agent (software that acts on its own without being prompted), up from 39% a year ago, so having it no longer differentiates anyone. A small group, including Vodafone, Lenovo and the Indian grocer Zepto, runs fully autonomous resolution at scale, but only on narrow, well-instrumented tasks with a person one click away. The rest sit in pilot purgatory: 74% of enterprises have pulled an autonomous customer-facing agent at least once, and a fifth of customers still say company chatbots do nothing for them. The window that is closing is not about buying tools; it is about fixing the knowledge base and escalation paths that every vendor's product depends on, and that only a third of support leaders say are ready.
This Fortnight
Salesforce closed its $3.6 billion purchase of Fin and put Agentforce at $1.5 billion in annual recurring revenue. Pricing has converged on around a dollar per resolved conversation across the major vendors. That consolidation gives you leverage in negotiation, but only if your contract defines a resolution the customer confirmed, not one the vendor assumed.
The strongest evidence yet that assisting agents pays: a field study of 5,179 support staff found AI-suggested replies lifted issues resolved per hour by 13.8%, with no drop in customer satisfaction. Junior agents gained 35%; seniors gained little. Point this spending at onboarding and your least experienced teams, not at everyone equally.
Independent production data undercut the vendor decks. A benchmark of 131 e-commerce merchants and 2.9 million tickets found 4.9% resolved fully by AI, while an audit of 33 vendors found advertised rates of 65 to 86% against 40 to 70% in the vendors' own case studies. Zendesk, unusually, published trial results showing its predictive routing helped in four of six deployments and did nothing in two; treat any vendor unwilling to show null results as a risk.
Customers have not moved even as spending doubled. Company-provided chatbots featured in 7% of customers' most recent service interactions, unchanged since 2022, while use of third-party AI tools nearly doubled. Your customers are adopting AI; they are not adopting yours, and 87% still require a route to a human.
Two areas lost momentum: claims automation and live call translation. False declines from automated dispute decisions cost merchants an estimated $201 billion in 2025, and retrofitting policy versioning into legacy insurance systems takes 300 to 900 hours; Zendesk announced the first native real-time voice translation for agent desktops, but analysts flagged unproven quality on sensitive calls and Microsoft's equivalent hit a hard EU data-residency block. Keep both in pilot unless your use case is low-stakes and outside the EU.
Coming Up
EU AI Act disclosure rules for customer-facing AI are already enforceable, with fines of €15 million or 3% of global turnover. A German court has separately held a company liable for what its chatbot told a customer, and the Air Canada precedent now appears in procurement documents. Have general counsel confirm this quarter that every autonomous customer interaction discloses itself and that a human can override any commitment the AI makes.
Zendesk shuts down its legacy scripted bots on December 10, 2026, and its replacement bills per verified resolution. The switch has exposed that old deflection metrics overstated performance by 15 to 25 points, partly because a conversation the customer abandoned for 72 hours was booked as resolved. Rebase your automation numbers on confirmed resolution before the migration, or you will discover the gap when the invoices arrive.
Real-time voice translation reaches general availability (out of beta) in mainstream contact-center platforms in the first quarter of 2027. Adoption intent among service leaders is heading toward 60%, but practitioners report hallucination (the AI confidently making things up) on 15 to 30% of realistic calls, and the largest outsourcers have kept translation itself at proof-of-concept. Budget for a scoped 2027 pilot on non-regulated calls; do not plan headcount around it.
What's Hard About This
Your knowledge base, not the AI model, sets the ceiling. The same agent on the same vendor platform has swung from 25% to 79% resolution on knowledge restructuring alone, and mature contact centers hit a 35 to 45% wall set by what information the AI can reach. A quarter of content hygiene will move your numbers more than any new model release, and only 32% of support leaders say their knowledge base is ready.
Keeping a human in the loop (a person reviews each AI output before it ships) is what makes deployments survive, and it is not free. Approval workflows lose money once reviewers miss more than 2% of errors on cases worth under $15, and a review of 106 experiments found people paired with an AI draft often did worse than the better human alone because the draft anchors them. Decide which decisions justify a reviewer's salary and which should be automated outright or left with a human; the middle ground is where costs hide.
Better governance finds failures; it does not prevent them. Organizations with fully mature safeguards roll back autonomous agents more often (81%) than average (74%) because their monitoring catches unauthorized actions and silent drift that others never see. Expect the governance bill to exceed the development bill, and treat a rollback as evidence the controls worked, not that the project failed.
Practices in this Domain (18)
Read the full technical briefing (1,966 words) →
Where AI Stands in Customer Operations
Customer operations is where enterprise AI has gone furthest, and where the evidence about what actually works is most brutal. Every tier-one platform — Zendesk, Salesforce, Microsoft, AWS, Genesys, Five9, Cisco — now ships intent detection, sentiment scoring, ticket routing, post-call summarisation, agent copilots and automated quality scoring as standard features rather than premium add-ons. Salesforce's State of Service puts AI agent adoption at 66% of service organisations, up from 39% a year earlier; Zendesk's Copilot triage has dropped to its Professional tier; call summarisation is bundled into base pricing almost everywhere. The domain's centre of gravity is a human-in-the-loop pattern that has now been validated causally rather than anecdotally: an NBER field study of 5,179 agents at a Fortune 500 software company found AI-suggested responses lifted issues resolved per hour by 13.8% (35% for juniors) with no satisfaction decline; TELUS Digital reports 15% from a 5,000-agent rollout; Octopus Energy's Magic Ink has summarised more than six million calls and now drafts roughly a third of outbound email with higher CSAT than unassisted mail. AI-assisted humans hold customer satisfaction at 84%, within a point or two of fully human service, while chatbot-only sits at 68–74%. That gap explains almost everything else about the domain.
Momentum is building on the augmentation side and stalling on the autonomy side, and the two are diverging rather than converging. The market has consolidated around autonomous resolution — Salesforce closed its $3.6 billion acquisition of Fin on 10 September and put Agentforce at $1.5 billion ARR; Intercom formalised $0.99-per-resolution pricing; Vodafone's SuperTOBi resolves 70% of 60 million monthly conversations; Lenovo runs autonomous agents across 500 million tickets a year. But independent production data refuses to match vendor decks. Chatarmin's benchmark of 131 e-commerce merchants and 2.9 million tickets found 4.9% resolved fully autonomously; EnderTuring's study of four mature contact centres found every deployment hitting a 35–45% ceiling set by what information the agent could reach, not by model quality; Drag's audit of 33 vendors found advertised resolution of 65–86% against 40–70% in the vendors' own case studies. Sinch's survey of 2,527 leaders — now cited so widely it functions as the domain's baseline — finds 74% of organisations have rolled back an autonomous customer-facing agent at least once. Klarna rehired humans after quality slid; Commonwealth Bank rehired 45 agents; a German court has held a company liable for what its chatbot told customers; EU AI Act Article 50 has been live since 2 August with fines of €15 million or 3% of turnover for undisclosed autonomous interaction. The only autonomous deployments surviving at scale are narrow, well-instrumented and gated: Zepto's 100,000 tickets a day with evaluation-first governance, CollageDepot's 65% auto-send at $0.79 a ticket, SteadyPay's FCA-regulated voice agents with twenty guardrails per turn.
What distinguishes this domain from its neighbours is that the binding constraint has shifted decisively from the model to the organisation around it, and the industry now has the numbers to prove it. Only 32% of support leaders rate their knowledge base AI-ready; 77% of technology leaders say a fifth or less of enterprise knowledge is agent-ready; 73% of autonomous-agent failures trace to stale or contradictory knowledge. Churn prediction has plateaued at the same wall: 76% of B2B SaaS firms have piloted AI health scoring but 22% have made it work, ML models need 500-plus historical churn events before they beat rule-based stacking, and 71% of customer-success leaders cannot explain why their own score predicts churn. Proactive engagement has near-universal investment intent and, as one pointed analysis this fortnight noted, still no independently audited case of AI causally lifting net revenue retention. Voice AI handles around 40% of tier-one calls at large US enterprises, yet SuccessKPI's survey of 400 contact centres finds 2.5% fully automated and 82% still in pilots. The pattern repeats across a dozen practices: the tooling is commoditised and the ROI is documented at the leading edge, while the median organisation is stuck between a pilot that worked and a production environment whose knowledge, data plumbing and escalation paths were never built for it.
What's New, 2026-09-06 to 2026-09-20
This window (6–20 September, a normal fortnight, though a few practices carry evidence from the last days of August) was dominated by consolidation at the top and sobering field data underneath. Salesforce completed the Fin deal and shipped seven pre-configured Agentforce agents with named customers (Anthropic 79%, Hibbett 90%); Zendesk pushed routing to voice, added Salesforce and HubSpot knowledge connectors, and announced the first native real-time voice translation in a mainstream CCaaS desktop (early access in October, GA in Q1 2027, 13 languages) — while Futurum immediately flagged unproven quality on sensitive calls and unresolved GDPR questions, and Microsoft's Dynamics 365 real-time voice hit a hard EU data-boundary block. Zendesk also did something rare: it published randomised A/B results for predictive routing showing gains in four of six deployments and no significant effect in two. Against that, the NBER study and TELUS's 5,000-agent rollout gave auto-draft its strongest causal evidence yet, Roland Berger reported overall customer-service AI adoption falling from 95% to 54% even as agent assist rose to the single most-used application, and Gartner found company-provided chatbots stuck at 7% of customers' most recent interactions, unchanged since 2022, while third-party GenAI use nearly doubled.
Two practices lost forward momentum this cycle and two settled. Returns, warranty and claims automation — until now advancing on strong unit economics — ran into quantified friction: false declines cost merchants an estimated $201 billion in 2025 with a third of affected shoppers churning, only 27% of merchants deploy AI fraud detection, and retrofitting effective-date versioning into legacy policy systems takes 300–900 hours, which makes coverage-sensitive claims automation infeasible for many carriers without architectural repair. Real-time voice translation likewise slipped: Everise scaled Krisp's accent conversion across 10,000 seats but kept translation itself at proof-of-concept, and practitioner reports put hallucination on realistic calls at 15–30%. Churn prediction and knowledge-base maintenance, by contrast, have stopped sliding and are now simply flat — mature tooling, well-understood failure modes (hand-built scorecards with untested weights; Gainsight's two uncoordinated scoring engines; knowledge decay and "AI slop" feedback loops), and no sign that either constraint is being solved faster than it is being documented. Elsewhere the news was corroborating rather than moving: Maccabi Healthcare's full IVR replacement at 45,000 calls a day, a Parloa rebuild cutting p95 audible latency from 1.9 seconds to 280 milliseconds, Aviva's £60 million-plus claims savings, and yet another restatement of the 74% rollback figure.
Key Tensions
The review gate is the product — and it has its own economics. Every deployment surviving at scale keeps a human between the model and the customer: Sinch's 74% rollback rate applies to agents that removed that step, and NIST SP 1353 now formalises the gate as a compliance control. But the gate is not free. Vendor analysis puts approval workflows underwater once reviewers miss more than 2% of errors on cases worth under $15; an MIT meta-analysis of 106 experiments found human-AI pairs underperforming the better human alone on decision tasks because a draft anchors rather than informs; Eesel's trial data shows 93% draft accuracy but only 12% of drafts sent unedited. The industry has chosen augmentation, and is only now costing what supervision actually requires.
Vendor metrics measure containment, production measures resolution. The gap between advertised and delivered performance is now the most-documented fact in the domain: Drag's 33-vendor audit (65–86% claimed, 40–70% in case studies), Zendesk's enterprise median of 41.2% against Decagon's 80%, Chatarmin's 4.9% full autonomous resolution across 2.9 million tickets. Mechanisms are specific — Zendesk's 72-hour inactivity threshold books a conversation as resolved even if the customer replies late, and a failed handoff with no agent available closes as an automated resolution. Gartner's finding that only 14% of service issues fully resolve via self-service, and Zendesk's own split of 'contained' from 'verified' resolutions, mark an industry being forced towards honest accounting.
Knowledge, not models, sets the ceiling. The same agent on the same vendor stack swings from 25% to 79% resolution on knowledge-base restructuring alone; Vodafone lifted first-time resolution from 15% to 60% without touching the model. Yet 32% of support leaders call their knowledge base AI-ready, 77% of technology leaders say a fifth or less of enterprise knowledge is agent-ready, and 94% of knowledge content goes untouched in a given month. EnderTuring's 35–45% ceiling across mature contact centres is an information-availability limit, which is why every frontier-model release changes less in this domain than a quarter of content hygiene would.
Governance surfaces failures rather than preventing them. Organisations with 'fully mature' safeguards roll back autonomous agents more often (81%) than the average (74%): better observability detects the authorisation gaps, cascading actions and silent drift that unmonitored deployments never see. Liability has hardened around that visibility — a German court ruling on chatbot statements, the Air Canada precedent now cited in procurement RFPs, EU AI Act Article 50 enforcement, Australia's Privacy Act requirements for post-deployment accuracy monitoring. The result is a market where 84% of AI teams spend more than half their time on safety infrastructure, and where governance spending now exceeds development spending.
The customer has not moved. Enterprise adoption has nearly doubled in a year, but Gartner finds company-provided chatbots at 7% of customers' latest interactions, flat since 2022; 64% of consumers prefer companies not use AI for service and 87% require a route to a human. The listening channel is collapsing in parallel — survey response rates have fallen to 5–15% and 30% of customers now stay silent after a bad experience, leaving no diagnostic trail. Verint's finding that 69% would switch to AI if it fully resolved their issue locates the problem in execution rather than principle, but that is cold comfort when production resolution sits in the forties.
Top 10 Evidence Items
- NBER Field Study: 13.8% Issues-per-Hour Gain from AI-Suggested Responses (5,179 Agents) (adoption-metric) — The causal NBER evidence anchors the whole human-in-the-loop case, distinguishing this domain's augmentation story from anecdote. https://stealthagents.com/research/ai-copilot-productivity-statistics-2026
- Sinch survey: 74% of autonomous AI agent deployments have rolled back (adoption-metric) — The 74% rollback figure is the single most-cited counterweight to autonomy hype and structures the whole tension around governance. https://sinch.com/blog/ai-agent-escalation/
- Drag vendor audit: advertised 65–86% autonomous resolution versus 40–70% in production (industry-report) — Quantifies the vendor-versus-production gap that the briefing calls the domain's most-documented fact. https://www.dragapp.com/blog/state-of-ai-support/
- MIT Meta-Analysis: Human-AI Review Pairs Underperform Humans Alone on Decision Tasks (opinion) — Supplies the uncomfortable academic counterpoint that even the celebrated human-in-the-loop pattern can underperform, complicating the augmentation narrative. https://www.cmswire.com/contact-center/why-is-your-contact-center-learning-about-bad-content-first/
- AI agents that can't escalate are costing you customers (Sinch rollback survey) (adoption-metric) — The finding that mature-governance organisations roll back more often is the sharpest illustration that observability surfaces failure rather than preventing it. https://sinch.com/blog/ai-agent-escalation/
- Zendesk Predictive Routing: built on evidence, not assumptions (case-study) — Zendesk's own disclosed null results are the rare case of a vendor publishing honest, non-cherry-picked A/B data. https://www.zendesk.com/blog/reports/predictive-routing-white-paper/
- AI knowledge base readiness: Where customer support AI fails (adoption-metric) — Grounds the knowledge-ceiling argument in a hard adoption number, showing the constraint has moved to the organisation rather than the model. https://www.customersuccesscollective.com/ai-knowledge-base-customer-support/
- Policy versioning as structural blocker for insurance claims automation: 300–900 hours to retrofit effective dating (opinion) — Shows a concrete architectural reason (legacy policy versioning) why a previously advancing practice has stalled. https://zyneto.com/blog/insurance-claims-automation
- German court establishes liability for autonomous chatbot customer communications (opinion) — The German liability ruling turns abstract governance risk into enforceable legal consequence for autonomous output. https://www.jdsupra.com/legalnews/who-s-liable-for-ai-hallucinations-in-9891424/
- Company-provided LLM chatbots at 7% customer adoption, flat since 2022 (Gartner, n=3,566) (adoption-metric) — The flat 7% adoption figure is the clearest evidence that customer behaviour hasn't moved despite enterprise-side momentum. https://chattermate.chat/ai-customer-service-adoption/