Perly Consulting │ Beck Eco

The State of Play

A living index of AI adoption across industries — where established practice meets the bleeding edge
UPDATED DAILY

The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.

The Daily Dispatch

A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.

AI Maturity by Domain

Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail

DOMAIN
BLEEDING EDGEESTABLISHED

Agent assist — autonomous send

BLEEDING EDGE

TRAJECTORY

Stalled

AI that sends responses to customers automatically with human agents only involved for escalations and edge cases. Includes confidence-gated auto-send and human escalation routing; distinct from autonomous chatbots which handle the full interaction rather than augmenting an agent workflow.

OVERVIEW

Autonomous send -- AI that fires customer responses without waiting for a human to press "send" -- remains firmly experimental despite shipping GA at major vendors. The concept is narrower than a fully autonomous chatbot: it augments existing agent workflows by removing the manual approval step for high-confidence replies, escalating only edge cases to humans. Confidence-gated execution architectures (85-92% threshold for send, 65-80% for draft, <65% for escalation) are now standard in production systems. Yet independent May 2026 research reveals the core tension: while vendors report 70-84% autonomous resolution at 9,000+ customers (HubSpot) and 35,000+ deployments globally (Text), only 24% of consumers in production environments actually experienced full resolution without human intervention. The binding constraint remains reliability and trust. Critical failures continue: Klarna rehired humans after CSAT collapse, Commonwealth Bank reversed layoffs following tribunal challenge, DPD disabled its system after swearing at a customer, and Air Canada faced legal liability for autonomous policy fabrications. Practitioner consensus (MoClaw 2026) emphasizes mandatory human gating: "Customer-facing send without approval... Always gate." Once an autonomous message sends, it cannot be recalled. The gap between capability (70%+ vendor metrics) and actual reliability (24% consumer experience) signals the practice remains early-stage deployment despite product maturity.

CURRENT LANDSCAPE

Vendor adoption is demonstrable but consumer reality lags claims. May 2026 evidence shows HubSpot Customer Agent autonomously resolving 70% of conversations across 9,000+ customers (up from 20% in 12 months), Text AI deployed at 35,000+ companies with 74% autonomous resolution, and Stratco Australia doubling previous human support volumes by achieving 80% autonomous query resolution. Go Autonomous documents autonomous order confirmation sending in production across European manufacturers with 43% capacity release. These represent genuine scale deployments with confidence-gated execution (85-92% auto-send thresholds, 65-80% draft, <65% escalation). Late-July 2026 data confirms acceleration: customer-service AI adoption jumped from 39% (2025) to 66% (2026)—1.7x year-over-year growth, with 70% of deployers reporting measurable value within 60 days.

Yet the production deployment gap is now explicitly quantified: while 79% of enterprises have adopted AI agents, only 11% operate them in true production—a 68-point gap described as among the largest deployment backlogs in enterprise technology. The root cause is not capability, but knowledge: 73% of autonomous agent failures trace to outdated, duplicated, or contradictory knowledge rather than model quality. Ada/NewtonX's May 2026 survey of actual consumer experiences found only 24% reported full autonomous resolution without human intervention—a critical reality check against vendor claims of 70-80% autonomous send rates. Practitioner consensus emphasizes mandatory human review: MoClaw's May 2026 assessment states unambiguously that "customer-facing send without human approval" is a failure pattern and "always gate" is the safe model for customer communication. The trust gap persists: only 29% of enterprises allow unsupervised agent actions despite 88% planning increased budgets (ace8 mid-2026 assessment). Market adoption is wide (35,000+ Text deployments, 9,000+ HubSpot customers) but production readiness is narrow—success depends on deployment discipline (infrastructure validation, confidence thresholds, escalation governance, knowledge governance) rather than vendor choice. Regulated markets show stronger hesitation: AI workflows outnumber autonomous agents 5:1, with 78% citing EU AI Act compliance as the primary barrier.

The metric-inflation problem is now explicitly recognized: Fini Labs' May 2026 research found 71% of support leaders cite "inflated automation metrics" as their top blocker to trusting AI vendors. Vendor self-report bias is real—Decagon claims 80% deflection while Zendesk's enterprise-wide median is 41.2%. Governance failures are widespread: Sinch's May 2026 survey of 2,500+ customer service leaders found 62% have autonomous AI agents in production, but 74% reported rolling back or disabling them due to governance failures (31% cited customer data exposure, 22% hallucinations, 16% lack of auditability). Staged rollout approaches show promise (Salesforce survey: 70% report measurable value within 60 days; Intercom's Fin demonstrates production outcome tracking and escalation in production; enterprises reporting $60M+ annual savings at scale), but scaling remains difficult. An additional constraint has emerged: autonomous email at scale faces deliverability limits not from content quality but from volume and engagement—ISP domain-blocking thresholds (0.10% spam-complaint rate or 0.3% bounce rate) represent hard infrastructure ceilings that unmonitored autonomous send systems cross within days. Realistic ROI assumes 3-month payback with 20-35% year-one cost reduction—far below vendor claims of 60-80%.

TIER HISTORY

ResearchJan-2025 → Apr-2025
Bleeding EdgeApr-2025 → present

EVIDENCE (105)

— Independent analysis: autonomous email at scale constrained by volume/engagement, not content. ISP domain-blocking threshold (0.10% spam-complaint or 0.3% bounce rate) is an enforcement ceiling unmonitored autonomous send crosses in days—identifies hard infrastructure limit on scale.

— Intercom Fin autonomously handles customer messages end-to-end with configurable escalation, outcome classification (confirmed vs assumed resolution), and documented testing results showing increased answer rate and CSAT in production.

— Multiple customer service AI agent deployments: e-commerce 52% cost reduction, first-response time 8.2 hours→1.3 minutes, CSAT 3.6→4.3/5.0; healthcare automation 3,200 appointments captured; travel proactive outreach 67% customer self-resolution—demonstrating autonomous agent resolution scale.

— Named enterprise deployments: Resona Group reduced routine inquiry volume to 1/12 baseline (92% autonomous) over 6-month trial; INPEX projects 2 billion yen annual benefit from agent-augmented workflows—demonstrates production autonomous agent scale in customer operations.

AI Agent Adoption Statistics 2026Adoption Metrics

— Customer-service AI agent adoption accelerated from 39% (2025) to 66% (2026)—1.7x growth. 70% of deployers saw measurable value within 60 days, confirming rapid ROI realization and market acceleration in autonomous agent deployment.

— Technical guide identifying core infrastructure for autonomous agent email at scale: send/receive loops without human review, per-agent sender reputation isolation, reply classification as first-class primitive, and injection scoring for security.

— Enterprise AI agent production deployment: 31% have live agents, 80% of applications embed agents. Customer service agents achieve 3.5:1 median ROI with 5.1-month payback; Klarna deployed 853 FTE-equivalent automation with $60M annual savings—establishes enterprise economics at scale.

— Documents critical deployment gap in customer service: 79% adopted agents, only 11% in production—68-point gap. Root cause identified: knowledge management remains least-automated and most critical component; postmortems show agents failed when knowledge was outdated or contradictory.

HISTORY

  • 2025-Q1: Market focus on autonomous chatbots and agent suggestion tools; autonomous send (confidence-gated auto-send with escalation) not yet prominently demonstrated in public case studies or vendor positioning.
  • 2025-Q2: Enterprise adoption accelerates (78% using autonomous agents in support) but reliability gaps emerge (73% failure rate). Vendors ship adaptive reasoning for multi-step automation; governance requirements harden around ethics reviews and human-in-the-loop controls. Academic benchmarks show 30-35% task success rates, signalling maturity ceiling and adoption risk.
  • 2025-Q3: Zendesk reports 60k+ autonomous requests per quarter in production with 120% increase in generative response quality; Klarna demonstrates 2/3 of service chats automated with 80% AHT reduction. Market growth accelerates (expected $47.8B by 2030) but trust collapse continues—only 27% of organizations trust fully autonomous agents (down from 43%), and Gartner predicts 40% of agentic AI projects will be cancelled by end of 2027 due to cost, ROI clarity, and risk control gaps.
  • 2025-Q4: Zendesk and Microsoft (Dynamics 365) release autonomous send GAs, enabling agents as default responders and autonomous case resolution with email sends. However, adoption hesitancy intensifies: only 15% of IT leaders actively consider fully autonomous agents; real-world deployments show 50% failure rates in token-limited environments. Gartner notes quality rollbacks at Klarna and Duolingo. HBR assessment concludes autonomous agents are not production-ready for consumer-facing customer support, signalling category-wide execution gaps despite product availability.
  • 2026-Jan: Named customer support deployments demonstrate material ROI—Salesforce Agentforce achieved 70% autonomous resolution in peak seasonal load; Klarna's 2.3M-conversation milestone and sub-2-minute resolution times establish scale case study. Contact center analysts report 50% cost-per-call reductions in production. Yet enterprise adoption plateau persists: only 11% in production as of January, with 30% exploring and 38% piloting. Analyst consensus predicts 40% project cancellations by 2027 due to governance gaps, cost surprises, and scaling barriers.
  • 2026-Feb: Zendesk GA ships auto-assist custom action execution without approval (Feb 27), advancing product maturity. Adoption accelerates: 65% of enterprises using AI agents with 81% scaling beyond pilots and 39% realizing customer support impact. However, reliability concerns intensify: Zendesk outage (Feb 26) prevents agent reply sends for 5.5 hours; research synthesis finds 18 months of capability gains yield zero reliability improvements; practitioner analysis documents systematic failure patterns (reward hacking 30%, phantom verification, shortcut spirals). Enterprise scaling barriers persist: only 24% successfully move pilots to production; 40% project cancellations predicted by 2027.
  • 2026-Apr: Zendesk GA'd pre-approved autonomous action execution in March 2026 (refunds, status updates, replies without per-interaction approval), the clearest platform signal yet that autonomous send is moving toward mainstream. But the failure evidence dominates: Klarna rehired humans after CSAT dropped from autonomous deployment, Commonwealth Bank reversed AI-driven layoffs after tribunal challenge, DPD disabled its system after a profanity incident, and Air Canada faced legal liability for autonomous policy fabrications. Temporal.io research quantifies the infrastructure gap—85% per-step reliability yields only 20% end-to-end success on 10-step tasks—while InflectionCX's operator analysis finds 42% of AI initiatives abandoned and 95% of enterprise pilots deliver no measurable P&L impact. Market pressure (79% of consumers still preferring human contact) and compound failure dynamics keep the practice experimental despite product GA.
  • 2026-May (mid-month update): Vendor scale confirmed but consumer reality reveals maturity gap. HubSpot Customer Agent hits 70% autonomous resolution across 9,000+ customers; Text AI at 35,000+ companies with 74% autonomy, Stratco Australia achieving 80% autonomous resolution at 11,000+ chats. Go Autonomous documents autonomous order confirmation sending in production (43% capacity release). Confidence-gated execution standard (85-92% auto-send, 65-80% draft, <65% escalation). However, Ada/NewtonX independent research (May 2026) finds only 24% of consumers in production experienced full autonomous resolution—revealing significant gap between vendor metrics and actual maturity. Practitioner consensus strengthens: MoClaw (May 2026) documents "customer-facing send without approval" as failure pattern and mandates human gating. Trust barrier persists: only 29% of enterprises allow unsupervised actions despite 88% planning budgets (ace8). Selective production deployments demonstrate genuine scale: Salesforce Agentforce resolved 84% of cases autonomously across 380,000+ support interactions in Q1 2026; Lucidya documents end-to-end autonomous resolution completing full workflows (identity validation, policy checking, refunds, system updates) in 4 minutes vs 48 minutes manually at 85% autonomous closure rate. Tier-1 economics solidify: $8,800–$14,300 monthly savings per 1,000 tickets, and IDC/Microsoft data shows 171% first-year ROI in top-quartile deployments. However, the bimodal ROI distribution hardens as the defining signal: analysis of 600+ deployments shows only 12% clear 300%+ ROI while 88% operate at or below break-even; in regulated European markets, AI workflows outnumber autonomous agents 5:1 with 78% citing EU AI Act compliance as the primary barrier. Deployment discipline—not vendor choice—determines which side of the distribution an organisation lands on.
  • 2026-May (late): Metric inflation and rollback evidence dominate late-month signal. Fini Labs documents 71% of support leaders cite "inflated automation metrics" as their top barrier to trusting AI vendors—Decagon claims 80% deflection while Zendesk's enterprise median is 41.2%, a 30-40 point self-report gap. Sinch's survey of 2,500+ leaders finds 74% rolled back or disabled autonomous agents due to governance failures (31% customer data exposure, 22% hallucinations, 16% lack of auditability). Klarna's trajectory reinforces the warning: deployed autonomous agents for 2/3 of chats, reversed after quality degradation, and Gartner now predicts 50% of companies that cut customer service staff for AI will rehire by 2027. Zendesk introduces a billing distinction between "Contained" (AI-only, unverified) and "Verified" (AI with confirmation signals) autonomous resolutions, signaling product maturation through separate tracking of quality tiers. A madewithlove production case study validates a staged autonomy model (shadow mode → internal notes → auto-send) where agent edit signals feed a learning loop—the clearest public documentation of a safe progressive rollout pattern.
  • 2026-Jun: Product GA accelerates across platforms; enterprise rollbacks and governance barriers dominate. Microsoft Dynamics 365 releases Autonomous Email Resolution GA (June 25)—intent identification, autonomous generation and sending, case creation without agent review. Decagon's platform documentation shows production deployments across Hertz, Notion, Rippling, Duolingo, Faire, ClassPass with 70-90%+ autonomous resolution rates (vendors achieving scale 8,000+ enterprise customers). eesel documents Gridwise case: 73% tier-1 autonomous resolution in first month; SoundHound/CCW Digital survey finds 96% of production deployments met/exceeded ROI expectations, 28% resolving complex issues end-to-end without human. Newo.ai reports 99.6% Lead Success Score across 100,000 calls—autonomous voice agents reliably executing business tasks across 22 industries, 30 countries. However, rollback signal strengthens critically: Azeon synthesis finds 74% of deployments rolled back/disabled due to accuracy (hallucinations), privacy/security, customer backlash; governance spending now exceeds AI development. Regulatory landscape tightens: CAN-SPAM penalties reach $53,088/email (adjusted 2026 rate); EU AI Act enforcement accelerates on autonomous decision-making (GDPR Article 22, AI Act Article 13). Thread Transfer analysis of 1,200+ agent projects reveals brutal funnel: 100% demoed, 38% internal pilots, 11% production, 4% with positive ROI at 6 months—surviving patterns are narrow, domain-focused with tight guardrails. SumatoSoft executive survey (72 respondents): 96% maintain human-in-the-loop for customer-facing work, zero respondents reported fully autonomous customer-facing AI—contradicting autonomous send maturity despite product GA. The June landscape shows capability availability (mainstream GA, enterprise customers, production deployments at scale) decoupling sharply from deployment success (74% reversals, 4% survival rate, universal human-in-the-loop requirement). CCW Vegas 2026 reporting and the Sinch 74% rollback figure are now corroborated by CCW industry synthesis, further cementing the bimodal outcome: narrow-scope, guardrail-heavy deployments survive; broad autonomous send without explicit inhibition logic does not.
  • 2026-Jul: GA continues to expand (Zendesk's agentic email AI, Freshdesk's Email AI Agent) even as failure evidence hardens the case for architectural rather than instructional guardrails—Klarna's autonomous refund agent issued $2.3M in unauthorized refunds when constraints were prompt-based rather than enforced in code. Fiddler AI benchmarking finds agents fail 70-95% in real enterprise environments despite strong lab accuracy, and practitioner guidance converges on a prepare-draft-approve-send pattern with mandatory human approval for irreversible, customer-facing sends.
  • 2026-Aug: Adoption metrics confirm scale (66% of service orgs now use AI agents, up from 39% in 2025) alongside a persistent production gap—79% adopted vs 11% in production per Knowmax, echoing named enterprise wins (Resona Group cut routine inquiry volume to 1/12 baseline at 92% autonomy; Klarna's 853 FTE-equivalent automation) against infrastructure-level constraints, including ISP domain-blocking thresholds that cap unmonitored autonomous email send within days of crossing spam/bounce limits.