The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail
AI that classifies ticket intent, detects sentiment and escalation risk, identifies language, and routes accordingly. Includes multi-label topic tagging and escalation prediction; distinct from ticket routing which assigns based on rules rather than understanding content.
Ticket intelligence — AI that classifies support tickets by intent, sentiment, and escalation risk — is technically mature but facing a customer acceptance crisis. Production systems achieve 87-92% sentiment accuracy and 85-95% intent classification rates on routine cases, with major vendors (Google, AWS, Zendesk, Sprinklr) shipping GA features used at scale (10B+ predictions/day). Yet the technology shows critical limitations: frontier LLMs plateau at 38-40% accuracy on fine-grained emotion detection, sarcasm detection drops to 63% accuracy (vs 80-90% on straightforward text), and real-world deployments expose a gap where 98% accuracy on structured tasks (password resets, order status) collapses to 61% on emotional or complex intents requiring empathy or policy interpretation. Production deployments face compounding challenges: LLM-based classification (800ms latency, $400K per project) underperforms simple rule+ML rebuilds (2ms latency, 89% accuracy), and naive intent-drift detection creates false-positive floods at scale. More concerning, independent data reveals fundamental customer rejection—64% of consumers prefer companies avoid AI service entirely, and 53% report they would switch competitors after poor AI experiences. Organisations continue investing (84% Fortune 500 adoption, Zendesk Professional tier now democratizes sentiment access as of July 2026), but this leading-edge practice is defined by the widening tension between technical capability and customer acceptance, with the majority struggling to translate sentiment/intent signals into measurable business outcomes.
Zendesk, IBM Watson, Google Cloud, AWS Comprehend, Sprinklr, and Salesforce all ship GA intent detection and sentiment analysis. August 2026 product momentum: Zendesk expanded Intelligent Triage (renamed "Topics") to Professional tier effective July 2026, democratizing sentiment analysis from premium-only to standard tier; Sprinklr operates 10B+ sentiment predictions/day across 100+ languages; AWS released entity-level sentiment for aspect-based routing (e.g., positive product, negative service in same ticket); Salesforce integrated email sentiment into Dynamics 365; Amazon Comprehend provides real-time entity-based sentiment detection. Real-world deployments reveal operational maturity dependencies: Fin AI achieved >98% accuracy on escalation routing; AssemblyAI cut first-response time from 15 minutes to 23 seconds; 72% of IT organizations deployed AI-assisted classification with 40-60% ticket deflation at scale. However, consultant benchmarks show deflection-rate variance: rule-based routing achieves 15-25%, while hybrid NLU systems reach 42-58%, with guardrails (confidence thresholds, output verification, PII masking) critical for production stability. A fintech company reduced sentiment-analysis costs from $12,000/month to $0.30/hour via fine-tuned Llama 3 8B while maintaining 95% accuracy, demonstrating cost-optimization and open-weight model viability.
The customer acceptance crisis persists despite tactical improvements. Independent research documents fundamental accuracy limitations: sarcasm detection degrades to 63% vs 80-90% on straightforward text; sentiment-rating mismatches occur in 20% of cases, requiring topic-level analysis; frontier LLMs plateau at 38-40% on fine-grained emotion detection. Production failures expose operational maturity requirements: LLM-based classification systems deployed at $400K cost and 800ms latency underperform simple rule+ML rebuilds (2ms latency, 89% accuracy); naive intent-drift detection creates 0.40 false alerts per endpoint per day, overwhelming SOC teams at scale. Customer sentiment data remains the binding constraint: 64% of consumers prefer companies avoid AI service entirely; 53% would switch competitors after poor experiences; 40% of agentic AI projects face cancellation within 12 months. The practice remains leading-edge because technical capability (87-92% sentiment accuracy, >98% on structured intents) is proven at scale, but organizational execution (only 6% reach 1-year payback) and customer acceptance remain severely constrained—despite 84% Fortune 500 adoption intent and Zendesk's Professional-tier democratization, deployment execution and outcome linkage remain the primary barriers to sustained adoption.
— Peer-reviewed benchmark (100K+ app reviews) finds 20% mismatch between star ratings and sentiment text. Highlights topic-level sentiment requirement and multilingual deployment challenge, with LLMs matching traditional models only on longer text.
— Zendesk Professional+ tier (July 2026) includes sentiment analysis on 5-point scale, topic/intent classification (~150 languages), language detection, entity extraction, with confidence levels enabling high-confidence automation and integration with routing/SLA policies.
— Named fintech deployment: fine-tuned Llama 3 8B reduced sentiment-analysis costs from $12,000/month to $0.30/hour while maintaining 95% accuracy. Production rollout demonstrates open-weight model fine-tuning viability and substantial ROI for sentiment at enterprise scale.
— Vendor-credible critical analysis documents four failure modes: thin verbatims starving sentiment of signal, inability to ask follow-ups, context/causality loss via aggregation, and sarcasm detection gaps. Essential negative signal on what text analytics cannot do.
— Post-mortem of failed LLM ticket classification project ($400K spent, 40% misclassified); successful rebuild with rules+TF-IDF achieved 89% accuracy. Critical negative signal: LLM-based classification costs 800ms vs 2ms for simple models, latency mismatch with production requirements.
— Peer-reviewed research (IJDSN 2025) documents sentiment analysis accuracy: 63% on sarcasm vs 80-90% on straightforward text. Critical limitation showing 20-30 point performance drop and missing ~30% of cases at sarcasm accuracy on 10K reviews/month.
— Analysis of 28,000 production sessions shows naive intent-drift detection generates 0.40 benign false alerts per endpoint/day (4K/day at scale). Documents production maturity challenge: intent detection works but requires behavioral baselining sophistication beyond topic classification.
— Consultant playbook with specific deflection benchmarks: rule-based 15-25%, hybrid NLU 42-58%. Documents guardrail necessity (88% classifier + confidence floor + output verification reduces hallucination from 5-15% to 2-4%), cost estimates $60-110K pilot with ~1-year payback.