The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail
AI that autonomously scores contract risk, generates assessment reports, and recommends accept/reject/negotiate decisions. Includes automated risk scoring and recommendation generation; distinct from risk flagging which highlights issues for human assessment rather than making recommendations.
Autonomous contract assessment -- AI that scores risk, generates reports, and recommends accept/reject/negotiate decisions -- has moved from bleeding-edge experiment to leading-edge production standard in high-volume triage, yet the practice faces a hard tier ceiling beyond routine work. Global adoption has normalized rapidly: 92% of lawyers across 10 countries now use AI daily; 87% of general counsel employ AI; 52% of in-house teams actively use or evaluate contract review AI (quadrupled since 2024). Deployments deliver measurable ROI for high-volume screening -- 40-60% efficiency gains, 87% time reductions (LegalFly: 2-hour to 15-minute reviews), 300-450% reported ROI. Leading-edge production deployments now demonstrate autonomous end-to-end workflows: Fortune 50 semiconductor company deployed agentic AI for autonomous contract ingestion, extraction, risk detection, and escalation; M&A dealmaker adoption rising from 16% to 48% within 3 years with 12% more material risks caught than manual review. Yet foundational barriers prevent tier advancement. Autonomous assessment capability diverges sharply from autonomous deployment: Harvey's Legal Agent Benchmark reveals 13.3% all-pass autonomy for end-to-end matters (vs 93.4% for isolated tasks), demonstrating the autonomy-maturity gap. Institutional knowledge barriers are structural: autonomous systems fail systematically without embedded company negotiation standards, risk tolerance, and escalation rules. Data infrastructure gaps affect 78% of enterprises -- 54% lack bidirectional system synchronization, forcing reliance on generic models rather than fine-tuned deployment. Regulatory frameworks and deployment realities impose binding constraints. EU AI Act (Annex III) classifies autonomous contract assessment as high-risk, mandating human oversight mechanisms and conformity assessment by December 2027; Article 14 design requirements mean oversight must be architectural (human-in-the-loop, not rubber-stamped approval), making fully autonomous assessment regulatory non-compliant in major markets. Accuracy ceilings remain: realistic benchmarks show 85% recall for risk flagging and 80% factual accuracy, not the 95%+ marketing claims vendors promote. Hallucination incidents have spiked: 1,200+ documented cases globally, with $145K in court sanctions in Q1 2026 alone. The tier-defining tension is structural: the practice excels at high-volume first-pass screening where human review is downstream, but regulatory requirements, accuracy-on-complexity ceilings, and institutional knowledge barriers prevent autonomous decision-making on contested or complex agreements without solving governance, fairness, and liability exposure.
Production deployments at scale demonstrate the economic momentum, though maturity diverges sharply between routine screening and autonomous decision-making. Icertis Vera (June 2026 GA) introduces portfolio-wide autonomous risk assessment against business events across 1/3 Fortune 100 customer base; Microsoft Cloud Operations achieves 2-hour-to-15-minute contract-to-PO cycles through autonomous Icertis-SAP Ariba integration. LegalMind AI's production deployment automates 70% of workload across 3,400 contracts monthly through an eight-step autonomous pipeline (normalization, extraction, template comparison, risk scoring, compliance checking, summary, queue prioritization), compressing 4.2-hour reviews to 38 minutes and reducing infrastructure costs 76%. Persistent Systems deployed agentic AI for a Fortune 50 semiconductor company managing 1,000+ HR and supplier contracts with autonomous ingestion, risk detection, and high-risk case routing without manual review intervention. Skopx case studies document mid-market deployments achieving 80% time reduction (2.1 hours to 25 minutes) and financial services achieving 97.3% regulatory compliance accuracy with 0.89 correlation to attorney assessments. Concord's engine processes 10k+ contracts monthly with 94% autonomous risk-spotting accuracy, compressing review from 92 minutes to 26 seconds per contract. Inkvex's independent validation on 327 real contracts confirmed 94% catch rate of high-severity flags with 99% catch on auto-renewal clauses and 95% on liability caps. LegalFly's autonomous assessment (Agristo case study) compresses 2-hour to 15-minute reviews (87.5% reduction) through playbook-based autonomous clause scanning and risk prioritization. Vendor ecosystem consolidation is deepening: Icertis serves 250+ Fortune 500 customers with $350M ARR and 30%+ Fortune 100 penetration; LinkSquares reports 1,300+ teams managing 13M contracts with 800k+ hours saved. M&A dealmaker adoption rising from 16% to 48% within 3 years, catching 12% more material risks (assignment restrictions, auto-renewal, liability caps) than control groups. Global adoption has normalized: Wolters Kluwer's 810-lawyer survey across 10 countries shows 92% use AI daily, 62% report 6-20% time savings, and 61% are confident in AI-driven workflows. July 2026 data reveals divergence between usage and confidence: independent survey of 534 lawyers across 75 countries shows 86% use AI weekly but only 18% report high confidence in purpose-built legal AI, with verification/traceability cited as critical missing capability.
However, autonomy maturity claims diverge sharply from deployment reality. Harvey's Legal Agent Benchmark (June 2026) demonstrates the autonomy gap: frontier LLMs achieve 93.4% accuracy on isolated contract tasks but only 13.3% all-pass autonomy on end-to-end matters where every step must succeed without human intervention. Stanford Law School research demonstrates AI outperforms law professors 75% on contract law reasoning, but this task-isolated clarity masks persistent autonomy barriers. Institutional knowledge is the core barrier: autonomous assessment systems systematically fail without embedded company negotiation standards, risk tolerance, and escalation rules. Law firm analysis documents failures when systems lack understanding of which customers justify exceptions, which issues require which department sign-off, and which fallback positions are acceptable by deal size. Accuracy calibration requirements are strict: realistic benchmarks show 85% recall for risk flagging and 80% factual accuracy, compared to vendor claims of 95%+ that do not survive real-world contract complexity. July 2026 rigorous analysis reveals clause-by-clause accuracy variance is the core barrier: precision drops to 44% on complex clauses (covenant, IP assignment, ROFR) while boilerplate approaches ceiling, disproving blended accuracy claims and establishing why autonomous assessment systematically fails on high-stakes terms. Axiom's survey of 500+ legal leaders across 8 countries shows only 31% at wide-scale autonomous deployment despite 96% adoption in some form; July 2026 follow-up research shows only 7% of legal teams moved beyond pilots to full autonomous agent implementation, with governance, data privacy, and workflow complexity cited as primary barriers. Data infrastructure gaps affect 78% of enterprises: 54% lack bidirectional system synchronization, 42% adopt AI but store contracts across fragmented repositories, forcing reliance on generic models rather than fine-tuned deployment. Conga's survey of 250 CLM professionals found 92% still require human review of AI outputs with governance and trust as biggest scaling barriers. Stanford AI Index 2026 benchmarks hallucination rates of 22-94% across 26 leading models, with best-performing models delivering incorrect answers in roughly 20% of responses—a reliability ceiling that directly constrains autonomous legal assessment viability for high-stakes decisions. Economic viability barriers are understated: analysis of production agentic systems shows agents passing all quality metrics but incurring per-success costs exceeding human baseline, making autonomous workflows economically underwater despite technical correctness—a failure mode obscured by technical evaluation frameworks. The autonomy-maturity divergence is not a technology problem but a structural reality: production systems marketed as autonomous are operationally dependent on institutional knowledge integration, human oversight gates, escalation rules, and data infrastructure maturity that bind them to hybrid human-AI architectures rather than genuine autonomy.
The advancement barriers, however, are hardening rather than softening. Brittney Ball's April 2026 research documents 1,200+ AI hallucination incidents in legal proceedings globally (roughly 10 per day), with $145K in court sanctions in Q1 2026 alone and indefinite attorney suspension for filing 57 defective AI-generated citations. Thomson Reuters analysis identifies the strategic risk: 80% of legal professionals see AI as transformational, yet only 38% expect near-term organizational change, and Gartner projects over 40% of agentic AI projects will be discontinued by 2027. Real-world deployment data shows 17-34% error rates in production despite 95%+ accuracy benchmarks; governance and infrastructure gaps prevent 78% of agentic pilots from reaching production. Regulatory constraints compound the ceiling. EU AI Act Article 14 establishes that human oversight is a design requirement, not a staffing workaround; organizations achieving 98%+ approval rates without meaningful human judgment are regulatory non-compliant. Academic research demonstrates the underlying tension: higher autonomy compresses viable agency in regulated contexts—organizations cannot simultaneously maximize autonomous decision-making and satisfy regulatory human oversight mandates. Independent benchmarking establishes realistic performance ceilings: 50-75% of clause changes autonomous in steady state (vs. vendor claims of higher coverage), with 70-80% playbook coverage meaning 20-30% of changes always require human judgment. The bias vulnerability identified in January 2026 law review research persists: autonomous scoring tools systematically favor corporations over individuals in negotiation, creating direct liability exposure. Contractual data-access restrictions force reliance on generic models rather than fine-tuned deployment. Autonomous decision-making on complex or disputed agreements remains out of scope for all but the most risk-tolerant teams.
— Comparative testing of 10 autonomous contract assessment platforms showing mature ecosystem with risk scoring, missing clause detection, and autonomous analysis capability.
— Market sizing: contract AI grows to USD 14.76B by 2031 (28.52% CAGR); identifies hallucinations and human review needs as restraint impact (-0.9% CAGR), acknowledging accuracy and governance barriers.
— Independent testing of 5 platforms on 25 real contracts: Ironclad 93% detection (81/87 issues). Demonstrates autonomous assessment accuracy and capability at scale.
— Independent MIT CSAIL/Harvard benchmarking: contract review error rates 6-13% standard, 15-22% specialized; exposes hallucination risks and accuracy limitations vs vendor claims.
— Named deployments: Salesforce $5M contract automation, Klarna $60M; critical negative signal: 22% of AI agent projects report negative ROI at 12 months, mostly scope creep.
— Flank Record launch: autonomous agentic contract system with real-time obligation monitoring and autonomous risk-profile extraction, representing production deployment of self-scoring agents.
— Technical analysis of systematic bias in LLM scoring (position, verbosity, self-preference, confidence bias); essential infrastructure gap for trustworthy autonomous contract assessment.
— 92% AI tool adoption but governance gap: 51% lack AI error policies, 96% cite liability/accountability concerns as expansion blocker—governance constraints bind autonomous deployment.