Perly Consulting │ Beck Eco

The State of Play

A living index of AI adoption across industries — where established practice meets the bleeding edge
UPDATED DAILY

The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.

The Daily Dispatch

A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.

AI Maturity by Domain

Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail

DOMAIN
BLEEDING EDGEESTABLISHED

Decision support & reasoning frameworks

BLEEDING EDGE

TRAJECTORY

Stalled

AI that helps individuals structure decisions, evaluate options, and apply reasoning frameworks to complex choices. Includes decision matrix generation and pro/con analysis; distinct from feature prioritisation which applies frameworks to product decisions rather than general personal choices.

OVERVIEW

AI-assisted decision support has proven it can work in narrow, data-rich contexts -- but the barrier to reliable deployment is no longer capability, it is organizational readiness. That tension now defines its bleeding-edge status. Cox Communications achieved 7x ROI on multi-agent sales decisioning; Kai eliminated 99.5% false positives in security triage; Anthropic handles 95% of internal analytics queries through Claude. Yet across 2,400+ enterprise AI initiatives, the failure rate sits at 80%, and across 6,000+ executives surveyed, 90% report no discernible productivity impact from deployed systems. The deepest barrier is not technical: it is organizational ability to make decision-making processes explicit and design consistent human-AI workflows. Harvard AI Institute director (HBR, June 2026) identifies the core constraint: "organizations lack the ability to make decision-making processes explicit"—decision types, flows, criteria, trade-off rules, escalation conditions. Without that articulation, even capable AI systems fail at scale. Simultaneously, fundamental LLM reasoning flaws constrain what can be delegated: models collapse under cognitive load (Stroop: 91%→1% accuracy), spend more tokens on failed tasks (inverse of human judgment), commit 6-12% safety violations on critical tasks. These are architectural, not tunable. Individuals seeking structured decision support find AI can generate useful frameworks and surface tradeoffs, but organizational governance, workflow design, and human judgment remain essential for decisions that matter. The field has shifted from "can we build it" to "why aren't deployed systems delivering value."

July 2026 evidence confirms operationalization of structured reasoning frameworks alongside persistent reasoning limitations. RCT evidence shows clinical decision-support can improve physician reasoning (Kenya +18%) when integrated with proper human-AI design, establishing that capability alone is insufficient—design and governance determine outcomes. Simultaneously, frontier reasoning models exhibit <25% instruction-following compliance during reasoning traces, creating control gaps that limit structured decision-support deployment in governed environments. Legal frameworks are crystallizing around decision accountability: courts establish that decision authority cannot be delegated to AI, requiring mandatory documentation of reasoning review and source verification. Structured decision frameworks themselves (39-item cognitive framework catalogs, multi-agent orchestration plugins) are moving from instructional content to operationalized tooling deployed at scale via enterprise platforms, signaling ecosystem maturation in bounded decision domains.

CURRENT LANDSCAPE

Credible deployments share two patterns: tightly scoped problems with rich, structured data, AND explicit governance frameworks. Cox Communications (15,000 employees, 7x first-year ROI on sales decisioning) and Kai (99.5% false-positive elimination across 2.5M security findings) show operational maturity in bounded decision domains. California government deployment of Claude across state agencies (policy deliberation, Medicaid workflows, cyber triage) signals institutional-scale adoption. Yet a large pragmatic trial (Nature Medicine, 9,600 patients, Kenya) found AI-assisted clinical decision support improved documentation but failed to reduce treatment failure (2.2% vs 2.0%, P=0.13)—core outcome was null. This outcome-documentation gap is emblematic: AI can improve process quality and reduce manual load, but translating that into measurable decision improvement remains elusive.

Professional services adoption jumped to 40% organisation-wide in 2026 from 22% in 2025, according to Thomson Reuters survey of 1,816 professionals across 27 countries—but 91% report AI value shortfall despite regular use, and 90% demand reasoning that can be explained and defended. Only 18% of firms track ROI, and 40% report client confusion over AI-use policies. Adoption is outrunning accountability, and the measurement story reveals the depth: NBER study of ~6,000 executives found 90% report zero productivity impact from deployed AI; PwC CEO survey shows 56% experienced no revenue/cost improvement. The field is shifting from deployment metrics to business outcomes (MindFinders framework: 5 critical performance questions), recognizing that deployment activation is not adoption, and adoption is not value realization. June 2026 evidence crystallizes the adoption bottleneck: 72% of CEOs expect AI to support human-directed decisions (not autonomous), identifying human governance as the design requirement; 55% of professionals cite lack of structured human-AI workflows as adoption barrier; 82% of individuals are AI-ready but only 15% of enterprises can leverage capability, quantifying organizational framework gap. Mechanistic constraints remain unresolved: models collapse under cognitive load (Stroop task: 91%→1% accuracy as list length grows), spend MORE tokens on failed tasks (inverse of human judgment), and commit 6-12% safety violations on critical tasks. All tested LLMs hit working memory saturation at identical thresholds (20-30 parallel branches) regardless of model size, indicating architectural ceiling. Semantic perturbations trigger 28-45% answer-flip rates; chain-of-thought explanations are post-hoc narratives rather than reasoning logs. Legal precedent is setting guardrails: Australian Federal Court ruled that executives using public AI tools informally breach duty of care; multiple jurisdictions establish that decision accountability cannot be delegated to AI, requiring mandatory governance documentation. Regulatory deadlines—CCPA enforcement in effect, EU AI Act arriving August 2026—are forcing governance readiness whether organisations feel prepared or not.

June 2026 evidence clarifies critical adoption barriers beyond capability. Multi-agent reasoning frameworks (RecursiveMAS using extended thinking) achieve 65.5% accuracy with practical cost profiles and named use cases in legal and medical triage, yet 95% of enterprise AI agent pilots fail due to organizational sequencing (top-down deployment, weak training, IT gatekeeping) rather than technical limits. Confirmation bias emerges as a core mechanism: when AI agrees with humans' initial incorrect answer, reliance drops 64.5%, indicating humans update beliefs based on validation rather than evidence quality. Prospective studies on "human-in-the-loop" decision-making reveal an Augmentation Trap: hybrid human+AI teams performing worse than AI alone on venture forecasting (0.74 vs. 0.04-0.45 rank correlation), suggesting human judgment introduces idiosyncratic noise that degrades AI's signal extraction. Tool-integrated reasoning systems dramatically outperform neural-only approaches (86-94% vs. 24-42% accuracy), establishing a theoretical boundary: beyond 20-30 parallel decision branches, pure AI reasoning collapses and tool delegation becomes mandatory. Yet even accurate systems fail in adoption: users distrust overconfident AI due to Expectancy Violation Theory—humans expect experts to hesitate and consider alternatives; instant certainty triggers defensive disengagement. The emerging pattern: decision-support systems fail not at reasoning accuracy but at the human-organizational interface—adoption requires re-sequencing implementation (bottom-up first), managing confirmation bias, and designing uncertainty visibility that preserves human critical evaluation.

TIER HISTORY

ResearchNov-2022 → Apr-2024
Bleeding EdgeApr-2024 → present

EVIDENCE (164)

— Empirical evaluation of frontier reasoning models on real-world decision-making under uncertainty; none beat market baseline, revealing poor calibration in high-stakes forecasting.

— Technical synthesis showing AI reasoning explanations (chain-of-thought) systematically misrepresent actual model computation, creating compliance and safety gaps in decision-support systems.

— Multi-year field study documents critical framework failure: frontline workers defend AI decisions they neither created nor understand—directly addresses the accountability-understanding gap in deployed decision support.

— Multiple independent sources (S&P Global, Gartner, MIT, BCG, IBM) document organizational barriers to AI decision-system adoption; identifies resource allocation and accountability as binding constraints, not technical capability.

— Empirical study showing accuracy fell 27%→9%, confidence rose 30%→76%, judgment suspension collapsed 44%→3%; directly addresses core reliability and validation concerns in decision-support frameworks.

— Peer-reviewed field experiment with 758 real BCG consultants: AI decision-support delivered +12.2% productivity and +40% quality for tasks within capability frontier, but 19pp reduced correctness outside frontier—directly measuring when reasoning frameworks work and fail.

— Legal framework establishing AI decision-support governance requirement: mandatory documentation of reasoning review, source verification, and user responsibility—decision accountability cannot be delegated to AI systems.

— 32 frontier models tested in clinical reasoning show information utilization collapses 57%→26% under uncertainty, revealing core limitation: high reasoning capability does not ensure good decision outcomes without proper information-seeking behavior.

HISTORY

  • 2022-H2: First identified research surge in AI reasoning benchmarks (temporal, step-by-step, knowledge-graph) and human-in-the-loop frameworks; major failures in practice (Zillow, model brittleness); ~85% enterprise project failure rate documented; theoretical work on reasoning fallibility and the need for AI to express uncertainty.

  • 2023-H1: Research focus shifted to human-AI interaction challenges: overreliance on AI suggestions despite explanations, user dropout when AI feedback is unhelpful, and widespread concerns about accountability and risk. Evidence of adoption barriers in personal decision support remained dominant; no large-scale personal reasoning framework deployments documented.

  • 2023-H2: Research concentrated on three critical areas: (1) human-centered design frameworks for DSSs (PAAI questionnaire with 700+ participant validation), (2) trust and accountability barriers blocking clinical deployment (liability concerns, accuracy standards), and (3) underlying AI reasoning capabilities (foundation model survey). Evaluation gaps documented—despite renewed interest, empirical evidence on AI-CDS effectiveness remained scarce. Governance and risk analysis highlighted over-reliance, bias, and dynamic environment brittleness as core failure modes.

  • 2024-Q1: Adoption accelerated dramatically: ~90% of enterprises deployed AI for autonomous decision-making. Simultaneously, fundamental technical limitations became clearer—Apple research confirmed critical reasoning flaws in LLMs (GSM-Symbolic benchmark), and AI systems scored only 30% on novel reasoning tasks (ARC). Bias in operational AI decision systems documented in education. Revenue impact quantified: 6% average annual loss from underperforming models. The execution gap widened: implementations failed at generalization, bias mitigation, and reliability despite widespread organizational trust. Empirical evidence on decision support effectiveness remained sparse, leaving large-scale deployments without measured impact validation.

  • 2024-Q2: Real-world deployment evidence emerged, revealing persistent failures despite adoption breadth. A Dutch court case study showed AI decision-support in legal proceedings, while a Harvard-led RCT in Wisconsin courts found AI recommendations failed to improve bail decisions and judges rejected them 30%+ of the time. Experimental research with 1,403 participants confirmed overreliance remains endemic despite explainability efforts—workers align with biased AI recommendations up to 90% in hiring contexts. A 600-executive survey found 48% of AI projects paused or rolled back due to privacy, regulatory, and integration challenges. The landscape shifted from "can we build AI decision systems" to "why are deployed systems failing to improve decisions"—adoption at scale masked persistent technical and organizational gaps.

  • 2024-Q3: Research clarified three persistent barriers to effective human-AI decision-making: achieving complementarity, managing human mental models, and design choices that prevent cognitive overload. Healthcare case studies identified prerequisite frameworks for responsible AI-DSS (bias mitigation, human-centric learning loops, incremental trust-building). Critical assessments documented permanent risks in high-stakes decision contexts (military targeting, legal proceedings) due to hallucinations, brittleness, and inability to ensure regulatory compliance. Analyst forecasts predicted 30% project abandonment post-proof-of-concept by end of 2025, with organizations struggling to realize value despite major investments. The field continued to reconcile widespread enterprise adoption with persistent deployment failures, unresolved bias risks, and absence of clear impact metrics on decision quality.

  • 2024-Q4: Critical research published on technical and organizational barriers to reliable AI reasoning. New findings revealed overreliance persists in clinical decision-making despite trust calibration efforts, with physicians exhibiting diagnostic errors from AI misalignment. Healthcare professionals identified systemic ethical concerns (bias, transparency gaps, accountability deficits) in AI-CDSS deployments. Fundamental research showed state-of-the-art reasoning models exhibit overthinking and discard correct reasoning paths, undermining the assumption that larger models improve decision support. Legal sector adoption continued despite significant concerns: 43% of legal professionals observed bias, 37% feared unreliability. Clinical adoption surveys showed 76% of physicians now use LLMs for decisions yet 97% vet outputs, indicating cautious rather than confident deployment. By year-end 2024, the field had reached consensus that the core challenge is not building reasoning systems but deploying them safely and measurably—technical limitations in AI reasoning were well-documented, but practical implementation remained the bottleneck. Organizations continued investing despite unresolved risks, suggesting adoption momentum has decoupled from evidence of effectiveness.

  • 2025-Q1: Enterprise adoption continued but real-world reliability challenges intensified. UK government abandoned multiple welfare-system AI pilots (A-cubed, Aigent) due to scalability and reliability concerns—explicit signal of deployment failure in public sector decision support. Healthcare outcomes improved in targeted deployments: UK Health Security Agency achieved 90% accuracy in TB screening with 85% reduction in manual review workload. Research on adoption barriers showed 450 physicians in China identified multiple adoption pathways depending on hospital type and organizational context. Critical research revealed that human oversight alone is insufficient to prevent discrimination: EU study found human decision-makers equally likely to follow biased AI recommendations regardless of fairness-algorithm design. Fundamental reasoning limitations persisted: AI reasoning models continued exhibiting data bias, lack of common sense, and transparency failures that undermine high-stakes decision-making. The gap between pilot success and production scaling widened: isolated cases showed operational gains, but public sector abandonment and persistent bias findings suggested the field remained pre-scale.

  • 2025-Q2: Evidence revealed critical implementation gaps despite continued investment. Dermatology study (223 physicians) found AI support yielded only 1% accuracy improvement with low reliance (10%), indicating adoption barriers persist even in favorable clinical contexts. Defense deployments (Project Maven, UK autonomous targeting, Iron Dome) demonstrated real-world AI-DSS use but in high-stakes, tightly constrained settings. ChatGPT testing showed AI mirrors human decision-making biases including overconfidence and gambler's fallacy in half of scenarios, suggesting AI amplifies cognitive flaws rather than mitigating them. Expert Delphi consensus identified 34 critical implementation factors for healthcare AI-DSS, yet organizational capacity to execute remained limited. Industry analysis showed only 26% of companies have working AI products and 4% achieve significant ROI; Gartner predicted 40%+ project cancellations by 2027 due to unclear value and costs. Parallel evidence of high adoption breadth (93% of leaders report GenAI competitive benefits) masked low implementation depth and persistent execution challenges.

  • 2025-Q3: Research clarified fundamental and persistent technical limitations in AI reasoning: models performed no better than humans on novel problems and replicated cognitive biases including overconfidence. MIT analysis of 300 deployments found 95% of AI pilots failed to deliver value, with vendor solutions succeeding ~67% versus internal builds 33%—exposing both adoption and execution challenges. Consumer trust surveys (YouGov, 10K respondents) showed 52% comfort with AI for daily personal decisions but only 39% for financial decisions; humans retained override preference in 55%+ of scenarios. New tools for bias detection (CMU AIR) and structured decision frameworks (MCDM-based ModelSelect with 50 case-study validation) promised incremental rigor improvements yet could not address fundamental reasoning limitations. Research documented that AI actively degraded decision quality: executives using generative AI made worse forecasts than without it, highlighting the risk of overconfidence in AI-enhanced reasoning. The field remained characterized by adoption momentum decoupled from evidence of effectiveness, with organizations continuing heavy investment despite quantified failures and persistent technical barriers.

  • 2025-Q4: Deployment evidence revealed domain-specific outcomes: IBM achieved $4.5B productivity impact from agentic AI deployed to 270K employees; marketing decision-intelligence platforms reached 26-75% adoption with measurable ROI; UK Health Security Agency's AI-assisted TB screening achieved 90% accuracy. Yet critical limitations emerged across high-stakes domains: medical data gaps (EMR design flaws, not algorithmic limitations) constrained clinical decision-support impact; Indian judges warned of AI-fabricated legal judgments and hallucinations; government pilots stalled due to scaling and budget challenges; legal professionals documented persistent bias (43% observing bias, 37% fearing unreliability). Medical educators flagged overreliance risks: GenAI tools threaten critical thinking skill development and reinforce training data biases. Technical advances in reasoning (GPT-5.1 integration, causal AI frameworks) continued, yet ethics scholars debated justified use of black-box AI in high-stakes domains. By year-end, the field had consolidated around differentiation by domain: operational value in narrow contexts (marketing, logistics) versus persistent barriers and documented risks in broader organizational and high-stakes deployment scenarios. Adoption momentum remained decoupled from evidence of effectiveness, with organizations continuing investment despite quantified failures and unresolved deployment barriers.

  • 2026-Jan: Enterprise transition to operationalization emphasized data governance and architectural foundations; 62% of enterprises planning evolution to AI decision intelligence amid persistent 70-85% project failure rates and 42% initiative abandonment in 2025. Causal AI emerged as next-frontier addressing 74% faithfulness gap in existing systems. Clinical research documented error reduction (78% decline in guideline violations) through hybrid frameworks, yet deployment barriers remained: only 12% of executives reported both cost and revenue benefits; physician studies highlighted that reasoning cues must target high-discretion tasks where AI can add genuine value.

  • 2026-Feb: Multi-AI orchestration demonstrated operational feasibility (ARPIA 13-min data-to-strategy pipeline, causaLens enterprise deployments); professional services adoption jumped to 40% (2025: 22%), yet only 18% track ROI and 40% report policy confusion. Systematic LLM reasoning failures (Reversal Curse, Robustness Fragility, Working Memory Leaks) documented, undermining reliance on AI reasoning chains. Commercial AI-CDSS solutions still lack transparent training data and algorithm disclosure. Regulatory deadlines (CCPA Jan 2026, EU AI Act Aug 2026) drove governance platform launches. Across 2,400+ enterprise AI initiatives, 80.3% failed (33.8% abandoned, 28.4% deliver no value), with 95% GenAI pilots failing to reach production. Execution and governance remain bottlenecks, not capability.

  • 2026-Mar: Fundamental research documented persistent reasoning failures: CRYSTAL benchmark shows models skip 50%+ of reasoning steps (58% accuracy but only 48% reasoning recovery); BrainBench reveals stochastic reasoning gaps (6-16pp consistency variance even in top models); Stanford taxonomy classifies failures as architectural rather than scale-addressable. Real-world failures documented: NZ courts ruled AI-hallucinated legal citations may amount to obstruction; Deloitte refunded AUD 440K for AI-generated errors. Governance frameworks consolidated: RegTech expert consensus establishes human accountability cannot be delegated; KPMG legal analysis requires mandatory documentation of decision review. Practical deployment barriers clarified: reasoning models show 5x cost premium with performance ceiling at medium-complexity tasks (above which accuracy collapses). Field consensus solidifying around decision-support constraints: execution challenges and governance requirements are primary blockers, not reasoning capability gaps.

  • 2026-Apr: New empirical research confirmed architectural reasoning limits are scale-invariant: testing across 7 models (8B–235B parameters) showed all collapse at 20-30 parallel branches, and semantic variants of problems trigger 28-45% answer-flip rates — confirming brittleness is structural, not addressable by larger models. CMU testing of 14 leading LLMs (GPT-4, Claude 3, Gemini) found all fail simple logical contradiction detection, revealing that benchmark performance masks fundamental reasoning gaps. Harvard research added a new dimension: relational complexity causes accuracy to collapse when decisions require weighing multiple interacting factors simultaneously, directly constraining multi-factor analysis in healthcare and strategic contexts. Legal accountability frameworks tightened: Australian Federal Court (ASIC v Bekier) established that executives using public AI tools informally without verification breach their duty of care, reinforcing that decision accountability cannot be delegated to AI and mandating documentation of reasoning review. Latest evidence on personal decision support reveals critical tensions: targeted healthcare deployments show measurable value (RCT with 367 participants: 7.4-point satisfaction improvement, 50.7% vs 24.2% acceptance of AI recommendations), yet passive reliance on AI reasoning systematically erodes confidence in independent judgment and sense of authorship (behavioral study, 1,923 adults). Empirical adoption barriers remain severe: 9% of professionals trust AI for complex decisions despite 88% organizational adoption; 80% of workers reject enterprise AI tools (WalkMe, 3,750 professionals). Reasoning failure modes are now precisely characterized: hallucination rates span 22-94% across models (Stanford 2026 AI Index), with a documented "Reliability Gap" where capability scales 2-3x annually while reliability only 1.2-1.5x. Governance frameworks are consolidating: the SPEC framework achieves 89% accuracy vs 15% for unbounded RAG on incomplete-evidence scenarios by bounding AI confidence to evidential sufficiency; the CFA Institute articulates an epistemic anchoring principle that decision authority must remain in evidence-based human inquiry to avoid "knowledge-collapse equilibrium."

  • 2026-May: Research intensified focus on cognitive and systemic failure modes in AI-assisted decision-making. Wharton's Cognitive Reflection Test (1,300+ participants) demonstrated the core paradox: AI-correct advice improves accuracy 25pp, but AI-wrong advice degrades 15pp below baseline—worse than no AI. Overconfidence persists even when users know AI errs 50% of the time, indicating cognitive surrender rather than rational reliance. Metacognitive stability revealed as critical: empirical testing across 11 frontier models shows 8 collapse under adversarial pressure (30.2pp accuracy drops), with only Anthropic's Constitutional AI showing near-immunity—suggesting alignment-specific training is prerequisite for trustworthy reasoning, not achievable through standard RLHF. Enterprise adoption gaps widened: <20% of AI pilots reach production due to missing trust infrastructure (audit trails, explainability, liability frameworks); yet production deployments in bounded domains (insurance claims, pharma R&D) show 50%+ efficiency gains when governance frameworks are load-bearing. Real-world interaction data (5.5M instances) reveals expert-domain performance plateaued at 14-16% dissatisfaction despite scale; raw models generate confident false theories when given deliberately flawed premises. Most critical finding: brief AI assistance (10 minutes) systematically impairs independent problem-solving through cognitive offloading, reducing retention and analytical skepticism—unintended consequence suggesting tool design actively erodes reasoning autonomy. Governance literature consolidated: AI system abandonment driven primarily by organizational dynamics and resource constraints, not ethics concerns; clinical AI remains confined to pilots not due to model limitations but institutional capacity gaps. The field's consensus strengthened: decision-support reliability requires orchestration of governance, human oversight, design discipline, and alignment-specific training—capability alone is necessary but insufficient.

  • 2026-Late May: Enterprise deployment momentum accelerated with defined governance patterns. SAP Sapphire 2026 unveiled 224 specialized Claude agents handling autonomous decisions across finance, HR, supply chain, and procurement for hundreds of thousands of enterprise customers globally—largest-scale production deployment of autonomous decision-making systems. CLR-voyance clinical reasoning system demonstrated production-scale maturity: 6+ months in hospital operation, 84.91% accuracy with physician-validated outcome rubrics, drafting thousands of inpatient notes—proving structured decision-support can reach institutional deployment. Yet critical limitations emerged: meta-analysis of 5 RCTs (12,657 participants) found AI clinical decision support produces only small, marginal improvements in diagnostic accuracy; pooled effect size narrowly above zero; strongest signal in radiology, weakest in complex reasoning domains. Governance failure patterns crystallized: survey of 650 enterprise leaders found 78% ran AI agent pilots but only 14% scaled to production; 90% of agent codebases fail EU AI Act compliance scans; governance architecture (audit trails, escalation paths, ownership documentation) emerged as primary blocker, not model performance. Personal effectiveness frameworks showed concrete signals: executive decision-support (persistent context, standing priorities, weekly accountability reviews) demonstrated production viability; structured reasoning patterns (Clarify-Then-Act, Plan-Then-Execute, Human-in-the-Loop gates) identified enterprise reliability requirements. Research confirmed psychological decision-making frameworks improve AI reasoning: TU Berlin study showed Recognition-Primed Decision-Making and Data-Frame Theory prompts boost healthcare advice accuracy 13% → 30%, establishing that human decision-making structures significantly enhance AI decision quality. Institutional barriers remained decisive: ICML paper documented that technical viability alone insufficient for scaling; decision-support systems require institutional alignment across approvals, oversight capacity, fiscal sustainability, and regulatory readiness. The field consolidated around a fundamental insight: decision-support reliability depends primarily on governance, human oversight, and institutional capacity—not model capability alone.

  • 2026-Jun: Consensus solidified around organizational readiness as the binding constraint. Thomson Reuters survey (1,816 professionals): 91% report AI value shortfall despite 74% regular use; 90% demand reasoning that can be explained and defended. NBER study of ~6,000 executives: ~90% report zero productivity impact; PwC CEO survey: 56% experienced no revenue/cost improvement, only 12% significant benefits—a $2.5T "validation gap" between investment and outcomes. But concrete deployments prove bounded decision-support works: Cox Communications achieved 7x first-year ROI on multi-agent sales decisioning (15,000 employees); Kai eliminated 99.5% false positives in security triage (2.5M findings). Nature Medicine RCT (9,600 patients, 16 sites, Kenya): AI-assisted clinical decision support improved documentation quality but failed to reduce treatment failure (2.2% vs 2.0%, P=0.13)—emblematic outcome-documentation gap. The deepest barrier is organizational: only 15% of enterprises can leverage 82% of AI-ready individuals; 55% of professionals cite lack of structured workflows as adoption barrier. Harvard AI Institute director (HBR): "core constraint is organizational ability to make decision-making processes explicit." Mechanistic research exposures: Stroop task (PNAS Nexus) shows models collapse under cognitive load (91%→1% accuracy); reasoning models spend MORE tokens on failed tasks (arXiv:2606.26502, Cohen's d 1.47–3.13), inverse of human behavior; frontier models commit 6-12% safety violations on critical tasks (Arachne, 81,000-person survey). Reasoning capability advances documented (OpenAI disproving 80-year-old math conjecture, TTE-Flash, Recursive Language Models), yet hybrid architectures combining LLM planners with deterministic execution engines prove necessary for reliable analytics (trade-off: reasoning depth vs. factual grounding). Albabtain's organizational design study: AI augments or substitutes judgment based on design choices (transparency, override friction, performance metrics)—augmentation requires low-friction overrides, not AI compliance metrics. 72% of CEOs expect AI to support human-directed decisions (not autonomous), with trust/security concerns (31%) and data quality (34%) cited as barriers. Government-scale adoption signals: California partnership provides Claude across all state agencies (policy deliberation, Medicaid workflows, cyber triage). The field's 2026 consensus: decision-support reliability requires governance, explicit workflow design, and human judgment integration—not model accuracy alone.

  • 2026-Jul: The validation gap deepened with harder numbers. NBER (~6,000 executives) and PwC CEO survey together document 90% reporting zero productivity impact and 56% seeing no revenue or cost improvement — a $2.5T blind spot. Against that backdrop, two production deployments documented genuine bounded ROI: Cox Communications 7x first-year return on multi-agent sales decisioning and Kai's 99.5% false-positive elimination in security triage. Architectural research confirmed models cannot self-regulate reasoning effort: reasoning tokens scale inversely with task success (Cohen's d 1.47–3.13), and the Stroop task collapse (91%→1% accuracy under cognitive load) is now confirmed as architectural rather than tunable. Organizational design research (Albabtain) established that augmentation vs. substitution of human judgment is determined by governance choices — low-friction overrides and judgment-rewarding metrics — not model capability. Mid-July evidence added both operational tooling and governance boundaries: a three-country RCT of 249 physicians found AI improved clinical reasoning (Kenya +18%) when paired with proper human-AI design, while peer-reviewed ACL research found frontier reasoning models achieve under 25% instruction-following compliance during their own reasoning traces — a control gap limiting governed deployment. Structured decision frameworks moved from instructional content toward operational tooling: an open-source Claude Code plugin implemented multi-agent debate to surface reasoning blind spots, and a 39-framework cognitive-technique catalog shipped via the Claude Code marketplace. Legal commentary reinforced that decision accountability cannot be delegated to AI, and a 1,000-dealmaker survey found 62% consider human-only decisions indefensible yet only 22% delegate final decisions to AI — underscoring the governance boundary.

  • 2026-Aug: New evidence sharpened the accountability-capability gap: a real-world benchmark found frontier reasoning models fail to beat market baseline on decision-making under uncertainty, chain-of-thought explanations were shown to systematically misrepresent actual model computation, and a peer-reviewed study found AI advice cut human accuracy from 27% to 9% while nearly tripling confidence — while an HBR field study documented frontline workers defending AI decisions they neither created nor understood. A 758-consultant BCG field experiment quantified the capability boundary precisely (+12.2% productivity/+40% quality inside the AI's competence frontier vs. a 19pp correctness drop outside it), and multi-source analysis (S&P, Gartner, MIT, BCG, IBM) confirmed 42% of companies abandon AI initiatives due to organizational rather than technical barriers.