Decision support & reasoning frameworks
200 evidence items
AI that helps individuals structure decisions, evaluate options, and apply reasoning frameworks to complex choices. Includes decision matrix generation and pro/con analysis; distinct from feature prioritisation which applies frameworks to product decisions rather than general personal choices.
Overview
AI-assisted decision support has proven it can work in narrow, data-rich contexts, but the barrier to reliable deployment is no longer capability—it is organizational readiness, judgment preservation, and accountability. Cox Communications achieved 7x ROI on multi-agent sales decisioning; Kai eliminated 99.5% false positives in security triage; Anthropic handles 95% of internal analytics queries through Claude. Yet across 2,400+ enterprise AI initiatives, failure rates sit at 80%, and 90% of executives surveyed report no discernible productivity impact. The barrier is not technical: it is organizational ability to articulate decision processes, design consistent human-AI workflows, and maintain human judgment throughout. Even when AI reasoning is demonstrably superior, humans reject it when it contradicts their prior answer; higher confidence in AI correlates with lower critical thinking; and decision briefings can be factually correct yet unsafe if they suppress uncertainty, misstate causality, or omit auditable reasoning. Simultaneously, fundamental LLM reasoning flaws constrain what can be delegated: models collapse under cognitive load (Stroop: 91%→1% accuracy), fail at sequential decision-fork reasoning (59.7% accuracy on frontier models), and commit 6-12% safety violations on critical tasks. These are architectural, not tunable. The field has shifted from "can we build it" to "how do we scale safe, accountable decision support whilst preserving human judgment."
Current Landscape
Credible deployments in 2026 share two characteristics: tightly scoped problems with rich structured data, and explicit governance frameworks. Cox Communications achieved 7x first-year ROI on sales decisioning; Kai eliminated 99.5% false positives across 2.5M security findings; Xylem onboarded 15,000 employees with 4,000 daily active users and estimated $70M revenue opportunity and $25M savings; One New Zealand deployed 50+ decision-support agents, reducing audit-planning time 60% and risk-matrix preparation from two days to under half a day. California government scaled Claude across state agencies for policy deliberation, Medicaid workflows and cyber triage. Yet a large pragmatic RCT (Nature Medicine, Kenya, 9,600 patients) found AI-assisted clinical decision support improved documentation but failed to reduce treatment failure. Professional services adoption reached 40% organisation-wide (up from 22%), but 91% report value shortfall, 90% demand explainability, and only 18% track ROI. Governance is the dominant failure mode: 87% of enterprises delayed deployments due to data governance risks, and 74% rolled back agents due to governance failure (data leakage 30.7%, hallucination 20.8%, auditability 16.8%). Technical debt accumulates rapidly—organizations spawn ungoverned agent variants without unified ownership of prompts, data sources or evaluation logic, creating untraceable errors. A September 2026 high-stakes failure: a US military intelligence analyst's AI synthesized classified and open-source intelligence to fabricate a nuclear-weapons claim, which the chatbot formatted as an authoritative report circulating across command channels; armed personnel prepared to board a ship before the error was caught. Governance frameworks are maturing to address this: decision briefings require five accountability dimensions (citation-to-claim entailment, causal-language discipline, uncertainty preservation, action appropriateness, human accountability); decision matrices allocate rights by ambiguity and risk; practitioner frameworks distinguish pattern-work delegable to AI from value-work requiring judgment, with safety architecture embedded. Yet adoption remains bottlenecked: 72% of CEOs expect AI to support human-directed decisions (not autonomous), identifying human governance as the design requirement; 55% of professionals cite lack of structured workflows as adoption barrier; and 82% report individual AI-readiness whilst only 15% of enterprises operationalize capability. A 3,700-participant study found AI recommendations amplified accuracy when aligned (64%→87%) but crashed below baseline when misaligned (34%), establishing that decision-support only functions reliably in bounded domains with explicit value alignment. Fundamental LLM constraints remain: models collapse under cognitive load (Stroop: 91%→1% accuracy), fail at decision-fork reasoning (frontier models 59.7% accuracy; larger reasoning budgets provide no improvement), and commit 6-12% safety violations on critical tasks. Legal frameworks are crystallizing: decision authority cannot be delegated to AI; executives using public AI tools informally breach duty of care; and mandatory documentation of reasoning review and source verification is required. Australia's Fair Work Commission condemned AI decision-support failures (40% of cases involved AI), mandating disclosure requirements, signalling regulatory consolidation around accountability and judgment preservation.
Tier History
Evidence (200)
— Benchmark showing frontier reasoning models answer only 59.7% of 502 decision-fork questions correctly, with larger reasoning budgets providing no improvement—exposing architectural constraint on AI-delegated decision-making.
— MIT CISR framework allocating decision rights between humans and agents on ambiguity × risk matrix, exemplified by One New Zealand's 50+ agents reducing audit-planning time 60% and risk-matrix preparation from 2 days to under half a day.
— Practitioner case documenting how a procurement contract-summarization agent spawned 27 variants in a year without unified ownership of prompts, data sources or evaluation logic, creating untraceable hallucinations and blocking scaling.
— High-stakes failure in which US military analyst's AI synthesized intelligence to fabricate a nuclear-weapons claim, formatted as authoritative report, came within minutes of triggering a ship boarding before error was caught.
— Peer-reviewed five-dimension accountability framework (citation-to-claim entailment, causal-language discipline, uncertainty preservation, action appropriateness, human accountability) establishing that factually correct briefings can still be unsafe without explicit governance.
195 more · latest 2026-09-15 →
— Named enterprise deployment onboarding 15,000 employees with 4,000 daily active users, estimated $70M revenue opportunity and $25M savings potential, demonstrating decision-support systems reaching organizational scale.
— Wharton analysis documenting judgment erosion from AI reliance, with Microsoft study showing higher AI confidence correlates with lower critical thinking, and BetterUp data quantifying 'workslop' costs ($186/employee/month).
— HEC Paris analysis of Sibony & Hazan framework showing human override of superior AI (physician case: AI 90% accurate, human alone 74%, human override 76%) destroys value, establishing fundamental design tension in human-AI decision-making.
— BCG survey of 70 C-suite leaders: 50% observing judgment/problem-solving skill decay with 3-5 year business impact expected; 90% report employees stopped checking AI work; Shell's protective redesign demonstrates remediation through deliberate reasoning preservation.
— Ada Lovelace Institute and NICE partnership: independent research on public-sector AI adoption barriers identifies decision-support risks (reliance, accountability, bias) and proposes sociotechnical governance framework for iterative evaluation.
— Harvard field experiment with 228 senior reviewers found automation bias overwhelming: reviewers rejected expert-approved projects following AI recommendations, and explanations paradoxically worsened bias; proposes CAST framework to preserve human judgment.
— Meta-analysis of MIT, RAND, S&P Global, and Gartner studies: 95% of AI pilots report zero ROI, 42% abandon before production; root cause is organizational governance and decision-framework gaps, not model capability.
— Production deployment across tens of thousands of users handling 36K+ historical queries; multi-agent decision-support system with full audit trail, context management, and adaptive output selection addressing enterprise trust and accessibility.
— Comprehensive rollback ledger documents 22 halted AI deployments with named organizations and specific governance failure modes (reward-hacking, sandbox escape, privilege escalation, moderation collapse) demonstrating autonomous decision-support unreliability.
— Poland Supreme Administrative Court case I FZ 104/26 documents AI hallucination in legal decision-support: counsel cited genuine case numbers but attributed fabricated propositions, exposing critical reasoning failure mode and professional liability in high-stakes decisions.
— PACT benchmark of 22 models across 12 regulated domains found even top models violate compliance rules 6-10% at baseline, escalating 65% under realistic workplace pressure, identifying architectural compliance-failure mode in high-stakes decision-support.
— Meta's CORAL framework demonstrates closed-loop decision-optimization at production scale across two billion-user platforms, with A/B-tested improvements in engagement and cost efficiency via constrained reasoning over persistent memory.
— AvePoint survey of 3,235 leaders: 87% delayed AI deployments due to data governance risks (not model quality), 40.7% canceled GenAI rollouts in 2026 (up 31.7% in 2025), exposing governance and ownership failures as primary scalability blockers.
— Revenue AI Report survey of 2,527 enterprises (Jan-Feb 2026): 74% rolled back/shut down deployed AI agents due to governance failure; paradoxically, 81% with 'mature guardrails' still rolled back, indicating governance maturity lag.
— Salesforce/Anthropic partnership operationalizes 37 prebuilt sales decision-support skills (deal health, pipeline review, meeting prep) with governance model routing actions through Salesforce business rules, demonstrating enterprise decision-support system maturity.
— Fair Work Commission tribunal documents AI decision-support failure; 40% of cases involved AI (adoption metric); institutional response mandates October 2026 disclosure requirement—governance driving adoption readiness.
— Framework distinguishing AI-powered drafting from human-led deciding; identifies three irreducible human elements (empathy, accountability, moral responsibility) establishing decision-support boundaries.
— General availability Claude Skill implementing structured reflection framework; vendor investment in decision-support reasoning architecture signals market maturity and recognition of governance-as-feature design requirement.
— University of Bayreuth study (n=3,700): AI recommendations amplified correct decisions (64%→87%) but crashed accuracy below baseline when misaligned (34%); conflict-of-interest disclosure offered minimal protection.
— Salesforce consulting framework documents staged agent maturity (Diagnose-Suggest-Act-Orchestrate) with human approval gates; named deployment achieved 50%+ case deflection after structured framework application.
— Production personal decision-support framework dividing pattern-work (AI) from value-work (human); safety architecture uses read-only Plaid access; documents failure modes and interrogation protocol for personal effectiveness.
— 30+ reproductions of constraint-preservation failure in Gemini: model violates explicit logical rules despite acknowledging them—fundamental reasoning architecture gap blocking governed decision-support deployment.
— US Census Bureau survey of 55% worker AI adoption with time-savings metrics; 31% report 1-2 hour savings, providing independent government baseline for real-world decision-support usage patterns.
— Testing across 9 LLMs reveals quality-dependent position bias and order-reversal can flip underlying preferences; fundamental brittleness in AI reasoning for decision-making across models.
— Maven handles 5,000+ targeting decisions per day with proven operational success; yet documents 8 institutional barriers limiting wider adoption—demonstrates deployment feasibility constrained by organizational factors.
— GoTo survey of 2,500 employees: 70% admit using AI for high-stakes decisions; 50% report overreliance; 39% believe overreliance erodes skills—documents widespread adoption with documented skill degradation.
— 81% of physicians use AI professionally; FDA narrowed oversight definition while shifting liability to clinicians; absence of human-centered governance design enables autonomy erosion in deployed decision systems.
— Real deployment: 1,744 U.S. hospitals running ambient AI for clinical decisions; burnout dropped 51.9% to 38.8%; but only 10% have formal AI oversight boards—production adoption without governance maturity.
— Survey of UK adults: 5M+ self-reported harms from incorrect AI advice; 43% use chatbots for personal decisions; 18% act without second opinion—widespread production use with documented negative outcomes.
— Nature Medicine study: explainability helps experts but creates automation bias in novices; one-size-fits-all explanations mislead users with less domain knowledge—interface design determines decision-support outcomes.
— Harvard/MIT study (1,280 participants) shows reinforcement-learning adaptive decision-support improves accuracy and reduces overreliance; demonstrates design choices can enable human-AI complementarity.
— Empirical evaluation of frontier reasoning models on real-world decision-making under uncertainty; none beat market baseline, revealing poor calibration in high-stakes forecasting.
— Technical synthesis showing AI reasoning explanations (chain-of-thought) systematically misrepresent actual model computation, creating compliance and safety gaps in decision-support systems.
— Multi-year field study documents critical framework failure: frontline workers defend AI decisions they neither created nor understand—directly addresses the accountability-understanding gap in deployed decision support.
— Multiple independent sources (S&P Global, Gartner, MIT, BCG, IBM) document organizational barriers to AI decision-system adoption; identifies resource allocation and accountability as binding constraints, not technical capability.
— Empirical study showing accuracy fell 27%→9%, confidence rose 30%→76%, judgment suspension collapsed 44%→3%; directly addresses core reliability and validation concerns in decision-support frameworks.
— Peer-reviewed field experiment with 758 real BCG consultants: AI decision-support delivered +12.2% productivity and +40% quality for tasks within capability frontier, but 19pp reduced correctness outside frontier—directly measuring when reasoning frameworks work and fail.
— Legal framework establishing AI decision-support governance requirement: mandatory documentation of reasoning review, source verification, and user responsibility—decision accountability cannot be delegated to AI systems.
— 32 frontier models tested in clinical reasoning show information utilization collapses 57%→26% under uncertainty, revealing core limitation: high reasoning capability does not ensure good decision outcomes without proper information-seeking behavior.
— Open-source Claude Code plugin implementing multi-agent orchestration through structured debate patterns; replaces single-model reasoning with disagreement and evaluation criteria to surface blind spots in decision-making.
— Peer-reviewed ACL study shows frontier reasoning models achieve <25% instruction-following compliance during reasoning traces—fundamental control gap limiting reliable structured decision-support deployment in governed environments.
— 1,000 dealmakers across 27 countries: 62% believe human-only decisions indefensible, but only 22% delegate final decisions to AI—establishes governance boundaries and shows decision authority requires human accountability for consequential choices.
— Synthesis of Stanford analysis of 51 successful enterprise AI deployments identifies executive sponsorship paired with clear measurable business objectives defined before model work as the most consistent adoption accelerator.
— 39 vetted cognitive and reasoning frameworks (First Principles, Bayesian Updating, Theory of Constraints) packaged with rigorous evaluation pipeline; production deployment via Claude Code marketplace—operationalizes structured decision frameworks at scale.
— RCT of 249 physicians across three countries shows AI improves reasoning (Kenya +18%) when integrated with proper human-AI design; demonstrates both deployment potential and critical role of design in clinical decision-support effectiveness.
— NBER study of ~6,000 executives: ~90% report no discernible impact on productivity from AI; PwC CEO survey shows 56% experienced no revenue/cost improvement with only 12% reporting significant benefits, documenting persistent gap between deployment and measurable decision-support outcomes.
— California partnered with Anthropic for Claude deployment across all state agencies for policy deliberation, Medicaid workflows, and cyber triage, demonstrating institutional-scale decision-support adoption with government endorsement, though outcome metrics remain unquantified.
— Survey of 2,500 CEOs reveals 72% expect AI to support or execute work under human direction and governance, not autonomous decisions; identifies trust/security concerns (31%) and data quality (34%) as barriers to safe decision-support deployment.
— Large reasoning models spend MORE tokens on failed tasks versus solved ones (inverse of human behavior), with effect size Cohen's d 1.47–3.13, meaning longer reasoning chains are not reliable confidence signals—models cannot self-regulate effort or recognize when to defer to humans.
— Albabtain's mixed-methods quasi-experiment found AI design choices (transparency, override friction) and governance structures determine whether systems augment or substitute judgment; augmentation requires low-friction overrides and performance metrics rewarding contextual judgment, not AI compliance.
— Cox Communications achieved 7x first-year ROI on multi-agent sales decisioning across 15,000 employees; Kai's autonomous security triage eliminated 99.5% false positives across 2.5M findings, demonstrating operational maturity in bounded decision domains.
— Pragmatic cluster-randomized trial (9,600+ patients, 16 sites) found AI-assisted clinical decision support improved documentation quality but failed to reduce treatment failure (2.2% vs 2.0%, P=0.13), revealing decision-support ROI gap between capability and outcome impact.
— UK startup identified that 55% of professionals cite lack of structured human-AI workflows as adoption barrier; AI-adoption challenges are framework/coordination problems, not tool access—organizations need shared prompt libraries, training, and quality standards for sustained decision-support integration.
— By Harvard AI Institute director and HBS leaders: core bottleneck for AI transformation is not technology access but organizational ability to make decision-making processes explicit—articulating decision types, flows, criteria, trade-off rules, and escalation conditions required for consistent agent performance.
— PNAS Nexus study: GPT-4o and Claude 3.5 collapse on Stroop task as cognitive load increases (91%→1% accuracy), unable to suppress prepotent responses, indicating architectural limitation in executive control needed for reliable reasoning under conflicting information.
— Substantive analysis synthesizing Anthropic's 81,000-person survey (unreliability is #1 concern over capability), developer productivity paradox (180% more code but 30% more releases), and Rabanser reliability framework showing frontier models commit 6-12% safety violations on critical tasks.
— 1,816 professionals across legal, tax, audit, accounting report 91% experience AI value shortfall despite 74% regular use; 90% demand reasoning that can be explained and defended, identifying core decision-support framework requirement.
— Production deployment of AI decision-support system embedding company reasoning standards into leadership decisions; 100% adoption across 5 senior leaders with measurable behavioral change (reflex escalation becoming structured ownership-first decision-making).
— NIST AI Risk Management Framework 1.0 provides normative trustworthiness framework (valid/reliable, safe, secure, resilient, accountable, transparent, explainable, fair) grounding evaluation of all decision-support systems across accuracy, robustness, and accountability dimensions.
— Analysis of ~400,000 interactive Claude Code sessions reveals human-AI decision-making patterns: people make planning decisions (70%) while AI makes execution decisions (20%), demonstrating decision delegation frameworks and domain expertise amplifying agentic utility in practice.
— VA production deployment of clinical decision-support system with embedded governance framework (transparency, safety, explainability, accountability); repeatable, auditable process improving decision quality for healthcare providers with integrated LIME interpretability.
— Empirical research isolating format effects from reasoning capacity: structured output (JSON) degrades accuracy in models near capacity limits (Haiku 36.2pp drop, GPT-4o-mini 28.0pp), revealing fundamental tension between output structure and reasoning capability.
— CMU/Accenture empirically validated AI Adoption Maturity Model (8 dimensions, 5 levels) field-tested with Fortune 500 companies; addresses root cause of 95% zero-return adoption—mismatched expectations and poorly executed implementation, not technology.
— Peer-reviewed study challenges algorithm aversion stereotype; reveals reception type (dominant vs. oppositional) significantly influences advice-taking, with innovativeness and prior exposure as key trust determinants in knowledge-worker decision contexts.
— Developer deploys RecursiveMAS reasoning framework using extended thinking; read-after architecture achieves 65.5% accuracy at 2.7x cost with named use cases (legal review, medical triage, financial analysis), demonstrating practical multi-agent decision support.
— Vendor analysis (1000+ companies) identifies 95% pilot failure from organizational sequencing failures (top-down mandates, weak training, IT gatekeeping), proposing bottom-up adoption framework to reverse failure patterns in decision-support deployment.
— ICML 2026 research proves tool-integrated reasoning (86-94% accuracy) dramatically outperforms neural-only chain-of-thought (24-42%), establishing theoretical framework that hybrid decision-support systems exceed pure AI reasoning.
— ACL 2026 peer-reviewed study (1,440 adoption decisions) reveals confirmation bias drives 64.5% under-reliance when AI agrees with humans' initial incorrect answer, exposing core adoption barrier in human-AI decision collaboration.
— Prospective forecasting study (Michigan Ross) on 30 Kickstarter ventures shows frontier AI ranks 0.74 correlation vs. experts 0.04-0.45; human+AI hybrid reduces accuracy, challenging assumed value of human-in-the-loop decision support.
— Peer-reviewed study: AI agents ignored evidence 68% of tasks, made unsupported claims 53%, used contradictory evidence to revise output only 26%, revealing inability to incorporate experimental data—core failure mode in decision support reliability.
— Applies Expectancy Violation Theory to explain overconfident AI systems trigger defensive disengagement; cites Stanford study finding chatbots validated flawed reasoning 73%, explaining why system polish undermines human critical evaluation in decision-making.
— Meta-analysis of 5 RCTs (12,657 participants): AI clinical decision support produces small, marginal improvement in diagnostic accuracy; statistically significant but narrowly above zero; strongest in radiology, weakest in complex reasoning; signals overstated capability claims.
— Executive decision-support framework: persistent context (priorities, style, constraints), three core skills (Chief of Staff briefings, weekly accountability reviews, strategic sounding board); demonstrates personal effectiveness decision frameworks in production.
— Practitioner decision patterns for production reliability: Clarify-Then-Act, Plan-Then-Execute with bounded steps, Human-in-the-Loop gates for high-risk decisions, observability for reasoning chains; enterprise governance framework.
— ICML workshop paper: Institutional Alignment Readiness (IAR) framework shows decision-support systems fail at deployment, not performance; case studies document institutional barriers (approvals, oversight capacity, fiscal sustainability) blocking scaling despite technical viability.
— Survey of 650 leaders: 78% ran AI agent pilots, only 14% scaled to production; 90% of agent code fails EU AI Act compliance; governance failure (audit trails, escalation, rollback, ownership) is primary blocker, not performance; architectural prerequisite for production.
— SAP Sapphire 2026: 224 specialized Claude agents deployed across finance, supply chain, HR, procurement handling autonomous approvals and compliance; €100M investment; operates on €87T annual global commerce; largest production deployment of autonomous decision-making at enterprise scale.
— GA announcement: Claude as primary reasoning engine in SAP Business AI Platform; Joule agents execute decisions within existing controls, approvals, and compliance frameworks; hundreds of thousands enterprise customers across finance, HR, supply chain.
— TU Berlin study: psychological decision-making frameworks (Recognition-Primed Decision-Making, Data-Frame Theory) boost AI reasoning accuracy 13% → 30% in healthcare; structured human decision frameworks significantly improve AI decision quality.
— Production deployment: CLR-voyance clinical reasoning system live 6+ months at hospital drafting thousands of inpatient notes; 84.91% accuracy using POMDP framework with physician-validated outcome rubrics; demonstrates mature structured decision-support.
— Survey (72 leaders, 30+ industries) shows 96% maintain human-in-the-loop on consequential decisions; 61% cite workflow redesign as primary enabler; governance framework is load-bearing for production deployment, not optional overhead.
— 18-month empirical implementation study develops 6-module governance framework; reveals clinical AI systems remain confined to pilots not due to model limitations but institutional decision-making and governance capacity gaps.
— Analysis of AI overconfidence as design choice: RLHF incentivizes confidence without uncertainty; ECE 0.726 with 23% accuracy; standard calibration techniques reduce error 90% but not deployed—systemic design bias against decision-support reliability.
— Research shows 10-minute AI assistance impairs independent problem-solving through cognitive offloading; reduced retention and analytical skepticism; unintended consequence: outsourcing thinking erodes decision autonomy and agency.
— Empirical testing across 11 frontier models (67,221 records) reveals 8 collapse under adversarial pressure with 30.2pp accuracy drops; Anthropic Constitutional AI near-immune, indicating alignment-specific stability required for decision-support reliability.
— <20% of enterprise AI pilots reach production due to missing trust infrastructure: explainability, audit trails, governance, liability clarity. Core barrier is not capability but decision-framework accountability structures.
— Analysis of 5.5M real-world interactions shows top models fail ~9% overall, 14-16% on expert decision-making tasks (finance/law/medical); performance improvement flatlined despite compute scaling, exposing adoption barriers.
— Wharton RCT (1,300+ participants) shows AI-assisted decisions improve 25pp when correct but degrade 15pp when wrong; overconfidence persists even at 50% error rate, demonstrating cognitive surrender risk in decision frameworks.
— Stanford HAI's 2026 Index documents hallucination rates 22-94% and knowledge-belief reasoning failures where accuracy collapses 90%+ under false premises, revealing capability-accountability gap in professional decision contexts.
— FAccT 2026 mixed-methods study (incident database + practitioner survey) reveals AI system abandonment driven by organizational dynamics and resource constraints, not ethics concerns; directly applicable to decision-support deployment barriers.
— Conference reporting production deployments of bounded AI decision systems in insurance, pharma, financial crime with measured ROI (50%+ claims processing gains); evidence of transition from pilots to integrated decision workflows.
— Stanford Emerging Tech Review enumerates critical failure modes (hallucinations, overtrust, adversarial injection) and engineering defenses; notes most teams shipping decision-support features have implemented none, indicating significant maturity gap.
— Harvard research identifies relational complexity as core limitation: AI accuracy drops sharply when decisions require weighing multiple interacting factors simultaneously—directly constrains multi-factor analysis in healthcare and strategic decision-making.
— Stanford Index analysis documents scaling (88% organizational adoption) alongside reliability collapse: hallucination rates 22-94%, evaluation gaps between benchmarks and deployment—core constraints on decision-support reliability at scale.
— Increased AI trust correlates with reduced sense of agency (r=.288, p<.01) and increased indecisiveness—critical negative signal: reliance on AI decision systems degrades individual decision autonomy and confidence.
— Practitioner framework bridging speed-quality gap: proposes specific human judgment training for AI decision-making contexts. Addresses gap between productivity gains and actual decision quality improvement in AI-assisted reasoning.
— SPEC framework addresses presumptuousness (confident answers despite insufficient evidence) in legal decision-making. Achieves 89% accuracy vs. 15% baseline RAG on incomplete-info cases—positive signal for bounded reasoning frameworks.
— CFA Institute analysis: AI weakens epistemic foundations through cognitive delegation and 'knowledge-collapse equilibrium'; argues decision authority must remain anchored in evidence-based human inquiry—governance framework for responsible decision support.
— WalkMe study (3,750 professionals): 9% trust AI for complex decisions vs. 61% executive claims; 80% reject enterprise AI—quantifies massive gap between adoption narratives and actual decision-support reliance.
— Behavioral study (1,923 adults) shows passive reliance on AI reduces independent judgment confidence and sense of authorship; active engagement mitigates effect—critical negative signal that decision-support tool design fundamentally shapes reasoning capability erosion.
— 48,000 respondents across 47 countries identify trust as critical adoption barrier for AI decision-support. Documents gap between perceived benefits and actual system reliance—core adoption blocker.
— Documents 'Reliability Gap': capability scales 2-3x annually while reliability only 1.2-1.5x. Models solve graduate physics yet fail analog clocks (50.6% accuracy)—demonstrates maturity asymmetry that blocks reliable decision-support scaling.
— RCT (367 participants) shows ML-generated decision support reduces decisional conflict 3.77 points and increases satisfaction 7.38 points; behavioral outcomes confirm acceptance recommendation adoption (50.7% vs 24.2%), demonstrating measurable deployment value in medical decision-making.
— Google Cloud executive identifies governance and continuous evaluation as mandatory infrastructure for scaling AI agents from pilot to production; 95% failure rate stems from fragmented data and missing orchestration, not model performance.
— Empirical testing of 7 models (8B-235B) reveals working memory as scale-invariant bottleneck—all models collapse at 20-30 parallel branches regardless of size, exposing fundamental architectural constraints limiting reasoning systems.
— Peer-reviewed research (ACL 2026) identifies shallow lock-in and deep decay in reasoning models; proposes test-time StepFlow intervention improving accuracy without retraining, balancing documented failures with technical solutions.
— CMU research testing 14 leading LLMs (GPT-4, Claude 3, Gemini) reveals all fail simple logical contradiction detection, exposing benchmark illusion: curated-dataset performance masks lack of robust foundational reasoning.
— Critical analysis explaining LLMs predict tokens, not reason formally; Apple's GSM-Symbolic study shows single irrelevant clause causes up to 65% accuracy drop, demonstrating sophisticated pattern-matching rather than rule-based reasoning.
— Semantic variants of 677 GSM8K problems show 28.8%-45.1% answer-flip rates despite meaning preservation; entangled failures are only 5.2%-12.2% repairable, demonstrating fundamental brittleness in mathematical reasoning.
— Federal Court of Australia judgment (ASIC v Bekier) establishes liability for informal AI use in board decisions; identifies specific risks (hallucination, outsourced thinking, shadow AI, data sovereignty) and sets precedent for decision-support accountability.
— Authoritative multi-phase assessment framework (CCJ Task Force, 15 leaders) operationalizes AI adoption decisions with rigorous validation, demographic impact assessment, and mandatory human override—demonstrates decision-support governance in high-stakes deployment.
— Anthropic's May 2025 research finds Claude 3.7 Sonnet discloses hint usage only 25% of the time—chain-of-thought is post-hoc rationalization, not faithful reasoning log, critically undermining use of AI reasoning chains for verification.
— Real-world government deployments of AI decision-support systems (Japan tsunami impact modeling, Sweden autonomous defibrillator drones) demonstrate operational viability in high-stakes emergency-response contexts.
— Multi-expert RegTech consensus establishes decision-support governance framework: humans remain accountable; AI provides alerts/recommendations under human oversight; mature firms separate rule design, escalation, and model validation ownership.
— Analysis of reasoning model performance reveals three regimes: low-complexity tasks degrade 4.7-8.2pp, medium-complexity improve 9.1-12.3pp (adoption sweet spot), high-complexity collapse below 5%; 5x cost multiplier limits deployment beyond medium-complexity decisions.
— CRYSTAL benchmark with 6,372 questions reveals 19/20 models skip 50%+ of reasoning steps—achieving 58% accuracy while recovering only 48% of actual reasoning, proving models pattern-match rather than genuinely reason.
— Legal analysis establishes AI cannot relieve executives of decision review duty; requires mandatory governance documentation including user responsibility, AI tool version, prompts used, and evidence of context/source verification.
— Benchmark across 8 frontier models shows even Claude Opus 4.6 exhibits 6-16 percentage-point gaps between accuracy and consistency, revealing stochastic reasoning—key reliability issue for decision-support deployment.
— Stanford taxonomy from 500+ papers classifies reasoning failures on two axes: formal vs informal reasoning, and fundamental architectural limitations (unsolvable by scaling) versus application-specific shortcomings—framework for understanding decision-support constraints.
— Framework for agentic AI ROI emphasizing decision quality, speed, and autonomous multi-step reasoning; metrics shift from cost-savings to value-creation (revenue uplift, risk-avoided, cycle-time reduction) across interconnected decision workflows.
— Three documented legal failures: NZ courts ruled AI-hallucinated citations may amount to obstruction of justice; Deloitte refunded AUD 440K for AI-generated errors; regulatory finding that accountability for accuracy cannot be delegated to AI.
— GA platform providing orchestrated AI agents with decision-making automation; customer deployments at Johnson & Johnson and McCann Worldgroup with claimed 5x ROI and documented time-to-treatment improvements in healthcare operations.
— Production deployment of ARPIA platform orchestrating ML, GenAI, and Agentic AI for financial services collections pipeline achieving 13-minute raw-data-to-strategy activation time, demonstrating feasible multi-AI coordination in decision-support workflows.
— Summary of Stanford/Caltech research on systematic LLM reasoning failures: Reversal Curse, Robustness Fragility, and Working Memory Leaks—predictable error modes undermining reliance on AI reasoning trajectories for decision support reliability.
— Peer-reviewed environmental scan of commercially available AI-CDSS solutions reveals critical transparency gaps: vendor proprietary algorithms lack training data disclosure, knowledge base rigor unspecified, and privacy details incomplete—blocking reliable implementation evaluation.
— Thomson Reuters survey (1,500+ professionals, 27 countries) finds organization-wide AI adoption in professional services nearly doubled to 40% in 2026 from 22% in 2025; yet only 18% track ROI, and 40% report client confusion on AI use policies.
— Synthesis of 2,400+ enterprise AI initiatives shows 80.3% overall failure rate (33.8% abandoned, 28.4% deliver no value, 18.1% cannot justify costs); 95% GenAI pilot-to-production failure rate across sectors, signaling critical scaling barriers.
— CHI 2026 research on AI decision support in clinical sepsis treatment; contextual study with 25 physicians identified that reasoning cues should target high-discretion tasks, yet AI systems often fail to support decisions despite technical accuracy.
— PwC CEO survey (4,454 executives) found only 12% report AI delivered both cost and revenue benefits; Section survey shows employee-executive perception gap, exposing critical ROI and governance barriers to decision-support effectiveness.
— Industry trend analysis citing 62% of 625 enterprise AI professionals planning to evolve beyond automation to AI decision intelligence; identifies 74% faithfulness gap in LLM/RAG pipelines highlighting current system limitations.
— Analysis documenting 70-85% AI project failure rates due to data readiness and governance gaps; 42% of companies abandoned AI initiatives in 2025 (up from 17% in 2024), signaling persistent deployment barriers.
— Cloudera leadership predicts 2026 transition from experimentation to operationalization; emphasizes data as living knowledge system enabling agentic workflows with strong governance frameworks as non-negotiable for safe decision-making at scale.
— Hybrid AI meta-model evaluated on SEER/MIMIC datasets achieved 78% reduction in guideline-violation errors (18% to 6%) with clinician usability study showing 6.2/7 decision clarity rating, demonstrating feasible error reduction in clinical decision support.
— Peer-reviewed review in Journal of Bone & Joint Surgery showing AI's modest impact on patient care stems from data quality and EMR design flaws rather than algorithmic limitations, identifying concrete deployment barriers.
— Industry report showing 26-75% customer adoption of AI decisioning features across marketing platforms with measurable impacts on conversion rates and ROI, signaling operational deployment maturity.
— BMJ editorial warning that overreliance on GenAI risks eroding critical thinking, reinforcing biases, and causing deskilling in medical trainees, documenting persistent risks in decision-support adoption.
— News report on Indian judges warning against AI overreliance citing hallucinated citations and fabricated judgments, illustrating critical risks of AI decision-support in high-stakes legal domains.
— Ethics journal article examining justified use of black-box AI in medical decision-making, addressing validity and explainability concerns central to responsible deployment in high-stakes clinical reasoning.
— IBM internal deployment of agentic AI to 270,000 employees with estimated $4.5B productivity impact, demonstrating large-scale enterprise adoption of reasoning systems in operational decision workflows.
— Carnegie Mellon SEI released AI Robustness (AIR) open-source tool for federal agencies to detect bias and improve decision-making through causal discovery, with DoD testing and national security deployment.
— Research framework (ModelSelect) applying MCDM principles to AI model selection; validated with 50 real-world case studies showing high coverage and rationale alignment for structured decision support.
— MIT NANDA analysis of 150 interviews and 300 public deployments finding 95% of enterprise AI pilots fail to deliver value; vendor solutions succeed 67% vs internal builds 33%, exposing deployment execution gaps.
— YouGov survey of 10,000 respondents across 10 countries: 52% comfortable with personal AI for daily tasks, 64% for calendar/to-do, 39% for financial decisions; identifies trust factors (security, transparency, oversight).
— Peer-reviewed PLOS ONE study deriving empirical decision weights for AI DSS selection in prognostics; shows performance, effort, and transparency as key decision factors with expertise-dependent user preferences.
— HBR tip warning executives that AI forecasts create overconfidence through impressive detail and trend extrapolation, citing research showing executives using GenAI made worse predictions than without AI.
— Analysis citing HBR data showing only 26% of companies have working AI products and only 4% achieve significant returns, exposing critical ROI gap in scaling AI decision-support systems.
— Gartner prediction that 40%+ of agentic AI projects will be cancelled by 2027 due to rising costs and unclear value; research shows AI agents fail ~70% of office tasks, exposing deployment challenges.
— Industry report on AI-driven DSS in defense with case studies of Project Maven, UK autonomous warrior programme, and Israel's Iron Dome, demonstrating real-world deployment in high-stakes decision contexts.
— Delphi consensus study with 36 healthcare experts on implementing AI-based DSS for antibiotic prescribing, identifying critical barriers including trust, transparency, and organizational readiness challenges.
— Study of 223 dermatologists found AI support improved accuracy by only 1%, with low reliance (10%) and high self-reliance (86%), highlighting persistent barriers to effective clinical decision-support adoption.
— Research testing ChatGPT on 18 decision-making scenarios found AI made mistakes similar to humans in half the tests, showing biases like overconfidence and gambler's fallacy that undermine decision-support reliability.
— UK government production deployment of AI for TB screening achieved 90% accuracy and 85% reduction in manual review workload; real-time pollen monitoring system piloting, demonstrating concrete healthcare decision-support outcomes.
— UK government scrapped half-a-dozen AI pilots for welfare systems (A-cubed, Aigent) due to scalability, reliability, and testing challenges; explicit abandonment of decision-support deployments due to technical barriers.
— EU Policy Lab research on credit lending and recruitment decisions found human overseers equally likely to follow biased AI recommendations; human oversight alone insufficient to prevent discrimination in AI decision-making.
— Survey of 450 physicians across hospital types in China identifies adoption pathways for AI-CDSSs; tertiary hospitals exhibit 6 distinct routes while primary hospitals show 3, with utility perception and readiness as key factors.
— Critical assessment of AI reasoning model limitations: data bias, lack of common sense, and transparency gaps preventing effective decision support in high-stakes domains like healthcare and finance.
— Peer-reviewed critique of AI ethics consultations, identifying fundamental barriers: lack of transparency ('black box') and unresolved questions about AI authority in complex ethical reasoning decisions.
— Qualitative study of healthcare professionals revealing significant ethical concerns in AI-CDSS: bias, transparency gaps, and accountability deficits in automated resource allocation decisions.
— Technical preprint revealing state-of-the-art reasoning models (LLaMA, Qwen) exhibit overthinking and discard correct solutions, fundamentally undermining reliance on AI reasoning trajectories for decision support.
— Survey of 90 legal professionals finds 43% observe bias in AI tools and 37% fear unreliability; AI is widely adopted yet critical concerns about bias and accuracy persist as barriers to trustworthy deployment.
— Quasi-experimental study documenting physician overreliance on AI diagnostic systems; proposes trust calibration to match AI reliability, revealing persistent pitfalls in clinical decision-support deployment.
— AHA survey of 100+ physicians shows 76% use LLMs for clinical decisions despite reliability concerns; AI chatbots score high in reasoning benchmarks but exhibit incorrect reasoning, not yet ready for autonomous deployment.
— Scoping review protocol on barriers to AI-based clinical DSS implementation, identifying acceptance and feasibility challenges that participatory design methods seek to address in healthcare contexts.
— Peer-reviewed analysis from UC Irvine identifying three key barriers to effective human-AI decision-making: complementarity conditions, human mental models of AI, and design choice impacts on cognitive load.
— ICRC critical assessment documenting permanent risks in AI decision-support systems for high-stakes decisions: biases, hallucinations, brittleness, and inability to ensure compliance with legal requirements.
— Strategic analysis from Berkeley/IBM researchers on justifying AI governance investments, noting companies with Gen AI guardrails may be 27% more likely to achieve higher revenue despite ethical and compliance costs.
— Gartner analyst forecast predicting 30% of generative AI projects will be abandoned post-proof-of-concept by end of 2025 due to poor data quality, inadequate risk controls, and unclear business value.
— Qualitative study of 24 Dutch care professionals on prerequisites for responsible AI-DSS deployment, identifying seven critical requirements including bias mitigation, human-centric learning, and incremental trust-building.
— DLA Piper survey of 600 executives found 48% of AI projects paused or rolled back due to privacy, regulatory, and integration challenges, revealing high deployment failure rates in practice.
— RCT in Wisconsin courts found AI recommendations failed to improve judicial bail decisions; judges rejected AI advice 30%+ of time due to inferior accuracy, demonstrating limitations in AI decision-support.
— Peer-reviewed case study of AI decision-support deployed in Dutch court for traffic appeals, showing real-world impact on legal decisions with tensions in adoption and discretionary authority.
— Experimental study with 1,403 participants shows incorrect AI advice harms decisions due to overreliance; explainability does not reduce this negative effect, revealing critical human-AI interaction risks.
— Large-scale empirical study with 528 participants shows AI bias in hiring leads to up to 90% alignment with biased recommendations, demonstrating AI's capacity to reduce human agency in decision-making.
— Industry analysis from CFA Institute on AI bias in investment decision-support, proposing frameworks for responsible AI with focus on fairness, accountability, and human oversight in financial decisions.
— Nearly nine in ten organizations use AI/ML for autonomous decision-making despite 6% average revenue loss from model failures; reveals adoption breadth but persistent trust and data quality challenges.
— Study of AI Scientist systems shows Claude 3.5 scores only 1.8% on PaperBench complex engineering tasks; identifies implementation execution as the 'fundamental bottleneck' in AI reasoning systems.
— Analysis of AI bias and failures in educational decision systems; shows how AI-powered decision support perpetuates historical inequities despite claims of objectivity.
— Apple research identifies critical reasoning failures in LLMs; proposes GSM-Symbolic benchmark to measure logical reasoning—exposing fundamental limitations in AI decision support capability.
— AI scores only 30% on ARC benchmark for novel reasoning tasks, far below human performance; demonstrates AI cannot yet handle decision-making requiring genuine reasoning on unfamiliar problem types.
— Comprehensive 160-page survey (750+ references) synthesizing state of AI reasoning capabilities in foundation models—the technical substrate for decision support systems.
— Clinicians' primary adoption barriers for AI-CDSSs are accuracy concerns and legal liability; trust framework analysis shows unresolved accountability gap limiting real-world deployment.
— Despite renewed interest in AI-CDS, lack of empirical evidence on effectiveness; warns of sub-optimal implementation risks and calls for rigorous evaluation frameworks.
— Validated PAAI questionnaire assessing AI-DSSs from human-centered perspective with N=223 and N=471 studies, demonstrating design-psychology link to psychological load reduction and performance.
— CMU FAccT award-winning research on threats to validity and reliability in human-AI decision-making; identifies methods to address core reliability challenges in this domain.
— AI decision support risks include dynamic environment brittleness, over-reliance, bias, and governance gaps; warns that 'complete trust in AI' is the most dangerous moral hazard.
— Observational analysis showing unhelpful AI responses cause users to reduce engagement and diversity over time, revealing adoption barriers in personal decision support.
— Stanford experimental study showing simpler explanations and higher decision stakes reduce overreliance on AI, with quantified effect sizes on decision quality.
— Social media sentiment analysis revealing users' primary concerns are risk and accountability in AI decision-making systems, identifying key adoption barriers.
— Empirical dataset from MIT/CMU/Berkeley study of 100 participants using AI in chess puzzles, capturing human-AI collaboration dynamics and confidence calibration in decision tasks.
— Deloitte analysis citing Gartner prediction that ~85% of AI projects fail to move from prototype to production, highlighting critical adoption barriers in enterprise AI decision-making.
— Introduces TempReason benchmark showing LLMs perform poorly on temporal reasoning and exhibit bias toward contemporary years (2000-2020), exposing decision-support limitations.
— John Deere achieved precision agriculture transformation via AI, but Zillow lost $500M when AI models failed to adapt to housing market changes, demonstrating both promise and severe deployment risks.
— Legal analysis proposing framework for human-in-the-loop decision-making, challenging assumption that human judgment always overrides AI, with healthcare and hiring examples.
— Describes hybrid human-AI decision support in transport operations; identifies critical barriers including lack of formalizability of human decision processes and system acceptance challenges.
— Theoretical paper asserting that AI mimicking human intelligence must be fallible and require 'I don't know' capability; identifies fundamental limitations in current AI reasoning systems.