{
  "slug": "decision-support-and-reasoning-frameworks",
  "name": "Decision support & reasoning frameworks",
  "tier": "bleeding-edge",
  "trend": "steady",
  "blockerType": null,
  "tools": [],
  "evidence": [
    {
      "title": "Frontier Models Fail at Decision-Fork Reasoning: Taste-Bench Benchmark",
      "url": "https://hyper.ai/en/papers/2609.25804",
      "date": "2026-09-23",
      "type": "research-paper",
      "added": "2026-09-27",
      "superseded_by": null,
      "window": null,
      "explanation": "Benchmark showing frontier reasoning models answer only 59.7% of 502 decision-fork questions correctly, with larger reasoning budgets providing no improvement—exposing architectural constraint on AI-delegated decision-making."
    },
    {
      "title": "AI Decision Matrix Framework and One New Zealand Deployment: Allocating Decision Rights",
      "url": "https://mitsloan.mit.edu/ideas-made-to-matter/a-framework-determining-when-ai-can-make-decisions",
      "date": "2026-09-22",
      "type": "news-coverage",
      "added": "2026-09-27",
      "superseded_by": null,
      "window": null,
      "explanation": "MIT CISR framework allocating decision rights between humans and agents on ambiguity × risk matrix, exemplified by One New Zealand's 50+ agents reducing audit-planning time 60% and risk-matrix preparation from 2 days to under half a day."
    },
    {
      "title": "Technical Debt in AI Decision-Support Systems: 27 Ungoverned Agent Variants and Untraceable Errors",
      "url": "https://www.techtarget.com/it-strategy/opinion/The-quiet-rise-of-AI-technical-debt",
      "date": "2026-09-21",
      "type": "opinion",
      "added": "2026-09-27",
      "superseded_by": null,
      "window": null,
      "explanation": "Practitioner case documenting how a procurement contract-summarization agent spawned 27 variants in a year without unified ownership of prompts, data sources or evaluation logic, creating untraceable hallucinations and blocking scaling."
    },
    {
      "title": "SOCPAC Intelligence Failure: AI Fabricated Nuclear Weapons Claim, Nearly Triggered Ship Boarding",
      "url": "https://www.techtimes.com/articles/327796/20260920/us-military-almost-boarded-chinese-ship-over-ai-hallucinated-nuclear-claim.htm",
      "date": "2026-09-20",
      "type": "news-coverage",
      "added": "2026-09-27",
      "superseded_by": null,
      "window": null,
      "explanation": "High-stakes failure in which US military analyst's AI synthesized intelligence to fabricate a nuclear-weapons claim, formatted as authoritative report, came within minutes of triggering a ship boarding before error was caught."
    },
    {
      "title": "Accountability Framework for AI-Generated Decision Briefings",
      "url": "https://www.techscience.com/cmc/v89n2/68847/html",
      "date": "2026-09-15",
      "type": "research-paper",
      "added": "2026-09-27",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed five-dimension accountability framework (citation-to-claim entailment, causal-language discipline, uncertainty preservation, action appropriateness, human accountability) establishing that factually correct briefings can still be unsafe without explicit governance."
    },
    {
      "title": "Xylem: 15,000-Employee Decision-Support Deployment at Enterprise Scale",
      "url": "https://www.ey.com/en_gl/insights/industrial-products/how-xylem-turned-ai-ambition-into-enterprise-execution",
      "date": "2026-09-15",
      "type": "case-study",
      "added": "2026-09-27",
      "superseded_by": null,
      "window": null,
      "explanation": "Named enterprise deployment onboarding 15,000 employees with 4,000 daily active users, estimated $70M revenue opportunity and $25M savings potential, demonstrating decision-support systems reaching organizational scale."
    },
    {
      "title": "Agency Decay: Evidence That AI Overreliance Erodes Human Judgment and Critical Thinking",
      "url": "https://knowledge.wharton.upenn.edu/article/is-overreliance-on-ai-causing-agency-decay/",
      "date": "2026-09-15",
      "type": "opinion",
      "added": "2026-09-27",
      "superseded_by": null,
      "window": null,
      "explanation": "Wharton analysis documenting judgment erosion from AI reliance, with Microsoft study showing higher AI confidence correlates with lower critical thinking, and BetterUp data quantifying 'workslop' costs ($186/employee/month)."
    },
    {
      "title": "Human Override Paradox: How Human Authority Erases AI's Accuracy Advantage",
      "url": "https://www.hec.edu/en/dare/tech-ai/giving-humans-final-say-can-make-ai-decisions-worse",
      "date": "2026-09-14",
      "type": "opinion",
      "added": "2026-09-27",
      "superseded_by": null,
      "window": null,
      "explanation": "HEC Paris analysis of Sibony & Hazan framework showing human override of superior AI (physician case: AI 90% accurate, human alone 74%, human override 76%) destroys value, establishing fundamental design tension in human-AI decision-making."
    },
    {
      "title": "How AI adoption erodes employee judgment and critical thinking skills",
      "url": "https://finance.yahoo.com/technology/ai/articles/ai-adoption-erodes-employee-judgment-090000724.html",
      "date": "2026-09-08",
      "type": "adoption-metric",
      "added": "2026-09-13",
      "superseded_by": null,
      "window": null,
      "explanation": "BCG survey of 70 C-suite leaders: 50% observing judgment/problem-solving skill decay with 3-5 year business impact expected; 90% report employees stopped checking AI work; Shell's protective redesign demonstrates remediation through deliberate reasoning preservation."
    },
    {
      "title": "Care and consideration | Ada Lovelace Institute",
      "url": "https://www.adalovelaceinstitute.org/report/care-and-consideration/",
      "date": "2026-09-08",
      "type": "industry-report",
      "added": "2026-09-13",
      "superseded_by": null,
      "window": null,
      "explanation": "Ada Lovelace Institute and NICE partnership: independent research on public-sector AI adoption barriers identifies decision-support risks (reliance, accountability, bias) and proposes sociotechnical governance framework for iterative evaluation."
    },
    {
      "title": "The More Powerful AI Becomes, the More Valuable Your Native Human Judgment Ability Is",
      "url": "https://eu.36kr.com/en/p/3972590737043721",
      "date": "2026-09-07",
      "type": "case-study",
      "added": "2026-09-13",
      "superseded_by": null,
      "window": null,
      "explanation": "Harvard field experiment with 228 senior reviewers found automation bias overwhelming: reviewers rejected expert-approved projects following AI recommendations, and explanations paradoxically worsened bias; proposes CAST framework to preserve human judgment."
    },
    {
      "title": "Enterprise AI Deployment Failures and Outcomes in 2026",
      "url": "https://intuitionlabs.ai/articles/enterprise-ai-deployment-failures-and-outcomes-2026",
      "date": "2026-09-05",
      "type": "industry-report",
      "added": "2026-09-13",
      "superseded_by": null,
      "window": null,
      "explanation": "Meta-analysis of MIT, RAND, S&P Global, and Gartner studies: 95% of AI pilots report zero ROI, 42% abandon before production; root cause is organizational governance and decision-framework gaps, not model capability."
    },
    {
      "title": "Conversational AI for Construction Data, Built on Claude and Bedrock",
      "url": "https://sombrainc.com/case-studies/conversational-ai-construction-data-claude-bedrock",
      "date": "2026-09-04",
      "type": "case-study",
      "added": "2026-09-13",
      "superseded_by": null,
      "window": null,
      "explanation": "Production deployment across tens of thousands of users handling 36K+ historical queries; multi-agent decision-support system with full audit trail, context management, and adaptive output selection addressing enterprise trust and accessibility."
    },
    {
      "title": "AI rollbacks: 18 deployments paused or reversed",
      "url": "https://aiweekly.co/ai-use-cases/rollbacks",
      "date": "2026-09-03",
      "type": "adoption-metric",
      "added": "2026-09-13",
      "superseded_by": null,
      "window": null,
      "explanation": "Comprehensive rollback ledger documents 22 halted AI deployments with named organizations and specific governance failure modes (reward-hacking, sandbox escape, privilege escalation, moderation collapse) demonstrating autonomous decision-support unreliability."
    },
    {
      "title": "Limitations of AI in Legal Software | TTMS",
      "url": "https://ttms.com/limitations-of-ai-in-legal-software-risks-of-incorrect-advice-defective-court-filings-and-missed-deadlines/",
      "date": "2026-09-03",
      "type": "case-study",
      "added": "2026-09-13",
      "superseded_by": null,
      "window": null,
      "explanation": "Poland Supreme Administrative Court case I FZ 104/26 documents AI hallucination in legal decision-support: counsel cited genuine case numbers but attributed fabricated propositions, exposing critical reasoning failure mode and professional liability in high-stakes decisions."
    },
    {
      "title": "Can Enterprise AI Assistants Be Trusted Under Pressure?",
      "url": "https://www.alphaxiv.org/abs/2609.pact-enterprise-ai-compliance-testing",
      "date": "2026-09-03",
      "type": "research-paper",
      "added": "2026-09-13",
      "superseded_by": null,
      "window": null,
      "explanation": "PACT benchmark of 22 models across 12 regulated domains found even top models violate compliance rules 6-10% at baseline, escalating 65% under realistic workplace pressure, identifying architectural compliance-failure mode in high-stakes decision-support."
    },
    {
      "title": "CORAL: An LLM-Native Harness for Production Recommender Systems",
      "url": "https://www.alphaxiv.org/abs/2609.02730",
      "date": "2026-09-02",
      "type": "research-paper",
      "added": "2026-09-13",
      "superseded_by": null,
      "window": null,
      "explanation": "Meta's CORAL framework demonstrates closed-loop decision-optimization at production scale across two billion-user platforms, with A/B-tested improvements in engagement and cost efficiency via constrained reasoning over persistent memory."
    },
    {
      "title": "The AI Pilot Paradox: Why Do 87% of Companies Delay AI Deployments?",
      "url": "https://www.avepoint.com/blog/strategy-blog/why-ai-pilots-fail",
      "date": "2026-09-01",
      "type": "adoption-metric",
      "added": "2026-09-13",
      "superseded_by": null,
      "window": null,
      "explanation": "AvePoint survey of 3,235 leaders: 87% delayed AI deployments due to data governance risks (not model quality), 40.7% canceled GenAI rollouts in 2026 (up 31.7% in 2025), exposing governance and ownership failures as primary scalability blockers."
    },
    {
      "title": "Rollback is a measured pattern, not an anecdote | Research",
      "url": "https://www.therevenueaireport.com/research/rollback",
      "date": "2026-08-31",
      "type": "adoption-metric",
      "added": "2026-09-13",
      "superseded_by": null,
      "window": null,
      "explanation": "Revenue AI Report survey of 2,527 enterprises (Jan-Feb 2026): 74% rolled back/shut down deployed AI agents due to governance failure; paradoxically, 81% with 'mature guardrails' still rolled back, indicating governance maturity lag."
    },
    {
      "title": "Salesforce puts 37 prebuilt sales skills inside Claude for pilot customers",
      "url": "https://ppc.land/salesforce-puts-37-prebuilt-sales-skills-inside-claude-for-pilot-customers/",
      "date": "2026-08-30",
      "type": "product-ga",
      "added": "2026-09-13",
      "superseded_by": null,
      "window": null,
      "explanation": "Salesforce/Anthropic partnership operationalizes 37 prebuilt sales decision-support skills (deal health, pipeline review, meeting prep) with governance model routing actions through Salesforce business rules, demonstrating enterprise decision-support system maturity."
    },
    {
      "title": "Sacked ALDI worker's use of AI as 'quasi-legal advisor' condemned",
      "url": "https://www.abc.net.au/news/2026-08-29/fair-work-commission-condemns-ai-legal-advice/107089766",
      "date": "2026-08-29",
      "type": "case-study",
      "added": "2026-08-30",
      "superseded_by": null,
      "window": null,
      "explanation": "Fair Work Commission tribunal documents AI decision-support failure; 40% of cases involved AI (adoption metric); institutional response mandates October 2026 disclosure requirement—governance driving adoption readiness."
    },
    {
      "title": "Drafting vs. Deciding: Where to Draw the Line with AI in Your Business",
      "url": "https://www.businessgrowthtalks.com/blog/drafting-vs-deciding-where-to-draw-the-line-with-ai-in-your-business/",
      "date": "2026-08-27",
      "type": "opinion",
      "added": "2026-08-30",
      "superseded_by": null,
      "window": null,
      "explanation": "Framework distinguishing AI-powered drafting from human-led deciding; identifies three irreducible human elements (empathy, accountability, moral responsibility) establishing decision-support boundaries."
    },
    {
      "title": "discernment-nudge - Claude Skill",
      "url": "https://www.aimcp.info/de/skills/d90e6fbb-f00b-440d-8aae-8fc0d6c5ec84",
      "date": "2026-08-26",
      "type": "product-ga",
      "added": "2026-08-30",
      "superseded_by": null,
      "window": null,
      "explanation": "General availability Claude Skill implementing structured reflection framework; vendor investment in decision-support reasoning architecture signals market maturity and recognition of governance-as-feature design requirement."
    },
    {
      "title": "AI financial advice can improve decisions or steer investors wrong, experiment finds",
      "url": "https://phys.org/news/2026-08-ai-financial-advice-decisions-investors.html",
      "date": "2026-08-25",
      "type": "research-paper",
      "added": "2026-08-30",
      "superseded_by": null,
      "window": null,
      "explanation": "University of Bayreuth study (n=3,700): AI recommendations amplified correct decisions (64%→87%) but crashed accuracy below baseline when misaligned (34%); conflict-of-interest disclosure offered minimal protection."
    },
    {
      "title": "Building AI Agents That Actually Work: A Deliberate Framework for Success",
      "url": "https://www.thundersf.com/blog/building-ai-agents-that-actually-work-a-deliberate-framework-for-success",
      "date": "2026-08-20",
      "type": "case-study",
      "added": "2026-08-30",
      "superseded_by": null,
      "window": null,
      "explanation": "Salesforce consulting framework documents staged agent maturity (Diagnose-Suggest-Act-Orchestrate) with human approval gates; named deployment achieved 50%+ case deflection after structured framework application."
    },
    {
      "title": "AI and human judgment - how they actually work together in personal finance",
      "url": "https://moneypatrol.ai/blog/ai-and-human-judgment",
      "date": "2026-08-20",
      "type": "case-study",
      "added": "2026-08-30",
      "superseded_by": null,
      "window": null,
      "explanation": "Production personal decision-support framework dividing pattern-work (AI) from value-work (human); safety architecture uses read-only Plaid access; documents failure modes and interrogation protocol for personal effectiveness."
    },
    {
      "title": "Potential Regression in Gemini 3.6, 3.1 pro, Mid-July 2026",
      "url": "https://discuss.ai.google.dev/t/potentital-regression-in-gemini-3-6-3-1-pro-mid-july-2026-ongoing-public-ga-web-api-ai-studio-pro-paid-tiers-effected/178084",
      "date": "2026-08-13",
      "type": "opinion",
      "added": "2026-08-16",
      "superseded_by": null,
      "window": null,
      "explanation": "30+ reproductions of constraint-preservation failure in Gemini: model violates explicit logical rules despite acknowledging them—fundamental reasoning architecture gap blocking governed decision-support deployment."
    },
    {
      "title": "About a Third of Workers Who Used AI in the Last Week Said They Completed Tasks One to Two Hours Faster",
      "url": "https://www.census.gov/library/stories/2026/08/ai-use-at-work.html",
      "date": "2026-08-11",
      "type": "adoption-metric",
      "added": "2026-08-16",
      "superseded_by": null,
      "window": null,
      "explanation": "US Census Bureau survey of 55% worker AI adoption with time-savings metrics; 31% report 1-2 hour savings, providing independent government baseline for real-world decision-support usage patterns."
    },
    {
      "title": "AI models make choices in part based on the order in which options are presented",
      "url": "https://techxplore.com/news/2026-08-ai-choices-based-options.html",
      "date": "2026-08-11",
      "type": "research-paper",
      "added": "2026-08-16",
      "superseded_by": null,
      "window": null,
      "explanation": "Testing across 9 LLMs reveals quality-dependent position bias and order-reversal can flip underlying preferences; fundamental brittleness in AI reasoning for decision-making across models."
    },
    {
      "title": "Confronting the Barriers to AI Diffusion in the U.S. Military",
      "url": "https://carnegieendowment.org/research/2026/08/confronting-the-barriers-to-ai-diffusion-in-the-us-military",
      "date": "2026-08-10",
      "type": "industry-report",
      "added": "2026-08-16",
      "superseded_by": null,
      "window": null,
      "explanation": "Maven handles 5,000+ targeting decisions per day with proven operational success; yet documents 8 institutional barriers limiting wider adoption—demonstrates deployment feasibility constrained by organizational factors."
    },
    {
      "title": "Is AI making employees less able? - IT-Online",
      "url": "https://it-online.co.za/2026/08/07/is-ai-making-employees-less-able/",
      "date": "2026-08-07",
      "type": "adoption-metric",
      "added": "2026-08-16",
      "superseded_by": null,
      "window": null,
      "explanation": "GoTo survey of 2,500 employees: 70% admit using AI for high-stakes decisions; 50% report overreliance; 39% believe overreliance erodes skills—documents widespread adoption with documented skill degradation."
    },
    {
      "title": "AI won't enhance physician autonomy. It will further diminish it",
      "url": "https://www.statnews.com/2026/08/07/medical-ai-doctors-autonomy/",
      "date": "2026-08-07",
      "type": "opinion",
      "added": "2026-08-16",
      "superseded_by": null,
      "window": null,
      "explanation": "81% of physicians use AI professionally; FDA narrowed oversight definition while shifting liability to clinicians; absence of human-centered governance design enables autonomy erosion in deployed decision systems."
    },
    {
      "title": "Beyond AI Scribes: Why Ambient Clinical Intelligence Is Health IT's Greatest Governance Test",
      "url": "https://hitconsultant.net/2026/08/06/akhila-akula-ambient-clinical-intelligence-health-it-test/",
      "date": "2026-08-06",
      "type": "adoption-metric",
      "added": "2026-08-16",
      "superseded_by": null,
      "window": null,
      "explanation": "Real deployment: 1,744 U.S. hospitals running ambient AI for clinical decisions; burnout dropped 51.9% to 38.8%; but only 10% have formal AI oversight boards—production adoption without governance maturity."
    },
    {
      "title": "Five million Brits harmed by wrong AI advice from lost cash to health scares",
      "url": "https://www.uswitch.com/media-centre/2026/08/five-million-brits-harmed-by-wrong-ai-advice-from-lost-cash-to-health-scares/",
      "date": "2026-08-05",
      "type": "adoption-metric",
      "added": "2026-08-16",
      "superseded_by": null,
      "window": null,
      "explanation": "Survey of UK adults: 5M+ self-reported harms from incorrect AI advice; 43% use chatbots for personal decisions; 18% act without second opinion—widespread production use with documented negative outcomes."
    },
    {
      "title": "Medical AI benefits vary by expertise, with nonexperts more easily misled",
      "url": "https://medicalxpress.com/news/2026-08-medical-ai-benefits-vary-expertise.html",
      "date": "2026-08-04",
      "type": "research-paper",
      "added": "2026-08-16",
      "superseded_by": null,
      "window": null,
      "explanation": "Nature Medicine study: explainability helps experts but creates automation bias in novices; one-size-fits-all explanations mislead users with less domain knowledge—interface design determines decision-support outcomes."
    },
    {
      "title": "Adaptive decision support can fight overreliance on AI - Tech Xplore",
      "url": "https://techxplore.com/news/2026-08-decision-overreliance-ai.html",
      "date": "2026-08-03",
      "type": "research-paper",
      "added": "2026-08-16",
      "superseded_by": null,
      "window": null,
      "explanation": "Harvard/MIT study (1,280 participants) shows reinforcement-learning adaptive decision-support improves accuracy and reduces overreliance; demonstrates design choices can enable human-AI complementarity."
    },
    {
      "title": "AI Agents 2026: New Benchmark Tests Real-World Decision Making",
      "url": "https://vfuturemedia.com/ai/ai-agents-2026-real-world-decision-benchmark/",
      "date": "2026-07-27",
      "type": "case-study",
      "added": "2026-08-02",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical evaluation of frontier reasoning models on real-world decision-making under uncertainty; none beat market baseline, revealing poor calibration in high-stakes forecasting."
    },
    {
      "title": "The Faithfulness Gap: Explainability in Medical AI Agents",
      "url": "https://heckelai.com/the-faithfulness-gap-explainability-in-medical-ai-agents/",
      "date": "2026-07-27",
      "type": "opinion",
      "added": "2026-08-02",
      "superseded_by": null,
      "window": null,
      "explanation": "Technical synthesis showing AI reasoning explanations (chain-of-thought) systematically misrepresent actual model computation, creating compliance and safety gaps in decision-support systems."
    },
    {
      "title": "When Employees Are Held Accountable for AI-Generated Decisions",
      "url": "https://hbr.org/2026/07/when-employees-are-held-accountable-for-ai-generated-decisions",
      "date": "2026-07-22",
      "type": "case-study",
      "added": "2026-08-02",
      "superseded_by": null,
      "window": null,
      "explanation": "Multi-year field study documents critical framework failure: frontline workers defend AI decisions they neither created nor understand—directly addresses the accountability-understanding gap in deployed decision support."
    },
    {
      "title": "Why 42% of companies abandoned their AI initiatives (and the hire that predicts survival)",
      "url": "https://galileosearch.com.au/insights/why-42-percent-of-companies-abandon-ai-initiatives",
      "date": "2026-07-20",
      "type": "adoption-metric",
      "added": "2026-08-02",
      "superseded_by": null,
      "window": null,
      "explanation": "Multiple independent sources (S&P Global, Gartner, MIT, BCG, IBM) document organizational barriers to AI decision-system adoption; identifies resource allocation and accountability as binding constraints, not technical capability."
    },
    {
      "title": "AI Advice Slashes Accuracy and Boosts Confidence: New Study",
      "url": "https://aitoolly.com/ai-news/article/2026-07-20-ai-advice-reduces-human-accuracy-threefold-while-doubling-confidence-levels-research-finds",
      "date": "2026-07-20",
      "type": "research-paper",
      "added": "2026-08-02",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical study showing accuracy fell 27%→9%, confidence rose 30%→76%, judgment suspension collapsed 44%→3%; directly addresses core reliability and validation concerns in decision-support frameworks."
    },
    {
      "title": "Why AI Adoption Depends on Management | Capital, Incentives, and Execution",
      "url": "https://fdiostanzo.com/posts/2026/20260719_ai_adoption_is_a_management",
      "date": "2026-07-19",
      "type": "opinion",
      "added": "2026-08-02",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed field experiment with 758 real BCG consultants: AI decision-support delivered +12.2% productivity and +40% quality for tasks within capability frontier, but 19pp reduced correctness outside frontier—directly measuring when reasoning frameworks work and fail."
    },
    {
      "title": "Your AI Agent Is Not Your Tool. It Is Your Employee Now.",
      "url": "https://lawandkoffee.substack.com/p/your-ai-agent-is-not-your-tool-it",
      "date": "2026-07-15",
      "type": "opinion",
      "added": "2026-07-19",
      "superseded_by": null,
      "window": null,
      "explanation": "Legal framework establishing AI decision-support governance requirement: mandatory documentation of reasoning review, source verification, and user responsibility—decision accountability cannot be delegated to AI systems."
    },
    {
      "title": "Information-seeking failures of large language models in agentic clinical reasoning",
      "url": "https://arxiv.org/abs/2607.10275v1",
      "date": "2026-07-11",
      "type": "research-paper",
      "added": "2026-07-19",
      "superseded_by": null,
      "window": null,
      "explanation": "32 frontier models tested in clinical reasoning show information utilization collapses 57%→26% under uncertainty, revealing core limitation: high reasoning capability does not ensure good decision outcomes without proper information-seeking behavior."
    },
    {
      "title": "Council of High Intelligence GitHub: Make Claude Code Challenge Its Own Answers",
      "url": "https://www.youtube.com/watch?v=UIQkiAZLLdY",
      "date": "2026-07-10",
      "type": "significant-repo",
      "added": "2026-07-19",
      "superseded_by": null,
      "window": null,
      "explanation": "Open-source Claude Code plugin implementing multi-agent orchestration through structured debate patterns; replaces single-model reasoning with disagreement and evaluation criteria to surface blind spots in decision-making."
    },
    {
      "title": "ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning",
      "url": "https://aclanthology.org/2026.findings-acl.1456/",
      "date": "2026-07-08",
      "type": "research-paper",
      "added": "2026-07-19",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed ACL study shows frontier reasoning models achieve <25% instruction-following compliance during reasoning traces—fundamental control gap limiting reliable structured decision-support deployment in governed environments."
    },
    {
      "title": "The Board Decision Gap: AI Transforms the Deal, But Boards Still Govern the Outcome",
      "url": "https://finance.yahoo.com/technology/ai/articles/board-decision-gap-ai-transforms-070000001.html",
      "date": "2026-07-08",
      "type": "adoption-metric",
      "added": "2026-07-19",
      "superseded_by": null,
      "window": null,
      "explanation": "1,000 dealmakers across 27 countries: 62% believe human-only decisions indefensible, but only 22% delegate final decisions to AI—establishes governance boundaries and shows decision authority requires human accountability for consequential choices."
    },
    {
      "title": "A Guide for CTOs to Get Started with Enterprise AI - WWT",
      "url": "https://www.wwt.com/wwt-research/cto-guide-to-ai",
      "date": "2026-07-08",
      "type": "industry-report",
      "added": "2026-07-19",
      "superseded_by": null,
      "window": null,
      "explanation": "Synthesis of Stanford analysis of 51 successful enterprise AI deployments identifies executive sponsorship paired with clear measurable business objectives defined before model work as the most consistent adoption accelerator."
    },
    {
      "title": "tjboudreaux/cc-thinking-skills | DeepWiki",
      "url": "https://deepwiki.com/tjboudreaux/cc-thinking-skills",
      "date": "2026-07-08",
      "type": "significant-repo",
      "added": "2026-07-19",
      "superseded_by": null,
      "window": null,
      "explanation": "39 vetted cognitive and reasoning frameworks (First Principles, Bayesian Updating, Theory of Constraints) packaged with rigorous evaluation pipeline; production deployment via Claude Code marketplace—operationalizes structured decision frameworks at scale."
    },
    {
      "title": "Randomized Trial Finds AI Can Improve Physicians' Clinical Reasoning, but Human–AI Design Determines Outcomes",
      "url": "https://www.policyedge.in/p/randomized-trial-finds-ai-can-improve-physicians-clinical-reasoning-but-humanai-design-determines-outcomes",
      "date": "2026-07-07",
      "type": "case-study",
      "added": "2026-07-19",
      "superseded_by": null,
      "window": null,
      "explanation": "RCT of 249 physicians across three countries shows AI improves reasoning (Kenya +18%) when integrated with proper human-AI design; demonstrates both deployment potential and critical role of design in clinical decision-support effectiveness."
    },
    {
      "title": "Forbes: The AI Validation Gap — $2.5T Blind Spot in Enterprise AI",
      "url": "https://www.forbes.com/councils/forbestechcouncil/2026/06/30/the-ai-validation-gap-the-25-trillion-blind-spot-in-enterprise-ai/",
      "date": "2026-06-30",
      "type": "opinion",
      "added": "2026-07-05",
      "superseded_by": null,
      "window": null,
      "explanation": "NBER study of ~6,000 executives: ~90% report no discernible impact on productivity from AI; PwC CEO survey shows 56% experienced no revenue/cost improvement with only 12% reporting significant benefits, documenting persistent gap between deployment and measurable decision-support outcomes."
    },
    {
      "title": "California Government Partnership: Institutional Decision-Support Adoption",
      "url": "https://www.gov.ca.gov/2026/06/29/governor-newsom-announces-a-first-of-its-kind-partnership-providing-anthropic-tools-to-state-agencies-and-improving-services-for-californians/",
      "date": "2026-06-29",
      "type": "case-study",
      "added": "2026-07-05",
      "superseded_by": null,
      "window": null,
      "explanation": "California partnered with Anthropic for Claude deployment across all state agencies for policy deliberation, Medicaid workflows, and cyber triage, demonstrating institutional-scale decision-support adoption with government endorsement, though outcome metrics remain unquantified."
    },
    {
      "title": "Cisco: 72% of CEOs Expect AI to Support Human-Directed Decisions",
      "url": "https://www.uctoday.com/productivity-automation/ceos-bullish-on-ai-but-fear-theyre-falling-behind-cisco-research-finds/",
      "date": "2026-06-29",
      "type": "adoption-metric",
      "added": "2026-07-05",
      "superseded_by": null,
      "window": null,
      "explanation": "Survey of 2,500 CEOs reveals 72% expect AI to support or execute work under human direction and governance, not autonomous decisions; identifies trust/security concerns (31%) and data quality (34%) as barriers to safe decision-support deployment."
    },
    {
      "title": "Humans Disengage, Reasoning Models Persist (arXiv:2606.26502)",
      "url": "https://24-ai.news/en/news/2026-06-28/arxiv-reasoning-tokens-failure-divergence/",
      "date": "2026-06-28",
      "type": "research-paper",
      "added": "2026-07-05",
      "superseded_by": null,
      "window": null,
      "explanation": "Large reasoning models spend MORE tokens on failed tasks versus solved ones (inverse of human behavior), with effect size Cohen's d 1.47–3.13, meaning longer reasoning chains are not reliable confidence signals—models cannot self-regulate effort or recognize when to defer to humans."
    },
    {
      "title": "Strategic Adoption of AI-Enabled Decision-Making Systems: Human Agency Design Study",
      "url": "https://commonplace.workforcefutures.net/paper/semantic_scholar:2ec753425f3d98cfdb126fa1f7ef22c693a25680",
      "date": "2026-06-28",
      "type": "research-paper",
      "added": "2026-07-05",
      "superseded_by": null,
      "window": null,
      "explanation": "Albabtain's mixed-methods quasi-experiment found AI design choices (transparency, override friction) and governance structures determine whether systems augment or substitute judgment; augmentation requires low-friction overrides and performance metrics rewarding contextual judgment, not AI compliance."
    },
    {
      "title": "Cox Communications and Kai: Production Decision-Support Deployments",
      "url": "https://www.my2cents.ai/news/2026-06-26-anthropic/",
      "date": "2026-06-26",
      "type": "case-study",
      "added": "2026-07-05",
      "superseded_by": null,
      "window": null,
      "explanation": "Cox Communications achieved 7x first-year ROI on multi-agent sales decisioning across 15,000 employees; Kai's autonomous security triage eliminated 99.5% false positives across 2.5M findings, demonstrating operational maturity in bounded decision domains."
    },
    {
      "title": "Nature Medicine: Large-Scale RCT of AI Clinical Decision Support in Kenya",
      "url": "https://medicalxpress.com/news/2026-06-ai-tool-clinician-decisions-real.html",
      "date": "2026-06-26",
      "type": "research-paper",
      "added": "2026-07-05",
      "superseded_by": null,
      "window": null,
      "explanation": "Pragmatic cluster-randomized trial (9,600+ patients, 16 sites) found AI-assisted clinical decision support improved documentation quality but failed to reduce treatment failure (2.2% vs 2.0%, P=0.13), revealing decision-support ROI gap between capability and outcome impact."
    },
    {
      "title": "Atheni AI: Structured Workflows as Missing Adoption Layer",
      "url": "https://tech.eu/2026/06/26/companies-bought-the-ai-now-they-need-people-to-use-it/",
      "date": "2026-06-26",
      "type": "case-study",
      "added": "2026-07-05",
      "superseded_by": null,
      "window": null,
      "explanation": "UK startup identified that 55% of professionals cite lack of structured human-AI workflows as adoption barrier; AI-adoption challenges are framework/coordination problems, not tool access—organizations need shared prompt libraries, training, and quality standards for sustained decision-support integration."
    },
    {
      "title": "HBR: Teach Your AI How You Make Decisions",
      "url": "https://hbr.org/2026/06/teach-your-ai-how-you-make-decisions?ab=HP-latest-3",
      "date": "2026-06-25",
      "type": "opinion",
      "added": "2026-07-05",
      "superseded_by": null,
      "window": null,
      "explanation": "By Harvard AI Institute director and HBS leaders: core bottleneck for AI transformation is not technology access but organizational ability to make decision-making processes explicit—articulating decision types, flows, criteria, trade-off rules, and escalation conditions required for consistent agent performance."
    },
    {
      "title": "Stroop Task Reveals AI Executive Control Failure Under Cognitive Load",
      "url": "https://www.psypost.org/advanced-ai-models-suffer-a-near-total-collapse-on-classic-psychology-test-as-cognitive-demands-increase/",
      "date": "2026-06-23",
      "type": "research-paper",
      "added": "2026-07-05",
      "superseded_by": null,
      "window": null,
      "explanation": "PNAS Nexus study: GPT-4o and Claude 3.5 collapse on Stroop task as cognitive load increases (91%→1% accuracy), unable to suppress prepotent responses, indicating architectural limitation in executive control needed for reliable reasoning under conflicting information."
    },
    {
      "title": "Arachne: AI's Reliability Gap is the Binding Constraint",
      "url": "https://arachnemag.substack.com/p/ais-reliability-gap",
      "date": "2026-06-23",
      "type": "opinion",
      "added": "2026-07-05",
      "superseded_by": null,
      "window": null,
      "explanation": "Substantive analysis synthesizing Anthropic's 81,000-person survey (unreliability is #1 concern over capability), developer productivity paradox (180% more code but 30% more releases), and Rabanser reliability framework showing frontier models commit 6-12% safety violations on critical tasks."
    },
    {
      "title": "Thomson Reuters 2026: AI Value Gap and Framework Requirements",
      "url": "https://www.lawnext.com/2026/06/thomson-reuters-future-of-professionals-report-warns-of-widening-gap-between-ai-adoption-and-ai-value.html",
      "date": "2026-06-22",
      "type": "adoption-metric",
      "added": "2026-07-05",
      "superseded_by": null,
      "window": null,
      "explanation": "1,816 professionals across legal, tax, audit, accounting report 91% experience AI value shortfall despite 74% regular use; 90% demand reasoning that can be explained and defended, identifying core decision-support framework requirement."
    },
    {
      "title": "AI-Enabled Leadership Operating System Case Study",
      "url": "https://crystalis.ai/case-study-ai-enabled-operating-system",
      "date": "2026-06-19",
      "type": "case-study",
      "added": "2026-06-21",
      "superseded_by": null,
      "window": null,
      "explanation": "Production deployment of AI decision-support system embedding company reasoning standards into leadership decisions; 100% adoption across 5 senior leaders with measurable behavioral change (reflex escalation becoming structured ownership-first decision-making)."
    },
    {
      "title": "AI Risks and Trustworthiness - AIRC - NIST AI Resource Center",
      "url": "https://airc.nist.gov/airmf-resources/airmf/3-sec-characteristics/",
      "date": "2026-06-18",
      "type": "industry-report",
      "added": "2026-06-21",
      "superseded_by": null,
      "window": null,
      "explanation": "NIST AI Risk Management Framework 1.0 provides normative trustworthiness framework (valid/reliable, safe, secure, resilient, accountable, transparent, explainable, fair) grounding evaluation of all decision-support systems across accuracy, robustness, and accountability dimensions."
    },
    {
      "title": "Agentic Coding and Persistent Returns to Expertise",
      "url": "https://www.anthropic.com/research/claude-code-expertise?ms=email",
      "date": "2026-06-11",
      "type": "research-paper",
      "added": "2026-06-21",
      "superseded_by": null,
      "window": null,
      "explanation": "Analysis of ~400,000 interactive Claude Code sessions reveals human-AI decision-making patterns: people make planning decisions (70%) while AI makes execution decisions (20%), demonstrating decision delegation frameworks and domain expertise amplifying agentic utility in practice."
    },
    {
      "title": "Case Study: Implementing a Trustworthy AI Framework for Veterans Affairs (VA)",
      "url": "https://www.ellumen.com/news/case-study-implementing-a-trustworthy-ai-framework-for-veterans-affairs-(va)",
      "date": "2026-06-11",
      "type": "case-study",
      "added": "2026-06-21",
      "superseded_by": null,
      "window": null,
      "explanation": "VA production deployment of clinical decision-support system with embedded governance framework (transparency, safety, explainability, accountability); repeatable, auditable process improving decision quality for healthcare providers with integrated LIME interpretability."
    },
    {
      "title": "Capacity, Not Format: Rethinking Structured Reasoning Failures",
      "url": "https://arxiv.org/abs/2606.09410",
      "date": "2026-06-08",
      "type": "research-paper",
      "added": "2026-06-21",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical research isolating format effects from reasoning capacity: structured output (JSON) degrades accuracy in models near capacity limits (Haiku 36.2pp drop, GPT-4o-mini 28.0pp), revealing fundamental tension between output structure and reasoning capability."
    },
    {
      "title": "SEI and Accenture Release AI Adoption Maturity Model to Help Organizations Scale AI with Predictable Outcomes",
      "url": "https://finance.yahoo.com/sectors/technology/articles/sei-accenture-release-ai-adoption-120500639.html",
      "date": "2026-06-08",
      "type": "product-ga",
      "added": "2026-06-21",
      "superseded_by": null,
      "window": null,
      "explanation": "CMU/Accenture empirically validated AI Adoption Maturity Model (8 dimensions, 5 levels) field-tested with Fortune 500 companies; addresses root cause of 95% zero-return adoption—mismatched expectations and poorly executed implementation, not technology."
    },
    {
      "title": "Knowledge workers' trust and reception of generative AI's advice in complex tasks",
      "url": "https://researchers.mq.edu.au/en/publications/knowledge-workers-trust-and-reception-of-generative-ais-advice-in/",
      "date": "2026-06-05",
      "type": "research-paper",
      "added": "2026-06-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed study challenges algorithm aversion stereotype; reveals reception type (dominant vs. oppositional) significantly influences advice-taking, with innovativeness and prior exposure as key trust determinants in knowledge-worker decision contexts."
    },
    {
      "title": "I read a multi-agent reasoning paper, built the Claude-native version, and measured everything",
      "url": "https://dev.to/bhj37193/i-read-a-multi-agent-reasoning-paper-built-the-claude-native-version-and-measured-everything-211p",
      "date": "2026-06-01",
      "type": "case-study",
      "added": "2026-06-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Developer deploys RecursiveMAS reasoning framework using extended thinking; read-after architecture achieves 65.5% accuracy at 2.7x cost with named use cases (legal review, medical triage, financial analysis), demonstrating practical multi-agent decision support."
    },
    {
      "title": "AI agent adoption hits 95% failure rate in enterprise pilots",
      "url": "https://www.techjournal.uk/p/ai-agent-adoption-hits-95-failure",
      "date": "2026-05-29",
      "type": "adoption-metric",
      "added": "2026-06-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Vendor analysis (1000+ companies) identifies 95% pilot failure from organizational sequencing failures (top-down mandates, weak training, IT gatekeeping), proposing bottom-up adoption framework to reverse failure patterns in decision-support deployment."
    },
    {
      "title": "The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary",
      "url": "https://arxiv.org/abs/2606.00376v1",
      "date": "2026-05-29",
      "type": "research-paper",
      "added": "2026-06-07",
      "superseded_by": null,
      "window": null,
      "explanation": "ICML 2026 research proves tool-integrated reasoning (86-94% accuracy) dramatically outperforms neural-only chain-of-thought (24-42%), establishing theoretical framework that hybrid decision-support systems exceed pure AI reasoning."
    },
    {
      "title": "AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?",
      "url": "https://arxiv.org/abs/2605.28255",
      "date": "2026-05-27",
      "type": "research-paper",
      "added": "2026-06-07",
      "superseded_by": null,
      "window": null,
      "explanation": "ACL 2026 peer-reviewed study (1,440 adoption decisions) reveals confirmation bias drives 64.5% under-reliance when AI agrees with humans' initial incorrect answer, exposing core adoption barrier in human-AI decision collaboration."
    },
    {
      "title": "AI Predicts Startup Success Better Than Expert Panels: Adding Humans Makes It Worse",
      "url": "https://www.techtimes.com/articles/317288/20260527/ai-predicts-startup-success-better-than-expert-panels-adding-humans-makes-it-worse.htm",
      "date": "2026-05-27",
      "type": "adoption-metric",
      "added": "2026-06-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Prospective forecasting study (Michigan Ross) on 30 Kickstarter ventures shows frontier AI ranks 0.74 correlation vs. experts 0.04-0.45; human+AI hybrid reduces accuracy, challenging assumed value of human-in-the-loop decision support."
    },
    {
      "title": "AI bots ignore evidence. Can we trust them with science?",
      "url": "https://www.sciencenews.org/article/ai-ignore-evidence-trust-science",
      "date": "2026-05-27",
      "type": "research-paper",
      "added": "2026-06-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed study: AI agents ignored evidence 68% of tasks, made unsupported claims 53%, used contradictory evidence to revise output only 26%, revealing inability to incorporate experimental data—core failure mode in decision support reliability."
    },
    {
      "title": "The model is accurate. The UX is clean. So why don't users trust it?",
      "url": "https://behavioralinsight.substack.com/p/the-model-is-accurate-the-ux-is-clean",
      "date": "2026-05-27",
      "type": "opinion",
      "added": "2026-06-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Applies Expectancy Violation Theory to explain overconfident AI systems trigger defensive disengagement; cites Stanford study finding chatbots validated flawed reasoning 73%, explaining why system polish undermines human critical evaluation in decision-making."
    },
    {
      "title": "Medical AI's diagnostic promise shrinks under clinical trial scrutiny",
      "url": "https://www.devdiscourse.com/article/technology/3916876-medical-ais-diagnostic-promise-shrinks-under-clinical-trial-scrutiny",
      "date": "2026-05-23",
      "type": "industry-report",
      "added": "2026-05-24",
      "superseded_by": null,
      "window": null,
      "explanation": "Meta-analysis of 5 RCTs (12,657 participants): AI clinical decision support produces small, marginal improvement in diagnostic accuracy; statistically significant but narrowly above zero; strongest in radiology, weakest in complex reasoning; signals overstated capability claims."
    },
    {
      "title": "Claude Skills for Executives: How to Build an AI Chief of Staff in 2026",
      "url": "https://www.claudecodehq.com/blog/claude-skills-executive-ai-chief-of-staff",
      "date": "2026-05-22",
      "type": "opinion",
      "added": "2026-05-24",
      "superseded_by": null,
      "window": null,
      "explanation": "Executive decision-support framework: persistent context (priorities, style, constraints), three core skills (Chief of Staff briefings, weekly accountability reviews, strategic sounding board); demonstrates personal effectiveness decision frameworks in production."
    },
    {
      "title": "Building AI Agents with Claude in 2026 - Tool Use, Workflows, and Automation Best Practices",
      "url": "https://www.blockchain-council.org/claude-ai/building-ai-agents-with-claude-2026-tool-use-workflows-automation-best-practices/",
      "date": "2026-05-22",
      "type": "opinion",
      "added": "2026-05-24",
      "superseded_by": null,
      "window": null,
      "explanation": "Practitioner decision patterns for production reliability: Clarify-Then-Act, Plan-Then-Execute with bounded steps, Human-in-the-Loop gates for high-risk decisions, observability for reasoning chains; enterprise governance framework."
    },
    {
      "title": "Beyond Model Readiness: Institutional Readiness for AI Deployment in Public Systems",
      "url": "https://arxiv.org/abs/2605.17203",
      "date": "2026-05-17",
      "type": "research-paper",
      "added": "2026-05-24",
      "superseded_by": null,
      "window": null,
      "explanation": "ICML workshop paper: Institutional Alignment Readiness (IAR) framework shows decision-support systems fail at deployment, not performance; case studies document institutional barriers (approvals, oversight capacity, fiscal sustainability) blocking scaling despite technical viability."
    },
    {
      "title": "Most AI Agents Will Pass Pilot and Fail Audit as Governance Gap Widens",
      "url": "https://theagenttimes.com/articles/most-ai-agents-will-pass-pilot-and-fail-audit-as-governance--af2bef73",
      "date": "2026-05-17",
      "type": "adoption-metric",
      "added": "2026-05-24",
      "superseded_by": null,
      "window": null,
      "explanation": "Survey of 650 leaders: 78% ran AI agent pilots, only 14% scaled to production; 90% of agent code fails EU AI Act compliance; governance failure (audit trails, escalation, rollback, ownership) is primary blocker, not performance; architectural prerequisite for production."
    },
    {
      "title": "SAP Bets 224 Agents and Claude Will Own the Enterprise",
      "url": "https://techfastforward.com/articles/sap-224-agents-claude-anthropic-autonomous-enterprise-sapphire-2026",
      "date": "2026-05-14",
      "type": "case-study",
      "added": "2026-05-24",
      "superseded_by": null,
      "window": null,
      "explanation": "SAP Sapphire 2026: 224 specialized Claude agents deployed across finance, supply chain, HR, procurement handling autonomous approvals and compliance; €100M investment; operates on €87T annual global commerce; largest production deployment of autonomous decision-making at enterprise scale."
    },
    {
      "title": "SAP and Anthropic Plan to Bring Claude to SAP Business AI Platform",
      "url": "https://news.sap.com/2026/05/sap-anthropic-to-bring-claude-sap-business-ai-platform/",
      "date": "2026-05-12",
      "type": "product-ga",
      "added": "2026-05-24",
      "superseded_by": null,
      "window": null,
      "explanation": "GA announcement: Claude as primary reasoning engine in SAP Business AI Platform; Joule agents execute decisions within existing controls, approvals, and compliance frameworks; hundreds of thousands enterprise customers across finance, HR, supply chain."
    },
    {
      "title": "Researchers Say 'Natural Decision-Making' Prompt Strategy Boosts AI Accuracy in Healthcare Advice",
      "url": "https://theaiinsider.tech/2026/05/12/researchers-say-natural-decision-making-prompt-strategy-boosts-ai-accuracy-in-healthcare-advice/",
      "date": "2026-05-12",
      "type": "research-paper",
      "added": "2026-05-24",
      "superseded_by": null,
      "window": null,
      "explanation": "TU Berlin study: psychological decision-making frameworks (Recognition-Primed Decision-Making, Data-Frame Theory) boost AI reasoning accuracy 13% → 30% in healthcare; structured human decision frameworks significantly improve AI decision quality."
    },
    {
      "title": "CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics",
      "url": "https://arxiv.org/abs/2605.09584v1",
      "date": "2026-05-10",
      "type": "case-study",
      "added": "2026-05-24",
      "superseded_by": null,
      "window": null,
      "explanation": "Production deployment: CLR-voyance clinical reasoning system live 6+ months at hospital drafting thousands of inpatient notes; 84.91% accuracy using POMDP framework with physician-validated outcome rubrics; demonstrates mature structured decision-support."
    },
    {
      "title": "AI Readiness Research: How Companies Make Software in 2026",
      "url": "https://sumatosoft.com/blog/research-business-ai-readiness",
      "date": "2026-05-08",
      "type": "case-study",
      "added": "2026-05-10",
      "superseded_by": null,
      "window": null,
      "explanation": "Survey (72 leaders, 30+ industries) shows 96% maintain human-in-the-loop on consequential decisions; 61% cite workflow redesign as primary enabler; governance framework is load-bearing for production deployment, not optional overhead."
    },
    {
      "title": "From Pilot Trap to Institutional Capacity: A Governance Framework for Sustainable Clinical AI Implementation in Health Systems",
      "url": "https://www.jmir.org/2026/1/e92680",
      "date": "2026-05-07",
      "type": "research-paper",
      "added": "2026-05-10",
      "superseded_by": null,
      "window": null,
      "explanation": "18-month empirical implementation study develops 6-module governance framework; reveals clinical AI systems remain confined to pilots not due to model limitations but institutional decision-making and governance capacity gaps."
    },
    {
      "title": "#c26: The Confidence Trap – Why AI Overconfidence Erodes Human Judgment and How Sovereign AI Offers an Exit",
      "url": "https://theaicommons.substack.com/p/c26-the-confidence-trap-why-ai-overconfidence",
      "date": "2026-05-07",
      "type": "opinion",
      "added": "2026-05-10",
      "superseded_by": null,
      "window": null,
      "explanation": "Analysis of AI overconfidence as design choice: RLHF incentivizes confidence without uncertainty; ECE 0.726 with 23% accuracy; standard calibration techniques reduce error 90% but not deployed—systemic design bias against decision-support reliability."
    },
    {
      "title": "Study Finds Brief AI Use May Hurt Problem Solving",
      "url": "https://creati.ai/ai-news/2026-05-07/study-finds-brief-ai-use-may-hurt-problem-solving/",
      "date": "2026-05-07",
      "type": "research-paper",
      "added": "2026-05-10",
      "superseded_by": null,
      "window": null,
      "explanation": "Research shows 10-minute AI assistance impairs independent problem-solving through cognitive offloading; reduced retention and analytical skepticism; unintended consequence: outsourcing thinking erodes decision autonomy and agency."
    },
    {
      "title": "The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure",
      "url": "https://arxiv.org/abs/2605.02398",
      "date": "2026-05-04",
      "type": "research-paper",
      "added": "2026-05-10",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical testing across 11 frontier models (67,221 records) reveals 8 collapse under adversarial pressure with 30.2pp accuracy drops; Anthropic Constitutional AI near-immune, indicating alignment-specific stability required for decision-support reliability."
    },
    {
      "title": "The AI Trust Problem: Why Enterprises Still Don't Deploy",
      "url": "https://valueaddvc.com/blog/the-ai-trust-problem-why-enterprises-still-dont-deploy",
      "date": "2026-05-04",
      "type": "opinion",
      "added": "2026-05-10",
      "superseded_by": null,
      "window": null,
      "explanation": "<20% of enterprise AI pilots reach production due to missing trust infrastructure: explainability, audit trails, governance, liability clarity. Core barrier is not capability but decision-framework accountability structures."
    },
    {
      "title": "Shadow AI risks: why raw language models fail expert tasks",
      "url": "https://www.ability.ai/blog/shadow-ai-risks-model-failures",
      "date": "2026-05-03",
      "type": "adoption-metric",
      "added": "2026-05-10",
      "superseded_by": null,
      "window": null,
      "explanation": "Analysis of 5.5M real-world interactions shows top models fail ~9% overall, 14-16% on expert decision-making tasks (finance/law/medical); performance improvement flatlined despite compute scaling, exposing adoption barriers."
    },
    {
      "title": "Thinking Fast, Slow, Artificially: AI and Your Brain",
      "url": "https://executiveeducation.wharton.upenn.edu/thought-leadership/wharton-at-work/2026/05/thinking-fast-slow-and-artificially/",
      "date": "2026-05-01",
      "type": "research-paper",
      "added": "2026-05-10",
      "superseded_by": null,
      "window": null,
      "explanation": "Wharton RCT (1,300+ participants) shows AI-assisted decisions improve 25pp when correct but degrade 15pp when wrong; overconfidence persists even at 50% error rate, demonstrating cognitive surrender risk in decision frameworks."
    },
    {
      "title": "The 2026 AI Index: Capability Without Accountability",
      "url": "https://cloudtweaks.com/2026/05/standford-2026-ai-index/",
      "date": "2026-05-01",
      "type": "industry-report",
      "added": "2026-05-10",
      "superseded_by": null,
      "window": null,
      "explanation": "Stanford HAI's 2026 Index documents hallucination rates 22-94% and knowledge-belief reasoning failures where accuracy collapses 90%+ under false premises, revealing capability-accountability gap in professional decision contexts."
    },
    {
      "title": "To Build or Not to Build? Factors that Lead to Non-Development or Abandonment of AI Systems",
      "url": "https://arxiv.org/abs/2604.28053v1",
      "date": "2026-04-30",
      "type": "research-paper",
      "added": "2026-05-10",
      "superseded_by": null,
      "window": null,
      "explanation": "FAccT 2026 mixed-methods study (incident database + practitioner survey) reveals AI system abandonment driven by organizational dynamics and resource constraints, not ethics concerns; directly applicable to decision-support deployment barriers."
    },
    {
      "title": "Appian World 2026: Where AI Finally Grew Up",
      "url": "https://xebia.com/articles/appian-world-2026-ai/",
      "date": "2026-04-30",
      "type": "industry-report",
      "added": "2026-05-10",
      "superseded_by": null,
      "window": null,
      "explanation": "Conference reporting production deployments of bounded AI decision systems in insurance, pharma, financial crime with measured ROI (50%+ claims processing gains); evidence of transition from pilots to integrated decision workflows."
    },
    {
      "title": "AI Failure Modes Are Now a Top-of-Stack Concern: An Engineering Defense Playbook",
      "url": "https://conectia.pro/en/blog/ai-failure-modes-engineering-defenses-stanford-2026",
      "date": "2026-04-28",
      "type": "industry-report",
      "added": "2026-05-10",
      "superseded_by": null,
      "window": null,
      "explanation": "Stanford Emerging Tech Review enumerates critical failure modes (hallucinations, overtrust, adversarial injection) and engineering defenses; notes most teams shipping decision-support features have implemented none, indicating significant maturity gap."
    },
    {
      "title": "New Study Sheds Light on the Kinds of Reasoning Tasks that AI Systems Still Struggle to Do Well",
      "url": "https://kempnerinstitute.harvard.edu/news/new-study-sheds-light-on-the-kinds-of-reasoning-tasks-that-ai-systems-still-struggle-to-do-well/",
      "date": "2026-04-24",
      "type": "research-paper",
      "added": "2026-04-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Harvard research identifies relational complexity as core limitation: AI accuracy drops sharply when decisions require weighing multiple interacting factors simultaneously—directly constrains multi-factor analysis in healthcare and strategic decision-making."
    },
    {
      "title": "Stanford's 2026 AI Index Highlights Rapid Growth and Widening Governance Gaps",
      "url": "https://www.jdsupra.com/legalnews/stanford-s-2026-ai-index-highlights-7540386/",
      "date": "2026-04-24",
      "type": "industry-report",
      "added": "2026-04-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Stanford Index analysis documents scaling (88% organizational adoption) alongside reliability collapse: hallucination rates 22-94%, evaluation gaps between benchmarks and deployment—core constraints on decision-support reliability at scale."
    },
    {
      "title": "The Impact of AI Reliance on Sense of Agency and Indecisiveness",
      "url": "https://www.ijfmr.com/research-paper.php?id=75548",
      "date": "2026-04-22",
      "type": "research-paper",
      "added": "2026-04-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Increased AI trust correlates with reduced sense of agency (r=.288, p<.01) and increased indecisiveness—critical negative signal: reliance on AI decision systems degrades individual decision autonomy and confidence."
    },
    {
      "title": "Why AI Training Doesn't Build Capability — and What Judgment Training Actually Looks Like for AI Decision Making",
      "url": "https://www.withum.ai/resources/why-ai-training-doesnt-build-capability-and-what-judgment-training-actually-looks-like-for-ai-decision-making/",
      "date": "2026-04-22",
      "type": "opinion",
      "added": "2026-04-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Practitioner framework bridging speed-quality gap: proposes specific human judgment training for AI decision-making contexts. Addresses gap between productivity gains and actual decision quality improvement in AI-assisted reasoning."
    },
    {
      "title": "Learning When Not to Decide: A Framework for Overcoming Factual Presumptuousness in AI Adjudication",
      "url": "https://arxiv.org/abs/2604.19895",
      "date": "2026-04-21",
      "type": "research-paper",
      "added": "2026-04-26",
      "superseded_by": null,
      "window": null,
      "explanation": "SPEC framework addresses presumptuousness (confident answers despite insufficient evidence) in legal decision-making. Achieves 89% accuracy vs. 15% baseline RAG on incomplete-info cases—positive signal for bounded reasoning frameworks."
    },
    {
      "title": "Perils of Declining Judgment in the Age of AI",
      "url": "https://rpc.cfainstitute.org/blogs/enterprising-investor/2026/essay-perils-declining-judgment-ai",
      "date": "2026-04-21",
      "type": "opinion",
      "added": "2026-04-26",
      "superseded_by": null,
      "window": null,
      "explanation": "CFA Institute analysis: AI weakens epistemic foundations through cognitive delegation and 'knowledge-collapse equilibrium'; argues decision authority must remain anchored in evidence-based human inquiry—governance framework for responsible decision support."
    },
    {
      "title": "80% Reject AI at Work: CRE Investor Impact 2026",
      "url": "https://www.theaiconsultingnetwork.com/blog/enterprise-ai-adoption-resistance-80-percent-workers-reject-cre-investors-2026",
      "date": "2026-04-19",
      "type": "adoption-metric",
      "added": "2026-04-26",
      "superseded_by": null,
      "window": null,
      "explanation": "WalkMe study (3,750 professionals): 9% trust AI for complex decisions vs. 61% executive claims; 80% reject enterprise AI—quantifies massive gap between adoption narratives and actual decision-support reliance."
    },
    {
      "title": "Overreliance on AI programs may undermine confidence at work",
      "url": "https://www.eurekalert.org/news-releases/1123722",
      "date": "2026-04-16",
      "type": "research-paper",
      "added": "2026-04-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Behavioral study (1,923 adults) shows passive reliance on AI reduces independent judgment confidence and sense of authorship; active engagement mitigates effect—critical negative signal that decision-support tool design fundamentally shapes reasoning capability erosion."
    },
    {
      "title": "Trust, attitudes and use of artificial intelligence: A global study 2025",
      "url": "https://kpmg.com/be/en/insights/technology/ai-insights/trust-attitudes-and-use-of-artificial-intelligence-a-global-study-2025.html",
      "date": "2026-04-15",
      "type": "industry-report",
      "added": "2026-04-26",
      "superseded_by": null,
      "window": null,
      "explanation": "48,000 respondents across 47 countries identify trust as critical adoption barrier for AI decision-support. Documents gap between perceived benefits and actual system reliance—core adoption blocker."
    },
    {
      "title": "Stanford's 2026 AI Index Has a Warning: We're Building Faster Than We Can Measure",
      "url": "https://shshell.com/blog/stanford-ai-index-2026-measurement-crisis",
      "date": "2026-04-14",
      "type": "industry-report",
      "added": "2026-04-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Documents 'Reliability Gap': capability scales 2-3x annually while reliability only 1.2-1.5x. Models solve graduate physics yet fail analog clocks (50.6% accuracy)—demonstrates maturity asymmetry that blocks reliable decision-support scaling."
    },
    {
      "title": "Web-Based Personalized Machine Learning Recommendations to Enhance Shared Decision-Making in Prostate-Specific Antigen Screening: Randomized Controlled Trial",
      "url": "https://aging.jmir.org/2026/1/e83238",
      "date": "2026-04-13",
      "type": "case-study",
      "added": "2026-04-26",
      "superseded_by": null,
      "window": null,
      "explanation": "RCT (367 participants) shows ML-generated decision support reduces decisional conflict 3.77 points and increases satisfaction 7.38 points; behavioral outcomes confirm acceptance recommendation adoption (50.7% vs 24.2%), demonstrating measurable deployment value in medical decision-making."
    },
    {
      "title": "Why 95% of AI Pilots Produce Zero ROI | Yasmeen Ahmad, Google Cloud",
      "url": "https://www.theaireport.ai/articles/why-95-of-ai-pilots-produce-zero-roi--yasmeen-ahmad-google-cloud",
      "date": "2026-04-09",
      "type": "industry-report",
      "added": "2026-04-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Google Cloud executive identifies governance and continuous evaluation as mandatory infrastructure for scaling AI agents from pilot to production; 95% failure rate stems from fragmented data and missing orchestration, not model performance."
    },
    {
      "title": "Beyond Accuracy: Diagnosing Algebraic Reasoning Failures in LLMs Across Nine Complexity Dimensions",
      "url": "https://arxiv.org/abs/2604.06799",
      "date": "2026-04-08",
      "type": "research-paper",
      "added": "2026-04-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical testing of 7 models (8B-235B) reveals working memory as scale-invariant bottleneck—all models collapse at 20-30 parallel branches regardless of size, exposing fundamental architectural constraints limiting reasoning systems."
    },
    {
      "title": "Reasoning Fails Where Step Flow Breaks",
      "url": "https://arxiv.org/abs/2604.06695",
      "date": "2026-04-08",
      "type": "research-paper",
      "added": "2026-04-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed research (ACL 2026) identifies shallow lock-in and deep decay in reasoning models; proposes test-time StepFlow intervention improving accuracy without retraining, balancing documented failures with technical solutions."
    },
    {
      "title": "CMU Study: Top LLMs Fail Simple Contradiction Tests, Lack True Reasoning",
      "url": "https://gentic.news/article/cmu-study-top-llms-fail-simple",
      "date": "2026-04-06",
      "type": "news-coverage",
      "added": "2026-04-12",
      "superseded_by": null,
      "window": null,
      "explanation": "CMU research testing 14 leading LLMs (GPT-4, Claude 3, Gemini) reveals all fail simple logical contradiction detection, exposing benchmark illusion: curated-dataset performance masks lack of robust foundational reasoning."
    },
    {
      "title": "Your AI Vendor Claims Their LLM Can Reason. Here's What's Actually Happening.",
      "url": "https://zaruko.com/insights/your-llm-cannot-reason",
      "date": "2026-04-06",
      "type": "opinion",
      "added": "2026-04-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Critical analysis explaining LLMs predict tokens, not reason formally; Apple's GSM-Symbolic study shows single irrelevant clause causes up to 65% accuracy drop, demonstrating sophisticated pattern-matching rather than rule-based reasoning."
    },
    {
      "title": "Fragile Reasoning: A Mechanistic Analysis of LLM Sensitivity to Meaning-Preserving Perturbations",
      "url": "https://papers.cool/arxiv/2604.01639",
      "date": "2026-04-02",
      "type": "research-paper",
      "added": "2026-04-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Semantic variants of 677 GSM8K problems show 28.8%-45.1% answer-flip rates despite meaning preservation; entangled failures are only 5.2%-12.2% repairable, demonstrating fundamental brittleness in mathematical reasoning."
    },
    {
      "title": "The Algorithmic Breach: Why Public AI is the Newest Liability in the Boardroom",
      "url": "https://athenaboard.com/2026/04/02/the-algorithmic-breach-why-public-ai-is-the-newest-liability-in-the-boardroom/",
      "date": "2026-04-02",
      "type": "case-study",
      "added": "2026-04-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Federal Court of Australia judgment (ASIC v Bekier) establishes liability for informal AI use in board decisions; identifies specific risks (hallucination, outsourced thinking, shadow AI, data sovereignty) and sets precedent for decision-support accountability."
    },
    {
      "title": "Assessing AI for Criminal Justice: A User Decision Framework",
      "url": "https://counciloncj.org/assessing-ai-for-criminal-justice-a-user-decision-framework/",
      "date": "2026-03-30",
      "type": "case-study",
      "added": "2026-04-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Authoritative multi-phase assessment framework (CCJ Task Force, 15 leaders) operationalizes AI adoption decisions with rigorous validation, demographic impact assessment, and mandatory human override—demonstrates decision-support governance in high-stakes deployment."
    },
    {
      "title": "Chain of Thought Faithfulness: Why LLM Reasoning Is Often a Narrative",
      "url": "https://explore.n1n.ai/blog/llm-cot-faithfulness-research-2026-03-30",
      "date": "2026-03-30",
      "type": "research-paper",
      "added": "2026-04-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Anthropic's May 2025 research finds Claude 3.7 Sonnet discloses hint usage only 25% of the time—chain-of-thought is post-hoc rationalization, not faithful reasoning log, critically undermining use of AI reasoning chains for verification."
    },
    {
      "title": "Cognitive government accelerated | Deloitte Insights",
      "url": "https://www.deloitte.com/us/en/insights/industry/government-public-sector-services/government-trends/2026/cognitive-government-ai-driven-decision-making.html",
      "date": "2026-03-29",
      "type": "industry-report",
      "added": "2026-04-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Real-world government deployments of AI decision-support systems (Japan tsunami impact modeling, Sweden autonomous defibrillator drones) demonstrate operational viability in high-stakes emergency-response contexts."
    },
    {
      "title": "Who owns compliance decisions in automated systems?",
      "url": "https://fintech.global/globalregtechsummitusa/who-owns-compliance-decisions-in-automated-systems/",
      "date": "2026-03-25",
      "type": "opinion",
      "added": "2026-03-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Multi-expert RegTech consensus establishes decision-support governance framework: humans remain accountable; AI provides alerts/recommendations under human oversight; mature firms separate rule design, escalation, and model validation ownership."
    },
    {
      "title": "Evaluating Reasoning Models: Think Tokens, Steps, and Accuracy Tradeoffs - RIO World",
      "url": "https://rioworld.org/evaluating-reasoning-models-think-tokens-steps-and-accuracy-tradeoffs",
      "date": "2026-03-25",
      "type": "research-paper",
      "added": "2026-03-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Analysis of reasoning model performance reveals three regimes: low-complexity tasks degrade 4.7-8.2pp, medium-complexity improve 9.1-12.3pp (adoption sweet spot), high-complexity collapse below 5%; 5x cost multiplier limits deployment beyond medium-complexity decisions."
    },
    {
      "title": "CRYSTAL Benchmark Exposes How AI Models Fake Reasoning",
      "url": "https://aibytes.blog/benchmarks/crystal-benchmark-exposes-how-ai-models-fake-reasoning",
      "date": "2026-03-23",
      "type": "research-paper",
      "added": "2026-03-29",
      "superseded_by": null,
      "window": null,
      "explanation": "CRYSTAL benchmark with 6,372 questions reveals 19/20 models skip 50%+ of reasoning steps—achieving 58% accuracy while recovering only 48% of actual reasoning, proving models pattern-match rather than genuinely reason."
    },
    {
      "title": "Business Judgement Rule in the use of AI: how governing bodies are liable for decisions",
      "url": "https://kpmg-law.de/en/business-judgement-rule-in-the-use-of-ai-how-governing-bodies-are-liable-for-decisions/",
      "date": "2026-03-19",
      "type": "opinion",
      "added": "2026-03-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Legal analysis establishes AI cannot relieve executives of decision review duty; requires mandatory governance documentation including user responsibility, AI tool version, prompts used, and evidence of context/source verification."
    },
    {
      "title": "BrainBench: Exposing the Commonsense Reasoning Gap in Large Language Models",
      "url": "https://arxiv.org/abs/2603.14761",
      "date": "2026-03-16",
      "type": "research-paper",
      "added": "2026-03-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Benchmark across 8 frontier models shows even Claude Opus 4.6 exhibits 6-16 percentage-point gaps between accuracy and consistency, revealing stochastic reasoning—key reliability issue for decision-support deployment."
    },
    {
      "title": "Are LLMs Really Smart? Dissecting AI's Reasoning Failures",
      "url": "https://www.sotaaz.com/post/llm-reasoning-failures-en",
      "date": "2026-03-14",
      "type": "research-paper",
      "added": "2026-03-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Stanford taxonomy from 500+ papers classifies reasoning failures on two axes: formal vs informal reasoning, and fundamental architectural limitations (unsolvable by scaling) versus application-specific shortcomings—framework for understanding decision-support constraints."
    },
    {
      "title": "How Companies Will Measure the ROI of Agentic AI in 2026?",
      "url": "https://www.dataexpertise.in/roi-of-agentic-ai-business-value-2026/",
      "date": "2026-03-14",
      "type": "opinion",
      "added": "2026-03-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Framework for agentic AI ROI emphasizing decision quality, speed, and autonomous multi-step reasoning; metrics shift from cost-savings to value-creation (revenue uplift, risk-avoided, cycle-time reduction) across interconnected decision workflows."
    },
    {
      "title": "AI under scrutiny: accuracy, accountability and risk - Simpson Grierson",
      "url": "https://www.simpsongrierson.com/insights-news/legal-updates/ai-under-scrutiny-accuracy-accountability-and-risk",
      "date": "2026-03-02",
      "type": "case-study",
      "added": "2026-03-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Three documented legal failures: NZ courts ruled AI-hallucinated citations may amount to obstruction of justice; Deloitte refunded AUD 440K for AI-generated errors; regulatory finding that accountability for accuracy cannot be delegated to AI."
    },
    {
      "title": "causaLens: Reliable Digital Workers",
      "url": "https://causalens.com",
      "date": "2026-02-19",
      "type": "product-ga",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "GA platform providing orchestrated AI agents with decision-making automation; customer deployments at Johnson & Johnson and McCann Worldgroup with claimed 5x ROI and documented time-to-treatment improvements in healthcare operations."
    },
    {
      "title": "Enterprise AI Orchestration White Paper 2026 | Free Download",
      "url": "https://arpia.ai/whitepaper-feb-2026",
      "date": "2026-02-18",
      "type": "case-study",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Production deployment of ARPIA platform orchestrating ML, GenAI, and Agentic AI for financial services collections pipeline achieving 13-minute raw-data-to-strategy activation time, demonstrating feasible multi-AI coordination in decision-support workflows."
    },
    {
      "title": "Large Language Model Reasoning Failures - Rivista AI",
      "url": "https://www.rivista.ai/2026/02/16/large-language-model-reasoning-failures/",
      "date": "2026-02-16",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Summary of Stanford/Caltech research on systematic LLM reasoning failures: Reversal Curse, Robustness Fragility, and Working Memory Leaks—predictable error modes undermining reliance on AI reasoning trajectories for decision support reliability."
    },
    {
      "title": "Do We Have Enough Information? Assessing AI Clinical Decision Support Systems for Implementation in Primary Care - PubMed",
      "url": "https://pubmed.ncbi.nlm.nih.gov/41685476/?fc=20220524062322&ff=20260213065701&v=2.18.0.post22+67771e2",
      "date": "2026-02-12",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Peer-reviewed environmental scan of commercially available AI-CDSS solutions reveals critical transparency gaps: vendor proprietary algorithms lack training data disclosure, knowledge base rigor unspecified, and privacy details incomplete—blocking reliable implementation evaluation."
    },
    {
      "title": "2026 AI in Professional Services Report: AI adoption has hit critical scale",
      "url": "https://www.thomsonreuters.com/en-us/posts/technology/ai-in-professional-services-report-2026/",
      "date": "2026-02-09",
      "type": "industry-report",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Thomson Reuters survey (1,500+ professionals, 27 countries) finds organization-wide AI adoption in professional services nearly doubled to 40% in 2026 from 22% in 2025; yet only 18% track ROI, and 40% report client confusion on AI use policies."
    },
    {
      "title": "AI Project Failure Statistics 2026: The Complete Picture",
      "url": "https://www.pertamapartners.com/insights/ai-project-failure-statistics-2026",
      "date": "2026-02-08",
      "type": "adoption-metric",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Synthesis of 2,400+ enterprise AI initiatives shows 80.3% overall failure rate (33.8% abandoned, 28.4% deliver no value, 18.1% cannot justify costs); 95% GenAI pilot-to-production failure rate across sectors, signaling critical scaling barriers."
    },
    {
      "title": "Intelligent Reasoning Cues: A Framework and Case Study of the Roles of AI Information in Complex Decisions",
      "url": "https://arxiv.org/abs/2602.00259v1",
      "date": "2026-01-30",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "CHI 2026 research on AI decision support in clinical sepsis treatment; contextual study with 25 physicians identified that reasoning cues should target high-discretion tasks, yet AI systems often fail to support decisions despite technical accuracy."
    },
    {
      "title": "AI ROI Reality Check in 2026: Can we Close the Adoption Gap?",
      "url": "https://www.uctoday.com/unified-communications/ai-roi-reality-check-2026/",
      "date": "2026-01-29",
      "type": "news-coverage",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "PwC CEO survey (4,454 executives) found only 12% report AI delivered both cost and revenue benefits; Section survey shows employee-executive perception gap, exposing critical ROI and governance barriers to decision-support effectiveness."
    },
    {
      "title": "Causal AI Decision Intelligence: Why It Will Emerge in 2026",
      "url": "https://thecuberesearch.com/why-causal-ai-decision-intelligence-2026/",
      "date": "2026-01-23",
      "type": "industry-report",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Industry trend analysis citing 62% of 625 enterprise AI professionals planning to evolve beyond automation to AI decision intelligence; identifies 74% faithfulness gap in LLM/RAG pipelines highlighting current system limitations."
    },
    {
      "title": "Scaling AI in 2026- The foundations most enterprises still lack",
      "url": "https://genphase.ai/insights/scaling-ai-in-2026-the-foundations-most-enterprises-still-lack/",
      "date": "2026-01-19",
      "type": "opinion",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Analysis documenting 70-85% AI project failure rates due to data readiness and governance gaps; 42% of companies abandoned AI initiatives in 2025 (up from 17% in 2024), signaling persistent deployment barriers."
    },
    {
      "title": "2026 Predictions: The Architecture, Governance, and AI Trends Every Enterprise Must Prepare For",
      "url": "https://www.cloudera.com/blog/business/2026-predictions-the-architecture-governance-and-ai-trends-every-enterprise-must-prepare-for.html",
      "date": "2026-01-08",
      "type": "industry-report",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Cloudera leadership predicts 2026 transition from experimentation to operationalization; emphasizes data as living knowledge system enabling agentic workflows with strong governance frameworks as non-negotiable for safe decision-making at scale."
    },
    {
      "title": "A hybrid reasoning framework for transparent clinical decision support",
      "url": "https://www.accscience.com/journal/AIH/articles/online_first/6109",
      "date": "2026-01-06",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Hybrid AI meta-model evaluated on SEER/MIMIC datasets achieved 78% reduction in guideline-violation errors (18% to 6%) with clinician usability study showing 6.2/7 decision clarity rating, demonstrating feasible error reduction in clinical decision support."
    },
    {
      "title": "AI-Based Medical Decision Support: Exploring the Data Gap",
      "url": "https://pubmed.ncbi.nlm.nih.gov/41417882/",
      "date": "2025-12-19",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Peer-reviewed review in Journal of Bone & Joint Surgery showing AI's modest impact on patient care stems from data quality and EMR design flaws rather than algorithmic limitations, identifying concrete deployment barriers."
    },
    {
      "title": "AI Decision Intelligence in Marketing: G2's 2026 Industry Report",
      "url": "https://learn.g2.com/ai-decision-intelligence-marketing-report",
      "date": "2025-12-18",
      "type": "industry-report",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Industry report showing 26-75% customer adoption of AI decisioning features across marketing platforms with measurable impacts on conversion rates and ROI, signaling operational deployment maturity."
    },
    {
      "title": "Overreliance on AI risks eroding new and future doctors' critical thinking while reinforcing existing bias",
      "url": "https://bmjgroup.com/overreliance-on-ai-risks-eroding-new-and-future-doctors-critical-thinking-while-reinforcing-existing-bias/",
      "date": "2025-12-03",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "BMJ editorial warning that overreliance on GenAI risks eroding critical thinking, reinforcing biases, and causing deskilling in medical trainees, documenting persistent risks in decision-support adoption."
    },
    {
      "title": "Judges Warn Against Overreliance on AI in Courts, Cite Hallucinated Judgments",
      "url": "https://ibcworldnews.com/2025/12/01/judges-warn-against-overreliance-on-ai-in-courts-cite-hallucinated-judgments/",
      "date": "2025-12-01",
      "type": "news-coverage",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "News report on Indian judges warning against AI overreliance citing hallucinated citations and fabricated judgments, illustrating critical risks of AI decision-support in high-stakes legal domains."
    },
    {
      "title": "On the Justified Use of AI Decision Support in Evidence-Based Medicine: Validity, Explainability, and Responsibility",
      "url": "https://www.cambridge.org/core/journals/cambridge-quarterly-of-healthcare-ethics/article/on-the-justified-use-of-ai-decision-support-in-evidencebased-medicine-validity-explainability-and-responsibility/BECCD28924E332C8529EF3C98831F18F",
      "date": "2025-11-21",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Ethics journal article examining justified use of black-box AI in medical decision-making, addressing validity and explainability concerns central to responsible deployment in high-stakes clinical reasoning."
    },
    {
      "title": "Enterprise AI Agents: Beyond Productivity",
      "url": "https://www.ibm.com/think/insights/enterprise-ai-agents",
      "date": "2025-11-21",
      "type": "case-study",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "IBM internal deployment of agentic AI to 270,000 employees with estimated $4.5B productivity impact, demonstrating large-scale enterprise adoption of reasoning systems in operational decision workflows."
    },
    {
      "title": "SEI Tool Helps Federal Agencies Detect AI Bias and Build Trust",
      "url": "https://www.cmu.edu/news/stories/archives/2025/september/sei-tool-helps-federal-agencies-detect-ai-bias-and-build-trust",
      "date": "2025-09-17",
      "type": "product-ga",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "Carnegie Mellon SEI released AI Robustness (AIR) open-source tool for federal agencies to detect bias and improve decision-making through causal discovery, with DoD testing and national security deployment."
    },
    {
      "title": "Evidence-Driven Decision Support for AI Model Selection in Research Software Engineering",
      "url": "https://arxiv.org/html/2512.11984v1",
      "date": "2025-09-16",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "Research framework (ModelSelect) applying MCDM principles to AI model selection; validated with 50 real-world case studies showing high coverage and rationale alignment for structured decision support."
    },
    {
      "title": "MIT report: 95% of generative AI pilots at companies are failing",
      "url": "https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/",
      "date": "2025-08-18",
      "type": "news-coverage",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "MIT NANDA analysis of 150 interviews and 300 public deployments finding 95% of enterprise AI pilots fail to deliver value; vendor solutions succeed 67% vs internal builds 33%, exposing deployment execution gaps."
    },
    {
      "title": "Global survey reveals growing consumer trust in personal AI assistants",
      "url": "https://www.zendesk.com/newsroom/press-releases/global-survey-reveals-growing-consumer-trust-in-personal-ai-assistants/",
      "date": "2025-07-30",
      "type": "adoption-metric",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "YouGov survey of 10,000 respondents across 10 countries: 52% comfortable with personal AI for daily tasks, 64% for calendar/to-do, 39% for financial decisions; identifies trust factors (security, transparency, oversight)."
    },
    {
      "title": "Decision factors for the selection of AI-based decision support systems",
      "url": "https://journals.plos.org/plosone/article?id=10.1371%2Fjournal.pone.0328411",
      "date": "2025-07-24",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "Peer-reviewed PLOS ONE study deriving empirical decision weights for AI DSS selection in prognostics; shows performance, effort, and transparency as key decision factors with expertise-dependent user preferences."
    },
    {
      "title": "Don't Let AI Distort Your Decision-Making",
      "url": "https://hbr.org/tip/2025/07/dont-let-ai-distort-your-decision-making",
      "date": "2025-07-11",
      "type": "opinion",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "HBR tip warning executives that AI forecasts create overconfidence through impressive detail and trend extrapolation, citing research showing executives using GenAI made worse predictions than without AI."
    },
    {
      "title": "Close the ROI Gap When Scaling AI",
      "url": "https://guidehouse.com/insights/financial-services/2025/close-the-roi-gap-when-scaling-ai",
      "date": "2025-06-30",
      "type": "industry-report",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Analysis citing HBR data showing only 26% of companies have working AI products and only 4% achieve significant returns, exposing critical ROI gap in scaling AI decision-support systems."
    },
    {
      "title": "Gartner forecasts a steep cancellation rate for overpriced AI projects",
      "url": "https://www.differentiated.io/daily-news/2025-06-30",
      "date": "2025-06-30",
      "type": "news-coverage",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Gartner prediction that 40%+ of agentic AI projects will be cancelled by 2027 due to rising costs and unclear value; research shows AI agents fail ~70% of office tasks, exposing deployment challenges."
    },
    {
      "title": "The Future of Command: AI-Driven Decision-Support Systems in Defense",
      "url": "https://brite.ikeinstitute.org/brite_innovation_review_issue_35_may_2025/ai_driven_defense",
      "date": "2025-05-14",
      "type": "industry-report",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Industry report on AI-driven DSS in defense with case studies of Project Maven, UK autonomous warrior programme, and Israel's Iron Dome, demonstrating real-world deployment in high-stakes decision contexts."
    },
    {
      "title": "Implementation of artificial intelligence-based decision support systems for antibiotic prescribing in hospitals: a Delphi study",
      "url": "https://pubmed.ncbi.nlm.nih.gov/40352326/",
      "date": "2025-04-25",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Delphi consensus study with 36 healthcare experts on implementing AI-based DSS for antibiotic prescribing, identifying critical barriers including trust, transparency, and organizational readiness challenges."
    },
    {
      "title": "Psychological Factors Influencing Appropriate Reliance on AI in Clinical Decision Support Systems",
      "url": "https://www.jmir.org/2025/1/e58660",
      "date": "2025-04-04",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Study of 223 dermatologists found AI support improved accuracy by only 1%, with low reliance (10%) and high self-reliance (86%), highlighting persistent barriers to effective clinical decision-support adoption."
    },
    {
      "title": "AI Thinks Like Us: ChatGPT Mirrors Human Decision Biases",
      "url": "https://connect.informs.org/discussion/ai-thinks-like-us-flaws-and-all-new-study-finds-chatgpt-mirrors-human-decision-biases-in-half-the-tests",
      "date": "2025-04-01",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Research testing ChatGPT on 18 decision-making scenarios found AI made mistakes similar to humans in half the tests, showing biases like overconfidence and gambler's fallacy that undermine decision-support reliability."
    },
    {
      "title": "UKHSA Advisory Board: Update on artificial intelligence in UKHSA",
      "url": "https://www.gov.uk/government/publications/ukhsa-advisory-board-meeting-papers-march-2025/ukhsa-advisory-board-update-on-artifical-intelligence-in-ukhsa",
      "date": "2025-03-11",
      "type": "case-study",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "UK government production deployment of AI for TB screening achieved 90% accuracy and 85% reduction in manual review workload; real-time pollen monitoring system piloting, demonstrating concrete healthcare decision-support outcomes."
    },
    {
      "title": "Welfare system AI pilots dropped over 'false starts'",
      "url": "https://www.magzter.com/stories/newspaper/The-Guardian/WELFARE-SYSTEM-AI-PILOTS-DROPPED-OVER-FALSE-STARTS",
      "date": "2025-01-27",
      "type": "news-coverage",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "UK government scrapped half-a-dozen AI pilots for welfare systems (A-cubed, Aigent) due to scalability, reliability, and testing challenges; explicit abandonment of decision-support deployments due to technical barriers."
    },
    {
      "title": "Understanding the impact of Human-AI interaction on discrimination",
      "url": "https://policy-lab.ec.europa.eu/news/understanding-impact-human-ai-interaction-discrimination-2025-01-10_en",
      "date": "2025-01-10",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "EU Policy Lab research on credit lending and recruitment decisions found human overseers equally likely to follow biased AI recommendations; human oversight alone insufficient to prevent discrimination in AI decision-making."
    },
    {
      "title": "The Willingness of Doctors to Adopt Artificial Intelligence–Driven Clinical Decision Support Systems",
      "url": "https://www.jmir.org/2025/1/e62768/",
      "date": "2025-01-07",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "Survey of 450 physicians across hospital types in China identifies adoption pathways for AI-CDSSs; tertiary hospitals exhibit 6 distinct routes while primary hospitals show 3, with utility perception and readiness as key factors."
    },
    {
      "title": "What are the main limitations of AI reasoning models?",
      "url": "https://zilliz.com/ai-faq/what-are-the-main-limitations-of-ai-reasoning-models",
      "date": "2025-01-05",
      "type": "opinion",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "Critical assessment of AI reasoning model limitations: data bias, lack of common sense, and transparency gaps preventing effective decision support in high-stakes domains like healthcare and finance."
    },
    {
      "title": "Transparency and Authority Concerns with Using AI to Make Ethical Decisions",
      "url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC12391615/",
      "date": "2024-12-22",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Peer-reviewed critique of AI ethics consultations, identifying fundamental barriers: lack of transparency ('black box') and unresolved questions about AI authority in complex ethical reasoning decisions."
    },
    {
      "title": "Ethical Implications of AI-Driven Clinical Decision Support Systems on Healthcare Resource Allocation",
      "url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC11662436/",
      "date": "2024-12-21",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Qualitative study of healthcare professionals revealing significant ethical concerns in AI-CDSS: bias, transparency gaps, and accountability deficits in automated resource allocation decisions."
    },
    {
      "title": "Large Reasoning Models Are Not Thinking Straight: On the Unreliability of Thinking Trajectories",
      "url": "https://arxiv.org/html/2507.00711v1",
      "date": "2024-12-05",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Technical preprint revealing state-of-the-art reasoning models (LLaMA, Qwen) exhibit overthinking and discard correct solutions, fundamentally undermining reliance on AI reasoning trajectories for decision support."
    },
    {
      "title": "Bias and Inaccuracy as Key Concerns With Legal AI Tools",
      "url": "https://www.artificiallawyer.com/2024/12/03/bias-inaccuracy-key-concerns-with-legal-ai-tools-survey/",
      "date": "2024-12-03",
      "type": "adoption-metric",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Survey of 90 legal professionals finds 43% observe bias in AI tools and 37% fear unreliability; AI is widely adopted yet critical concerns about bias and accuracy persist as barriers to trustworthy deployment."
    },
    {
      "title": "Safe Use of AI-Based Diagnostic Decision Support Systems with Trust Calibration",
      "url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC11612524/",
      "date": "2024-11-27",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Quasi-experimental study documenting physician overreliance on AI diagnostic systems; proposes trust calibration to match AI reliability, revealing persistent pitfalls in clinical decision-support deployment."
    },
    {
      "title": "Are AI Chatbots Ready to Aid in Clinical Decision-Making?",
      "url": "https://www.aha.org/aha-center-health-innovation-market-scan/2024-10-15-are-ai-chatbots-ready-aid-clinical-decision-making",
      "date": "2024-10-15",
      "type": "news-coverage",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "AHA survey of 100+ physicians shows 76% use LLMs for clinical decisions despite reliability concerns; AI chatbots score high in reasoning benchmarks but exhibit incorrect reasoning, not yet ready for autonomous deployment."
    },
    {
      "title": "Challenges and Facilitation Approaches for the Participatory Design of AI-based Clinical Decision Support Systems",
      "url": "https://www.researchprotocols.org/2024/1/e58185/",
      "date": "2024-09-05",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "Scoping review protocol on barriers to AI-based clinical DSS implementation, identifying acceptance and feasibility challenges that participatory design methods seek to address in healthcare contexts."
    },
    {
      "title": "Three Challenges for AI-Assisted Decision-Making",
      "url": "https://pubmed.ncbi.nlm.nih.gov/37439761/",
      "date": "2024-09-04",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "Peer-reviewed analysis from UC Irvine identifying three key barriers to effective human-AI decision-making: complementarity conditions, human mental models of AI, and design choice impacts on cognitive load."
    },
    {
      "title": "The risks and inefficacies of AI systems in military targeting support",
      "url": "https://blogs.icrc.org/law-and-policy/2024/09/04/the-risks-and-inefficacies-of-ai-systems-in-military-targeting-support/",
      "date": "2024-09-04",
      "type": "opinion",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "ICRC critical assessment documenting permanent risks in AI decision-support systems for high-stakes decisions: biases, hallucinations, brittleness, and inability to ensure compliance with legal requirements."
    },
    {
      "title": "On the ROI of AI Ethics and Governance Investments",
      "url": "https://cmr.berkeley.edu/2024/07/on-the-roi-of-ai-ethics-and-governance-investments-from-loss-aversion-to-value-generation/",
      "date": "2024-07-29",
      "type": "industry-report",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "Strategic analysis from Berkeley/IBM researchers on justifying AI governance investments, noting companies with Gen AI guardrails may be 27% more likely to achieve higher revenue despite ethical and compliance costs."
    },
    {
      "title": "Research Firm Doubles Down on AI Disillusionment",
      "url": "https://virtualizationreview.com/Articles/2024/07/29/ai-disillusionment.aspx",
      "date": "2024-07-29",
      "type": "industry-report",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "Gartner analyst forecast predicting 30% of generative AI projects will be abandoned post-proof-of-concept by end of 2025 due to poor data quality, inadequate risk controls, and unclear business value."
    },
    {
      "title": "AI-Assisted Decision-Making in Long-Term Care - JMIR Nursing",
      "url": "https://nursing.jmir.org/2024/1/e55962/",
      "date": "2024-07-25",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "Qualitative study of 24 Dutch care professionals on prerequisites for responsible AI-DSS deployment, identifying seven critical requirements including bias mitigation, human-centric learning, and incremental trust-building."
    },
    {
      "title": "Nearly half of business AI projects abandoned midway, study finds",
      "url": "https://www.jpost.com/business-and-innovation/article-807391",
      "date": "2024-06-24",
      "type": "adoption-metric",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "DLA Piper survey of 600 executives found 48% of AI projects paused or rolled back due to privacy, regulatory, and integration challenges, revealing high deployment failure rates in practice."
    },
    {
      "title": "Does AI help humans make better decisions? - Harvard Gazette",
      "url": "https://news.harvard.edu/gazette/story/2024/06/does-ai-help-humans-make-better-decisions-artificial-intelligence-law/",
      "date": "2024-06-14",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "RCT in Wisconsin courts found AI recommendations failed to improve judicial bail decisions; judges rejected AI advice 30%+ of time due to inferior accuracy, demonstrating limitations in AI decision-support."
    },
    {
      "title": "Justitia ex machina: The impact of an AI system on legal decision-making and discretionary authority",
      "url": "https://research.tue.nl/en/publications/justitia-ex-machina-the-impact-of-an-ai-system-on-legal-decision-",
      "date": "2024-06-04",
      "type": "case-study",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "Peer-reviewed case study of AI decision-support deployed in Dutch court for traffic appeals, showing real-world impact on legal decisions with tensions in adoption and discretionary authority."
    },
    {
      "title": "Explainability does not mitigate the negative impact of AI bias in hiring decisions",
      "url": "https://pubmed.ncbi.nlm.nih.gov/38679619/",
      "date": "2024-04-28",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "Experimental study with 1,403 participants shows incorrect AI advice harms decisions due to overreliance; explainability does not reduce this negative effect, revealing critical human-AI interaction risks."
    },
    {
      "title": "No Thoughts Just AI: Biased LLM Recommendations Limit Human Agency in Consequential Decision-Making",
      "url": "https://arxiv.org/html/2509.04404v1",
      "date": "2024-04-28",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "Large-scale empirical study with 528 participants shows AI bias in hiring leads to up to 90% alignment with biased recommendations, demonstrating AI's capacity to reduce human agency in decision-making."
    },
    {
      "title": "The good, bad and ugly of bias in AI | CFA Institute",
      "url": "https://www.cfainstitute.org/insights/articles/good-bad-and-ugly-of-bias-in-ai",
      "date": "2024-04-25",
      "type": "industry-report",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "Industry analysis from CFA Institute on AI bias in investment decision-support, proposing frameworks for responsible AI with focus on fairness, accountability, and human oversight in financial decisions."
    },
    {
      "title": "Organizations Bullish on AI Adoption Despite Yearly Losses",
      "url": "https://tdwi.org/articles/2024/03/20/report-organizations-bullish-on-ai-adoption.aspx",
      "date": "2024-03-20",
      "type": "adoption-metric",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Nearly nine in ten organizations use AI/ML for autonomous decision-making despite 6% average revenue loss from model failures; reveals adoption breadth but persistent trust and data quality challenges."
    },
    {
      "title": "AI Scientists Fail Without Strong Implementation Capability",
      "url": "https://arxiv.org/html/2506.01372v1",
      "date": "2024-03-03",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Study of AI Scientist systems shows Claude 3.5 scores only 1.8% on PaperBench complex engineering tasks; identifies implementation execution as the 'fundamental bottleneck' in AI reasoning systems."
    },
    {
      "title": "Artificial Intelligence Is Not Immune to Sociopolitical Failures",
      "url": "https://www.tc.columbia.edu/articles/2024/february/artificial-intelligence-is-not-immune-to-sociopolitical-failures/",
      "date": "2024-02-13",
      "type": "opinion",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Analysis of AI bias and failures in educational decision systems; shows how AI-powered decision support perpetuates historical inequities despite claims of objectivity."
    },
    {
      "title": "Apple Study Reveals Critical Flaws in AI's Logical Reasoning Abilities",
      "url": "https://theoutpost.ai/news-story/apple-study-reveals-critical-flaws-in-ai-s-logical-reasoning-abilities-6840/",
      "date": "2024-02-10",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Apple research identifies critical reasoning failures in LLMs; proposes GSM-Symbolic benchmark to measure logical reasoning—exposing fundamental limitations in AI decision support capability."
    },
    {
      "title": "The IQ Test That AI Can't Pass",
      "url": "https://www.johndcook.com/blog/2024/01/16/the-iq-test-ai-cant-pass/",
      "date": "2024-01-16",
      "type": "opinion",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "AI scores only 30% on ARC benchmark for novel reasoning tasks, far below human performance; demonstrates AI cannot yet handle decision-making requiring genuine reasoning on unfamiliar problem types."
    },
    {
      "title": "A Survey of Reasoning with Foundation Models",
      "url": "https://arxiv.org/abs/2312.11562",
      "date": "2023-12-17",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2023-H2",
      "explanation": "Comprehensive 160-page survey (750+ references) synthesizing state of AI reasoning capabilities in foundation models—the technical substrate for decision support systems."
    },
    {
      "title": "Artificial intelligence and clinical decision support: clinicians' perspectives on trust, trustworthiness, and liability",
      "url": "https://pubmed.ncbi.nlm.nih.gov/37218368/",
      "date": "2023-11-27",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2023-H2",
      "explanation": "Clinicians' primary adoption barriers for AI-CDSSs are accuracy concerns and legal liability; trust framework analysis shows unresolved accountability gap limiting real-world deployment."
    },
    {
      "title": "The need to strengthen the evaluation of the impact of Artificial Intelligence-based decision support systems on healthcare provision",
      "url": "https://pubmed.ncbi.nlm.nih.gov/37579545/",
      "date": "2023-10-25",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2023-H2",
      "explanation": "Despite renewed interest in AI-CDS, lack of empirical evidence on effectiveness; warns of sub-optimal implementation risks and calls for rigorous evaluation frameworks."
    },
    {
      "title": "Psychological assessment of AI-based decision support systems: tool development and expected benefits",
      "url": "https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2023.1249322/full",
      "date": "2023-09-25",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2023-H2",
      "explanation": "Validated PAAI questionnaire assessing AI-DSSs from human-centered perspective with N=223 and N=471 studies, demonstrating design-psychology link to psychological load reduction and performance."
    },
    {
      "title": "Award-Winning Research Highlights Challenges and Opportunities for More Reliable Human-AI Decision-Making",
      "url": "https://www.mccormick.northwestern.edu/computer-science/news-events/news/articles/2023/award-winning-research-highlights-challenges-and-opportunities-for-more-reliable-human-ai-decision-making.html",
      "date": "2023-08-17",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2023-H2",
      "explanation": "CMU FAccT award-winning research on threats to validity and reliability in human-AI decision-making; identifies methods to address core reliability challenges in this domain."
    },
    {
      "title": "When A.I. (Artificial Intelligence) Fails - GRC 20/20 Research, LLC",
      "url": "https://grc2020.com/2023/07/31/when-a-i-artificial-intelligence-fails/",
      "date": "2023-07-31",
      "type": "industry-report",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2023-H2",
      "explanation": "AI decision support risks include dynamic environment brittleness, over-reliance, bias, and governance gaps; warns that 'complete trust in AI' is the most dangerous moral hazard."
    },
    {
      "title": "Feedback Effect in User Interaction with Intelligent Assistants: Delayed Engagement, Adaption and Drop-out",
      "url": "https://arxiv.org/abs/2303.10255",
      "date": "2023-03-17",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2023-H1",
      "explanation": "Observational analysis showing unhelpful AI responses cause users to reduce engagement and diversity over time, revealing adoption barriers in personal decision support."
    },
    {
      "title": "AI Overreliance Is a Problem. Are Explanations a Solution?",
      "url": "https://hai.stanford.edu/news/ai-overreliance-problem-are-explanations-solution",
      "date": "2023-03-13",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2023-H1",
      "explanation": "Stanford experimental study showing simpler explanations and higher decision stakes reduce overreliance on AI, with quantified effect sizes on decision quality."
    },
    {
      "title": "How Do Users Feel When They Use Artificial Intelligence for Decision-Making",
      "url": "https://ideas.repec.org/a/spr/infosf/v25y2023i3d10.1007_s10796-022-10293-2.html",
      "date": "2023-02-02",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2023-H1",
      "explanation": "Social media sentiment analysis revealing users' primary concerns are risk and accountability in AI decision-making systems, identifying key adoption barriers."
    },
    {
      "title": "Data on human decision, feedback, and confidence during an artificial intelligence-assisted decision-making task",
      "url": "https://pubmed.ncbi.nlm.nih.gov/36691561/",
      "date": "2023-01-09",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2023-H1",
      "explanation": "Empirical dataset from MIT/CMU/Berkeley study of 100 participants using AI in chess puzzles, capturing human-AI collaboration dynamics and confidence calibration in decision tasks."
    },
    {
      "title": "How Industries Evolve: Why AI Adoption Remains Low",
      "url": "https://peter.evans-greenwood.com/2022/12/21/why-hasnt-ai-delivered-on-its-promise/",
      "date": "2022-12-21",
      "type": "opinion",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2022-H2",
      "explanation": "Deloitte analysis citing Gartner prediction that ~85% of AI projects fail to move from prototype to production, highlighting critical adoption barriers in enterprise AI decision-making."
    },
    {
      "title": "Towards Benchmarking Temporal Reasoning in Large Language Models",
      "url": "https://ar5iv.labs.arxiv.org/html/2306.08952",
      "date": "2022-12-16",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2022-H2",
      "explanation": "Introduces TempReason benchmark showing LLMs perform poorly on temporal reasoning and exhibit bias toward contemporary years (2000-2020), exposing decision-support limitations."
    },
    {
      "title": "AI in Investment Decisions: Success and Catastrophic Failures",
      "url": "https://fortune.com/2022/12/02/ai-artificial-intelligence-investment-decisions-uniqueness-risks-benefits/amp",
      "date": "2022-12-02",
      "type": "news-coverage",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2022-H2",
      "explanation": "John Deere achieved precision agriculture transformation via AI, but Zillow lost $500M when AI models failed to adapt to housing market changes, demonstrating both promise and severe deployment risks."
    },
    {
      "title": "Human-Machine Joint Decision-Making: A Regulatory Framework Analysis",
      "url": "https://clsbluesky.law.columbia.edu/2022/12/02/debevoise-plimpton-discusses-the-myth-of-artificial-intelligence-errors/",
      "date": "2022-12-02",
      "type": "opinion",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2022-H2",
      "explanation": "Legal analysis proposing framework for human-in-the-loop decision-making, challenging assumption that human judgment always overrides AI, with healthcare and hiring examples."
    },
    {
      "title": "AI Decision Support in Public Transport: Technical and Socio-technical Barriers",
      "url": "https://ouci.dntb.gov.ua/en/works/9ZwYkJ87/",
      "date": "2022-12-01",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2022-H2",
      "explanation": "Describes hybrid human-AI decision support in transport operations; identifies critical barriers including lack of formalizability of human decision processes and system acceptance challenges."
    },
    {
      "title": "Consistent Reasoning Paradox: Fallibility as a Feature of Intelligent Systems",
      "url": "https://arxiv.org/html/2408.02357v1",
      "date": "2022-11-25",
      "type": "research-paper",
      "added": "2026-03-19",
      "superseded_by": null,
      "window": "2022-H2",
      "explanation": "Theoretical paper asserting that AI mimicking human intelligence must be fallible and require 'I don't know' capability; identifies fundamental limitations in current AI reasoning systems."
    }
  ],
  "tierHistory": [
    {
      "tier": "research",
      "from": "2022-11-01",
      "to": "2024-04-01"
    },
    {
      "tier": "bleeding-edge",
      "from": "2024-04-01",
      "to": null
    }
  ],
  "trendHistory": [
    {
      "trend": "steady",
      "blockerType": null,
      "from": "2026-09-26",
      "to": null
    }
  ],
  "description": "AI that helps individuals structure decisions, evaluate options, and apply reasoning frameworks to complex choices. Includes decision matrix generation and pro/con analysis; distinct from feature prioritisation which applies frameworks to product decisions rather than general personal choices.",
  "overview": "AI-assisted decision support has proven it can work in narrow, data-rich contexts, but the barrier to reliable deployment is no longer capability—it is organizational readiness, judgment preservation, and accountability. Cox Communications achieved 7x ROI on multi-agent sales decisioning; Kai eliminated 99.5% false positives in security triage; Anthropic handles 95% of internal analytics queries through Claude. Yet across 2,400+ enterprise AI initiatives, failure rates sit at 80%, and 90% of executives surveyed report no discernible productivity impact. The barrier is not technical: it is organizational ability to articulate decision processes, design consistent human-AI workflows, and maintain human judgment throughout. Even when AI reasoning is demonstrably superior, humans reject it when it contradicts their prior answer; higher confidence in AI correlates with lower critical thinking; and decision briefings can be factually correct yet unsafe if they suppress uncertainty, misstate causality, or omit auditable reasoning. Simultaneously, fundamental LLM reasoning flaws constrain what can be delegated: models collapse under cognitive load (Stroop: 91%→1% accuracy), fail at sequential decision-fork reasoning (59.7% accuracy on frontier models), and commit 6-12% safety violations on critical tasks. These are architectural, not tunable. The field has shifted from \"can we build it\" to \"how do we scale safe, accountable decision support whilst preserving human judgment.\"",
  "currentLandscape": "Credible deployments in 2026 share two characteristics: tightly scoped problems with rich structured data, and explicit governance frameworks. Cox Communications achieved 7x first-year ROI on sales decisioning; Kai eliminated 99.5% false positives across 2.5M security findings; Xylem onboarded 15,000 employees with 4,000 daily active users and estimated $70M revenue opportunity and $25M savings; One New Zealand deployed 50+ decision-support agents, reducing audit-planning time 60% and risk-matrix preparation from two days to under half a day. California government scaled Claude across state agencies for policy deliberation, Medicaid workflows and cyber triage. Yet a large pragmatic RCT (Nature Medicine, Kenya, 9,600 patients) found AI-assisted clinical decision support improved documentation but failed to reduce treatment failure. Professional services adoption reached 40% organisation-wide (up from 22%), but 91% report value shortfall, 90% demand explainability, and only 18% track ROI. Governance is the dominant failure mode: 87% of enterprises delayed deployments due to data governance risks, and 74% rolled back agents due to governance failure (data leakage 30.7%, hallucination 20.8%, auditability 16.8%). Technical debt accumulates rapidly—organizations spawn ungoverned agent variants without unified ownership of prompts, data sources or evaluation logic, creating untraceable errors. A September 2026 high-stakes failure: a US military intelligence analyst's AI synthesized classified and open-source intelligence to fabricate a nuclear-weapons claim, which the chatbot formatted as an authoritative report circulating across command channels; armed personnel prepared to board a ship before the error was caught. Governance frameworks are maturing to address this: decision briefings require five accountability dimensions (citation-to-claim entailment, causal-language discipline, uncertainty preservation, action appropriateness, human accountability); decision matrices allocate rights by ambiguity and risk; practitioner frameworks distinguish pattern-work delegable to AI from value-work requiring judgment, with safety architecture embedded. Yet adoption remains bottlenecked: 72% of CEOs expect AI to support human-directed decisions (not autonomous), identifying human governance as the design requirement; 55% of professionals cite lack of structured workflows as adoption barrier; and 82% report individual AI-readiness whilst only 15% of enterprises operationalize capability. A 3,700-participant study found AI recommendations amplified accuracy when aligned (64%→87%) but crashed below baseline when misaligned (34%), establishing that decision-support only functions reliably in bounded domains with explicit value alignment. Fundamental LLM constraints remain: models collapse under cognitive load (Stroop: 91%→1% accuracy), fail at decision-fork reasoning (frontier models 59.7% accuracy; larger reasoning budgets provide no improvement), and commit 6-12% safety violations on critical tasks. Legal frameworks are crystallizing: decision authority cannot be delegated to AI; executives using public AI tools informally breach duty of care; and mandatory documentation of reasoning review and source verification is required. Australia's Fair Work Commission condemned AI decision-support failures (40% of cases involved AI), mandating disclosure requirements, signalling regulatory consolidation around accountability and judgment preservation.",
  "history": "- **2022-H2:** First identified research surge in AI reasoning benchmarks (temporal, step-by-step, knowledge-graph) and human-in-the-loop frameworks; major failures in practice (Zillow, model brittleness); ~85% enterprise project failure rate documented; theoretical work on reasoning fallibility and the need for AI to express uncertainty.\n- **2023-H1:** Research focus shifted to human-AI interaction challenges: overreliance on AI suggestions despite explanations, user dropout when AI feedback is unhelpful, and widespread concerns about accountability and risk. Evidence of adoption barriers in personal decision support remained dominant; no large-scale personal reasoning framework deployments documented.\n- **2023-H2:** Research concentrated on three critical areas: (1) human-centered design frameworks for DSSs (PAAI questionnaire with 700+ participant validation), (2) trust and accountability barriers blocking clinical deployment (liability concerns, accuracy standards), and (3) underlying AI reasoning capabilities (foundation model survey). Evaluation gaps documented—despite renewed interest, empirical evidence on AI-CDS effectiveness remained scarce. Governance and risk analysis highlighted over-reliance, bias, and dynamic environment brittleness as core failure modes.\n- **2024-Q1:** Adoption accelerated dramatically: ~90% of enterprises deployed AI for autonomous decision-making. Simultaneously, fundamental technical limitations became clearer—Apple research confirmed critical reasoning flaws in LLMs (GSM-Symbolic benchmark), and AI systems scored only 30% on novel reasoning tasks (ARC). Bias in operational AI decision systems documented in education. Revenue impact quantified: 6% average annual loss from underperforming models. The execution gap widened: implementations failed at generalization, bias mitigation, and reliability despite widespread organizational trust. Empirical evidence on decision support effectiveness remained sparse, leaving large-scale deployments without measured impact validation.\n- **2024-Q2:** Real-world deployment evidence emerged, revealing persistent failures despite adoption breadth. A Dutch court case study showed AI decision-support in legal proceedings, while a Harvard-led RCT in Wisconsin courts found AI recommendations failed to improve bail decisions and judges rejected them 30%+ of the time. Experimental research with 1,403 participants confirmed overreliance remains endemic despite explainability efforts—workers align with biased AI recommendations up to 90% in hiring contexts. A 600-executive survey found 48% of AI projects paused or rolled back due to privacy, regulatory, and integration challenges. The landscape shifted from \"can we build AI decision systems\" to \"why are deployed systems failing to improve decisions\"—adoption at scale masked persistent technical and organizational gaps.\n- **2024-Q3:** Research clarified three persistent barriers to effective human-AI decision-making: achieving complementarity, managing human mental models, and design choices that prevent cognitive overload. Healthcare case studies identified prerequisite frameworks for responsible AI-DSS (bias mitigation, human-centric learning loops, incremental trust-building). Critical assessments documented permanent risks in high-stakes decision contexts (military targeting, legal proceedings) due to hallucinations, brittleness, and inability to ensure regulatory compliance. Analyst forecasts predicted 30% project abandonment post-proof-of-concept by end of 2025, with organizations struggling to realize value despite major investments. The field continued to reconcile widespread enterprise adoption with persistent deployment failures, unresolved bias risks, and absence of clear impact metrics on decision quality.\n- **2024-Q4:** Critical research published on technical and organizational barriers to reliable AI reasoning. New findings revealed overreliance persists in clinical decision-making despite trust calibration efforts, with physicians exhibiting diagnostic errors from AI misalignment. Healthcare professionals identified systemic ethical concerns (bias, transparency gaps, accountability deficits) in AI-CDSS deployments. Fundamental research showed state-of-the-art reasoning models exhibit overthinking and discard correct reasoning paths, undermining the assumption that larger models improve decision support. Legal sector adoption continued despite significant concerns: 43% of legal professionals observed bias, 37% feared unreliability. Clinical adoption surveys showed 76% of physicians now use LLMs for decisions yet 97% vet outputs, indicating cautious rather than confident deployment. By year-end 2024, the field had reached consensus that the core challenge is not building reasoning systems but deploying them safely and measurably—technical limitations in AI reasoning were well-documented, but practical implementation remained the bottleneck. Organizations continued investing despite unresolved risks, suggesting adoption momentum has decoupled from evidence of effectiveness.\n- **2025-Q1:** Enterprise adoption continued but real-world reliability challenges intensified. UK government abandoned multiple welfare-system AI pilots (A-cubed, Aigent) due to scalability and reliability concerns—explicit signal of deployment failure in public sector decision support. Healthcare outcomes improved in targeted deployments: UK Health Security Agency achieved 90% accuracy in TB screening with 85% reduction in manual review workload. Research on adoption barriers showed 450 physicians in China identified multiple adoption pathways depending on hospital type and organizational context. Critical research revealed that human oversight alone is insufficient to prevent discrimination: EU study found human decision-makers equally likely to follow biased AI recommendations regardless of fairness-algorithm design. Fundamental reasoning limitations persisted: AI reasoning models continued exhibiting data bias, lack of common sense, and transparency failures that undermine high-stakes decision-making. The gap between pilot success and production scaling widened: isolated cases showed operational gains, but public sector abandonment and persistent bias findings suggested the field remained pre-scale.\n- **2025-Q2:** Evidence revealed critical implementation gaps despite continued investment. Dermatology study (223 physicians) found AI support yielded only 1% accuracy improvement with low reliance (10%), indicating adoption barriers persist even in favorable clinical contexts. Defense deployments (Project Maven, UK autonomous targeting, Iron Dome) demonstrated real-world AI-DSS use but in high-stakes, tightly constrained settings. ChatGPT testing showed AI mirrors human decision-making biases including overconfidence and gambler's fallacy in half of scenarios, suggesting AI amplifies cognitive flaws rather than mitigating them. Expert Delphi consensus identified 34 critical implementation factors for healthcare AI-DSS, yet organizational capacity to execute remained limited. Industry analysis showed only 26% of companies have working AI products and 4% achieve significant ROI; Gartner predicted 40%+ project cancellations by 2027 due to unclear value and costs. Parallel evidence of high adoption breadth (93% of leaders report GenAI competitive benefits) masked low implementation depth and persistent execution challenges.\n- **2025-Q3:** Research clarified fundamental and persistent technical limitations in AI reasoning: models performed no better than humans on novel problems and replicated cognitive biases including overconfidence. MIT analysis of 300 deployments found 95% of AI pilots failed to deliver value, with vendor solutions succeeding ~67% versus internal builds 33%—exposing both adoption and execution challenges. Consumer trust surveys (YouGov, 10K respondents) showed 52% comfort with AI for daily personal decisions but only 39% for financial decisions; humans retained override preference in 55%+ of scenarios. New tools for bias detection (CMU AIR) and structured decision frameworks (MCDM-based ModelSelect with 50 case-study validation) promised incremental rigor improvements yet could not address fundamental reasoning limitations. Research documented that AI actively degraded decision quality: executives using generative AI made worse forecasts than without it, highlighting the risk of overconfidence in AI-enhanced reasoning. The field remained characterized by adoption momentum decoupled from evidence of effectiveness, with organizations continuing heavy investment despite quantified failures and persistent technical barriers.\n\n- **2025-Q4:** Deployment evidence revealed domain-specific outcomes: IBM achieved $4.5B productivity impact from agentic AI deployed to 270K employees; marketing decision-intelligence platforms reached 26-75% adoption with measurable ROI; UK Health Security Agency's AI-assisted TB screening achieved 90% accuracy. Yet critical limitations emerged across high-stakes domains: medical data gaps (EMR design flaws, not algorithmic limitations) constrained clinical decision-support impact; Indian judges warned of AI-fabricated legal judgments and hallucinations; government pilots stalled due to scaling and budget challenges; legal professionals documented persistent bias (43% observing bias, 37% fearing unreliability). Medical educators flagged overreliance risks: GenAI tools threaten critical thinking skill development and reinforce training data biases. Technical advances in reasoning (GPT-5.1 integration, causal AI frameworks) continued, yet ethics scholars debated justified use of black-box AI in high-stakes domains. By year-end, the field had consolidated around differentiation by domain: operational value in narrow contexts (marketing, logistics) versus persistent barriers and documented risks in broader organizational and high-stakes deployment scenarios. Adoption momentum remained decoupled from evidence of effectiveness, with organizations continuing investment despite quantified failures and unresolved deployment barriers.\n\n- **2026-Jan:** Enterprise transition to operationalization emphasized data governance and architectural foundations; 62% of enterprises planning evolution to AI decision intelligence amid persistent 70-85% project failure rates and 42% initiative abandonment in 2025. Causal AI emerged as next-frontier addressing 74% faithfulness gap in existing systems. Clinical research documented error reduction (78% decline in guideline violations) through hybrid frameworks, yet deployment barriers remained: only 12% of executives reported both cost and revenue benefits; physician studies highlighted that reasoning cues must target high-discretion tasks where AI can add genuine value.\n- **2026-Feb:** Multi-AI orchestration demonstrated operational feasibility (ARPIA 13-min data-to-strategy pipeline, causaLens enterprise deployments); professional services adoption jumped to 40% (2025: 22%), yet only 18% track ROI and 40% report policy confusion. Systematic LLM reasoning failures (Reversal Curse, Robustness Fragility, Working Memory Leaks) documented, undermining reliance on AI reasoning chains. Commercial AI-CDSS solutions still lack transparent training data and algorithm disclosure. Regulatory deadlines (CCPA Jan 2026, EU AI Act Aug 2026) drove governance platform launches. Across 2,400+ enterprise AI initiatives, 80.3% failed (33.8% abandoned, 28.4% deliver no value), with 95% GenAI pilots failing to reach production. Execution and governance remain bottlenecks, not capability.\n\n- **2026-Mar:** Fundamental research documented persistent reasoning failures: CRYSTAL benchmark shows models skip 50%+ of reasoning steps (58% accuracy but only 48% reasoning recovery); BrainBench reveals stochastic reasoning gaps (6-16pp consistency variance even in top models); Stanford taxonomy classifies failures as architectural rather than scale-addressable. Real-world failures documented: NZ courts ruled AI-hallucinated legal citations may amount to obstruction; Deloitte refunded AUD 440K for AI-generated errors. Governance frameworks consolidated: RegTech expert consensus establishes human accountability cannot be delegated; KPMG legal analysis requires mandatory documentation of decision review. Practical deployment barriers clarified: reasoning models show 5x cost premium with performance ceiling at medium-complexity tasks (above which accuracy collapses). Field consensus solidifying around decision-support constraints: execution challenges and governance requirements are primary blockers, not reasoning capability gaps.\n- **2026-Apr:** New empirical research confirmed architectural reasoning limits are scale-invariant: testing across 7 models (8B–235B parameters) showed all collapse at 20-30 parallel branches, and semantic variants of problems trigger 28-45% answer-flip rates — confirming brittleness is structural, not addressable by larger models. CMU testing of 14 leading LLMs (GPT-4, Claude 3, Gemini) found all fail simple logical contradiction detection, revealing that benchmark performance masks fundamental reasoning gaps. Harvard research added a new dimension: relational complexity causes accuracy to collapse when decisions require weighing multiple interacting factors simultaneously, directly constraining multi-factor analysis in healthcare and strategic contexts. Legal accountability frameworks tightened: Australian Federal Court (ASIC v Bekier) established that executives using public AI tools informally without verification breach their duty of care, reinforcing that decision accountability cannot be delegated to AI and mandating documentation of reasoning review. Latest evidence on personal decision support reveals critical tensions: targeted healthcare deployments show measurable value (RCT with 367 participants: 7.4-point satisfaction improvement, 50.7% vs 24.2% acceptance of AI recommendations), yet passive reliance on AI reasoning systematically erodes confidence in independent judgment and sense of authorship (behavioral study, 1,923 adults). Empirical adoption barriers remain severe: 9% of professionals trust AI for complex decisions despite 88% organizational adoption; 80% of workers reject enterprise AI tools (WalkMe, 3,750 professionals). Reasoning failure modes are now precisely characterized: hallucination rates span 22-94% across models (Stanford 2026 AI Index), with a documented \"Reliability Gap\" where capability scales 2-3x annually while reliability only 1.2-1.5x. Governance frameworks are consolidating: the SPEC framework achieves 89% accuracy vs 15% for unbounded RAG on incomplete-evidence scenarios by bounding AI confidence to evidential sufficiency; the CFA Institute articulates an epistemic anchoring principle that decision authority must remain in evidence-based human inquiry to avoid \"knowledge-collapse equilibrium.\"\n\n- **2026-May:** Research intensified focus on cognitive and systemic failure modes in AI-assisted decision-making. Wharton's Cognitive Reflection Test (1,300+ participants) demonstrated the core paradox: AI-correct advice improves accuracy 25pp, but AI-wrong advice degrades 15pp below baseline—worse than no AI. Overconfidence persists even when users know AI errs 50% of the time, indicating cognitive surrender rather than rational reliance. Metacognitive stability revealed as critical: empirical testing across 11 frontier models shows 8 collapse under adversarial pressure (30.2pp accuracy drops), with only Anthropic's Constitutional AI showing near-immunity—suggesting alignment-specific training is prerequisite for trustworthy reasoning, not achievable through standard RLHF. Enterprise adoption gaps widened: <20% of AI pilots reach production due to missing trust infrastructure (audit trails, explainability, liability frameworks); yet production deployments in bounded domains (insurance claims, pharma R&D) show 50%+ efficiency gains when governance frameworks are load-bearing. Real-world interaction data (5.5M instances) reveals expert-domain performance plateaued at 14-16% dissatisfaction despite scale; raw models generate confident false theories when given deliberately flawed premises. Most critical finding: brief AI assistance (10 minutes) systematically impairs independent problem-solving through cognitive offloading, reducing retention and analytical skepticism—unintended consequence suggesting tool design actively erodes reasoning autonomy. Governance literature consolidated: AI system abandonment driven primarily by organizational dynamics and resource constraints, not ethics concerns; clinical AI remains confined to pilots not due to model limitations but institutional capacity gaps. The field's consensus strengthened: decision-support reliability requires orchestration of governance, human oversight, design discipline, and alignment-specific training—capability alone is necessary but insufficient.\n\n- **2026-Late May:** Enterprise deployment momentum accelerated with defined governance patterns. SAP Sapphire 2026 unveiled 224 specialized Claude agents handling autonomous decisions across finance, HR, supply chain, and procurement for hundreds of thousands of enterprise customers globally—largest-scale production deployment of autonomous decision-making systems. CLR-voyance clinical reasoning system demonstrated production-scale maturity: 6+ months in hospital operation, 84.91% accuracy with physician-validated outcome rubrics, drafting thousands of inpatient notes—proving structured decision-support can reach institutional deployment. Yet critical limitations emerged: meta-analysis of 5 RCTs (12,657 participants) found AI clinical decision support produces only small, marginal improvements in diagnostic accuracy; pooled effect size narrowly above zero; strongest signal in radiology, weakest in complex reasoning domains. Governance failure patterns crystallized: survey of 650 enterprise leaders found 78% ran AI agent pilots but only 14% scaled to production; 90% of agent codebases fail EU AI Act compliance scans; governance architecture (audit trails, escalation paths, ownership documentation) emerged as primary blocker, not model performance. Personal effectiveness frameworks showed concrete signals: executive decision-support (persistent context, standing priorities, weekly accountability reviews) demonstrated production viability; structured reasoning patterns (Clarify-Then-Act, Plan-Then-Execute, Human-in-the-Loop gates) identified enterprise reliability requirements. Research confirmed psychological decision-making frameworks improve AI reasoning: TU Berlin study showed Recognition-Primed Decision-Making and Data-Frame Theory prompts boost healthcare advice accuracy 13% → 30%, establishing that human decision-making structures significantly enhance AI decision quality. Institutional barriers remained decisive: ICML paper documented that technical viability alone insufficient for scaling; decision-support systems require institutional alignment across approvals, oversight capacity, fiscal sustainability, and regulatory readiness. The field consolidated around a fundamental insight: decision-support reliability depends primarily on governance, human oversight, and institutional capacity—not model capability alone.\n- **2026-Jun:** Consensus solidified around organizational readiness as the binding constraint. Thomson Reuters survey (1,816 professionals): 91% report AI value shortfall despite 74% regular use; 90% demand reasoning that can be explained and defended. NBER study of ~6,000 executives: ~90% report zero productivity impact; PwC CEO survey: 56% experienced no revenue/cost improvement, only 12% significant benefits—a $2.5T \"validation gap\" between investment and outcomes. But concrete deployments prove bounded decision-support works: Cox Communications achieved 7x first-year ROI on multi-agent sales decisioning (15,000 employees); Kai eliminated 99.5% false positives in security triage (2.5M findings). Nature Medicine RCT (9,600 patients, 16 sites, Kenya): AI-assisted clinical decision support improved documentation quality but failed to reduce treatment failure (2.2% vs 2.0%, P=0.13)—emblematic outcome-documentation gap. The deepest barrier is organizational: only 15% of enterprises can leverage 82% of AI-ready individuals; 55% of professionals cite lack of structured workflows as adoption barrier. Harvard AI Institute director (HBR): \"core constraint is organizational ability to make decision-making processes explicit.\" Mechanistic research exposures: Stroop task (PNAS Nexus) shows models collapse under cognitive load (91%→1% accuracy); reasoning models spend MORE tokens on failed tasks (arXiv:2606.26502, Cohen's d 1.47–3.13), inverse of human behavior; frontier models commit 6-12% safety violations on critical tasks (Arachne, 81,000-person survey). Reasoning capability advances documented (OpenAI disproving 80-year-old math conjecture, TTE-Flash, Recursive Language Models), yet hybrid architectures combining LLM planners with deterministic execution engines prove necessary for reliable analytics (trade-off: reasoning depth vs. factual grounding). Albabtain's organizational design study: AI augments or substitutes judgment based on design choices (transparency, override friction, performance metrics)—augmentation requires low-friction overrides, not AI compliance metrics. 72% of CEOs expect AI to support human-directed decisions (not autonomous), with trust/security concerns (31%) and data quality (34%) cited as barriers. Government-scale adoption signals: California partnership provides Claude across all state agencies (policy deliberation, Medicaid workflows, cyber triage). The field's 2026 consensus: decision-support reliability requires governance, explicit workflow design, and human judgment integration—not model accuracy alone.\n- **2026-Jul:** The validation gap deepened with harder numbers. NBER (~6,000 executives) and PwC CEO survey together document 90% reporting zero productivity impact and 56% seeing no revenue or cost improvement — a $2.5T blind spot. Against that backdrop, two production deployments documented genuine bounded ROI: Cox Communications 7x first-year return on multi-agent sales decisioning and Kai's 99.5% false-positive elimination in security triage. Architectural research confirmed models cannot self-regulate reasoning effort: reasoning tokens scale inversely with task success (Cohen's d 1.47–3.13), and the Stroop task collapse (91%→1% accuracy under cognitive load) is now confirmed as architectural rather than tunable. Organizational design research (Albabtain) established that augmentation vs. substitution of human judgment is determined by governance choices — low-friction overrides and judgment-rewarding metrics — not model capability. Mid-July evidence added both operational tooling and governance boundaries: a three-country RCT of 249 physicians found AI improved clinical reasoning (Kenya +18%) when paired with proper human-AI design, while peer-reviewed ACL research found frontier reasoning models achieve under 25% instruction-following compliance during their own reasoning traces — a control gap limiting governed deployment. Structured decision frameworks moved from instructional content toward operational tooling: an open-source Claude Code plugin implemented multi-agent debate to surface reasoning blind spots, and a 39-framework cognitive-technique catalog shipped via the Claude Code marketplace. Legal commentary reinforced that decision accountability cannot be delegated to AI, and a 1,000-dealmaker survey found 62% consider human-only decisions indefensible yet only 22% delegate final decisions to AI — underscoring the governance boundary.\n- **2026-Aug:** New evidence sharpened the accountability-capability gap: a real-world benchmark found frontier reasoning models fail to beat market baseline on decision-making under uncertainty, chain-of-thought explanations were shown to systematically misrepresent actual model computation, and a peer-reviewed study found AI advice cut human accuracy from 27% to 9% while nearly tripling confidence — while an HBR field study documented frontline workers defending AI decisions they neither created nor understood. A 758-consultant BCG field experiment quantified the capability boundary precisely (+12.2% productivity/+40% quality inside the AI's competence frontier vs. a 19pp correctness drop outside it), and multi-source analysis (S&P, Gartner, MIT, BCG, IBM) confirmed 42% of companies abandon AI initiatives due to organizational rather than technical barriers. Mid-to-late August deployments revealed both operational maturity and architectural constraints: ambient clinical decision-support reached 1,744 U.S. hospitals with measurable burnout reduction (51.9%→38.8%) yet only 10% have formal governance, while Maven demonstrates 5,000+ daily targeting decisions in military operations constrained by organizational diffusion barriers. Research confirmed three binding failure modes—position bias renders decisions order-sensitive across all tested models; interface design that helps experts actively misleads novices in explainability; constraint-preservation failures (Gemini 30+ cases) show models violate explicit logical rules—indicating architectural limits rather than trainable gaps. Widespread personal adoption (55% of U.S. workers report AI time-savings, yet 70% use AI for high-stakes decisions with 50% reporting overreliance) paired with documented harms (5M+ UK citizens reporting financial/health losses) establishes adoption outpacing governance maturity as the binding constraint on tier advancement. A US Census Bureau survey independently corroborated the adoption baseline (55% weekly worker use, 31% reporting 1-2 hour time savings), while a 9-LLM study confirmed AI decision quality is sensitive to option-presentation order, and 81% of physicians now use AI professionally even as FDA narrowed oversight and shifted liability onto clinicians—reinforcing that governance gaps, not model capability, remain the binding constraint. Late-August evidence added labor-tribunal and framework signals: Australia's Fair Work Commission condemned a dismissed worker's reliance on AI as a \"quasi-legal advisor\" and will mandate AI-use disclosure from October 2026; a University of Bayreuth experiment (n=3,700) found AI financial advice raised decision accuracy from 64% to 87% when aligned but crashed it to 34% when misaligned, with conflict-of-interest disclosure offering little protection; and new drafting-versus-deciding frameworks (a GA Claude Skill for structured reflection, and a Salesforce-consulting staged-autonomy model achieving 50%+ case deflection) continued formalizing where AI assistance should stop short of the final decision.\n- **2026-Sep:** Judgment-erosion and governance-rollback evidence converged. BCG's C-suite survey (70 leaders) found 50% already observe judgment/problem-solving skill decay and 90% report employees have stopped checking AI work; a Harvard field experiment (228 senior reviewers) documented automation bias overwhelming expert judgment, with explanations paradoxically worsening rather than correcting the bias. Rollback evidence hardened: a meta-analysis of MIT/RAND/S&P Global/Gartner studies found 95% of AI pilots report zero ROI and 42% abandon before production; AvePoint's survey of 3,235 leaders found 87% delayed deployments over governance risk (40.7% canceled GenAI rollouts, up from 31.7% in 2025); a separate 2,527-enterprise survey found 74% rolled back deployed AI agents, with \"mature guardrails\" firms rolling back at an even higher 81%. PACT's benchmark of 22 models across 12 regulated domains found compliance violations escalate from 6-10% at baseline to 65% under workplace pressure, and a Polish Supreme Administrative Court case documented AI hallucination (fabricated propositions attached to real case numbers) creating professional liability. Against this backdrop, Meta's CORAL framework demonstrated closed-loop decision-optimization at production scale across two billion-user platforms, and Salesforce embedded 37 prebuilt sales decision-support skills inside Claude for pilot customers, routing actions through business-rule governance — showing capable, governed systems remain the exception rather than the norm. Further evidence pointed the same way: Taste-Bench found frontier models answer only 59.7% of decision-fork questions correctly regardless of reasoning budget, and a SOCPAC case saw an AI-fabricated nuclear claim come within minutes of triggering a ship boarding. Scale deployments continued (Xylem 15,000 employees; One New Zealand's 50+ agents), alongside a five-dimension accountability framework for AI briefings and a Microsoft review of 120+ overreliance papers.",
  "historyEntries": [
    {
      "period": "2022-H2",
      "text": "First identified research surge in AI reasoning benchmarks (temporal, step-by-step, knowledge-graph) and human-in-the-loop frameworks; major failures in practice (Zillow, model brittleness); ~85% enterprise project failure rate documented; theoretical work on reasoning fallibility and the need for AI to express uncertainty."
    },
    {
      "period": "2023-H1",
      "text": "Research focus shifted to human-AI interaction challenges: overreliance on AI suggestions despite explanations, user dropout when AI feedback is unhelpful, and widespread concerns about accountability and risk. Evidence of adoption barriers in personal decision support remained dominant; no large-scale personal reasoning framework deployments documented."
    },
    {
      "period": "2023-H2",
      "text": "Research concentrated on three critical areas: (1) human-centered design frameworks for DSSs (PAAI questionnaire with 700+ participant validation), (2) trust and accountability barriers blocking clinical deployment (liability concerns, accuracy standards), and (3) underlying AI reasoning capabilities (foundation model survey). Evaluation gaps documented—despite renewed interest, empirical evidence on AI-CDS effectiveness remained scarce. Governance and risk analysis highlighted over-reliance, bias, and dynamic environment brittleness as core failure modes."
    },
    {
      "period": "2024-Q1",
      "text": "Adoption accelerated dramatically: ~90% of enterprises deployed AI for autonomous decision-making. Simultaneously, fundamental technical limitations became clearer—Apple research confirmed critical reasoning flaws in LLMs (GSM-Symbolic benchmark), and AI systems scored only 30% on novel reasoning tasks (ARC). Bias in operational AI decision systems documented in education. Revenue impact quantified: 6% average annual loss from underperforming models. The execution gap widened: implementations failed at generalization, bias mitigation, and reliability despite widespread organizational trust. Empirical evidence on decision support effectiveness remained sparse, leaving large-scale deployments without measured impact validation."
    },
    {
      "period": "2024-Q2",
      "text": "Real-world deployment evidence emerged, revealing persistent failures despite adoption breadth. A Dutch court case study showed AI decision-support in legal proceedings, while a Harvard-led RCT in Wisconsin courts found AI recommendations failed to improve bail decisions and judges rejected them 30%+ of the time. Experimental research with 1,403 participants confirmed overreliance remains endemic despite explainability efforts—workers align with biased AI recommendations up to 90% in hiring contexts. A 600-executive survey found 48% of AI projects paused or rolled back due to privacy, regulatory, and integration challenges. The landscape shifted from \"can we build AI decision systems\" to \"why are deployed systems failing to improve decisions\"—adoption at scale masked persistent technical and organizational gaps."
    },
    {
      "period": "2024-Q3",
      "text": "Research clarified three persistent barriers to effective human-AI decision-making: achieving complementarity, managing human mental models, and design choices that prevent cognitive overload. Healthcare case studies identified prerequisite frameworks for responsible AI-DSS (bias mitigation, human-centric learning loops, incremental trust-building). Critical assessments documented permanent risks in high-stakes decision contexts (military targeting, legal proceedings) due to hallucinations, brittleness, and inability to ensure regulatory compliance. Analyst forecasts predicted 30% project abandonment post-proof-of-concept by end of 2025, with organizations struggling to realize value despite major investments. The field continued to reconcile widespread enterprise adoption with persistent deployment failures, unresolved bias risks, and absence of clear impact metrics on decision quality."
    },
    {
      "period": "2024-Q4",
      "text": "Critical research published on technical and organizational barriers to reliable AI reasoning. New findings revealed overreliance persists in clinical decision-making despite trust calibration efforts, with physicians exhibiting diagnostic errors from AI misalignment. Healthcare professionals identified systemic ethical concerns (bias, transparency gaps, accountability deficits) in AI-CDSS deployments. Fundamental research showed state-of-the-art reasoning models exhibit overthinking and discard correct reasoning paths, undermining the assumption that larger models improve decision support. Legal sector adoption continued despite significant concerns: 43% of legal professionals observed bias, 37% feared unreliability. Clinical adoption surveys showed 76% of physicians now use LLMs for decisions yet 97% vet outputs, indicating cautious rather than confident deployment. By year-end 2024, the field had reached consensus that the core challenge is not building reasoning systems but deploying them safely and measurably—technical limitations in AI reasoning were well-documented, but practical implementation remained the bottleneck. Organizations continued investing despite unresolved risks, suggesting adoption momentum has decoupled from evidence of effectiveness."
    },
    {
      "period": "2025-Q1",
      "text": "Enterprise adoption continued but real-world reliability challenges intensified. UK government abandoned multiple welfare-system AI pilots (A-cubed, Aigent) due to scalability and reliability concerns—explicit signal of deployment failure in public sector decision support. Healthcare outcomes improved in targeted deployments: UK Health Security Agency achieved 90% accuracy in TB screening with 85% reduction in manual review workload. Research on adoption barriers showed 450 physicians in China identified multiple adoption pathways depending on hospital type and organizational context. Critical research revealed that human oversight alone is insufficient to prevent discrimination: EU study found human decision-makers equally likely to follow biased AI recommendations regardless of fairness-algorithm design. Fundamental reasoning limitations persisted: AI reasoning models continued exhibiting data bias, lack of common sense, and transparency failures that undermine high-stakes decision-making. The gap between pilot success and production scaling widened: isolated cases showed operational gains, but public sector abandonment and persistent bias findings suggested the field remained pre-scale."
    },
    {
      "period": "2025-Q2",
      "text": "Evidence revealed critical implementation gaps despite continued investment. Dermatology study (223 physicians) found AI support yielded only 1% accuracy improvement with low reliance (10%), indicating adoption barriers persist even in favorable clinical contexts. Defense deployments (Project Maven, UK autonomous targeting, Iron Dome) demonstrated real-world AI-DSS use but in high-stakes, tightly constrained settings. ChatGPT testing showed AI mirrors human decision-making biases including overconfidence and gambler's fallacy in half of scenarios, suggesting AI amplifies cognitive flaws rather than mitigating them. Expert Delphi consensus identified 34 critical implementation factors for healthcare AI-DSS, yet organizational capacity to execute remained limited. Industry analysis showed only 26% of companies have working AI products and 4% achieve significant ROI; Gartner predicted 40%+ project cancellations by 2027 due to unclear value and costs. Parallel evidence of high adoption breadth (93% of leaders report GenAI competitive benefits) masked low implementation depth and persistent execution challenges."
    },
    {
      "period": "2025-Q3",
      "text": "Research clarified fundamental and persistent technical limitations in AI reasoning: models performed no better than humans on novel problems and replicated cognitive biases including overconfidence. MIT analysis of 300 deployments found 95% of AI pilots failed to deliver value, with vendor solutions succeeding ~67% versus internal builds 33%—exposing both adoption and execution challenges. Consumer trust surveys (YouGov, 10K respondents) showed 52% comfort with AI for daily personal decisions but only 39% for financial decisions; humans retained override preference in 55%+ of scenarios. New tools for bias detection (CMU AIR) and structured decision frameworks (MCDM-based ModelSelect with 50 case-study validation) promised incremental rigor improvements yet could not address fundamental reasoning limitations. Research documented that AI actively degraded decision quality: executives using generative AI made worse forecasts than without it, highlighting the risk of overconfidence in AI-enhanced reasoning. The field remained characterized by adoption momentum decoupled from evidence of effectiveness, with organizations continuing heavy investment despite quantified failures and persistent technical barriers."
    },
    {
      "period": "2025-Q4",
      "text": "Deployment evidence revealed domain-specific outcomes: IBM achieved $4.5B productivity impact from agentic AI deployed to 270K employees; marketing decision-intelligence platforms reached 26-75% adoption with measurable ROI; UK Health Security Agency's AI-assisted TB screening achieved 90% accuracy. Yet critical limitations emerged across high-stakes domains: medical data gaps (EMR design flaws, not algorithmic limitations) constrained clinical decision-support impact; Indian judges warned of AI-fabricated legal judgments and hallucinations; government pilots stalled due to scaling and budget challenges; legal professionals documented persistent bias (43% observing bias, 37% fearing unreliability). Medical educators flagged overreliance risks: GenAI tools threaten critical thinking skill development and reinforce training data biases. Technical advances in reasoning (GPT-5.1 integration, causal AI frameworks) continued, yet ethics scholars debated justified use of black-box AI in high-stakes domains. By year-end, the field had consolidated around differentiation by domain: operational value in narrow contexts (marketing, logistics) versus persistent barriers and documented risks in broader organizational and high-stakes deployment scenarios. Adoption momentum remained decoupled from evidence of effectiveness, with organizations continuing investment despite quantified failures and unresolved deployment barriers."
    },
    {
      "period": "2026-Jan",
      "text": "Enterprise transition to operationalization emphasized data governance and architectural foundations; 62% of enterprises planning evolution to AI decision intelligence amid persistent 70-85% project failure rates and 42% initiative abandonment in 2025. Causal AI emerged as next-frontier addressing 74% faithfulness gap in existing systems. Clinical research documented error reduction (78% decline in guideline violations) through hybrid frameworks, yet deployment barriers remained: only 12% of executives reported both cost and revenue benefits; physician studies highlighted that reasoning cues must target high-discretion tasks where AI can add genuine value."
    },
    {
      "period": "2026-Feb",
      "text": "Multi-AI orchestration demonstrated operational feasibility (ARPIA 13-min data-to-strategy pipeline, causaLens enterprise deployments); professional services adoption jumped to 40% (2025: 22%), yet only 18% track ROI and 40% report policy confusion. Systematic LLM reasoning failures (Reversal Curse, Robustness Fragility, Working Memory Leaks) documented, undermining reliance on AI reasoning chains. Commercial AI-CDSS solutions still lack transparent training data and algorithm disclosure. Regulatory deadlines (CCPA Jan 2026, EU AI Act Aug 2026) drove governance platform launches. Across 2,400+ enterprise AI initiatives, 80.3% failed (33.8% abandoned, 28.4% deliver no value), with 95% GenAI pilots failing to reach production. Execution and governance remain bottlenecks, not capability."
    },
    {
      "period": "2026-Mar",
      "text": "Fundamental research documented persistent reasoning failures: CRYSTAL benchmark shows models skip 50%+ of reasoning steps (58% accuracy but only 48% reasoning recovery); BrainBench reveals stochastic reasoning gaps (6-16pp consistency variance even in top models); Stanford taxonomy classifies failures as architectural rather than scale-addressable. Real-world failures documented: NZ courts ruled AI-hallucinated legal citations may amount to obstruction; Deloitte refunded AUD 440K for AI-generated errors. Governance frameworks consolidated: RegTech expert consensus establishes human accountability cannot be delegated; KPMG legal analysis requires mandatory documentation of decision review. Practical deployment barriers clarified: reasoning models show 5x cost premium with performance ceiling at medium-complexity tasks (above which accuracy collapses). Field consensus solidifying around decision-support constraints: execution challenges and governance requirements are primary blockers, not reasoning capability gaps."
    },
    {
      "period": "2026-Apr",
      "text": "New empirical research confirmed architectural reasoning limits are scale-invariant: testing across 7 models (8B–235B parameters) showed all collapse at 20-30 parallel branches, and semantic variants of problems trigger 28-45% answer-flip rates — confirming brittleness is structural, not addressable by larger models. CMU testing of 14 leading LLMs (GPT-4, Claude 3, Gemini) found all fail simple logical contradiction detection, revealing that benchmark performance masks fundamental reasoning gaps. Harvard research added a new dimension: relational complexity causes accuracy to collapse when decisions require weighing multiple interacting factors simultaneously, directly constraining multi-factor analysis in healthcare and strategic contexts. Legal accountability frameworks tightened: Australian Federal Court (ASIC v Bekier) established that executives using public AI tools informally without verification breach their duty of care, reinforcing that decision accountability cannot be delegated to AI and mandating documentation of reasoning review. Latest evidence on personal decision support reveals critical tensions: targeted healthcare deployments show measurable value (RCT with 367 participants: 7.4-point satisfaction improvement, 50.7% vs 24.2% acceptance of AI recommendations), yet passive reliance on AI reasoning systematically erodes confidence in independent judgment and sense of authorship (behavioral study, 1,923 adults). Empirical adoption barriers remain severe: 9% of professionals trust AI for complex decisions despite 88% organizational adoption; 80% of workers reject enterprise AI tools (WalkMe, 3,750 professionals). Reasoning failure modes are now precisely characterized: hallucination rates span 22-94% across models (Stanford 2026 AI Index), with a documented \"Reliability Gap\" where capability scales 2-3x annually while reliability only 1.2-1.5x. Governance frameworks are consolidating: the SPEC framework achieves 89% accuracy vs 15% for unbounded RAG on incomplete-evidence scenarios by bounding AI confidence to evidential sufficiency; the CFA Institute articulates an epistemic anchoring principle that decision authority must remain in evidence-based human inquiry to avoid \"knowledge-collapse equilibrium.\""
    },
    {
      "period": "2026-May",
      "text": "Research intensified focus on cognitive and systemic failure modes in AI-assisted decision-making. Wharton's Cognitive Reflection Test (1,300+ participants) demonstrated the core paradox: AI-correct advice improves accuracy 25pp, but AI-wrong advice degrades 15pp below baseline—worse than no AI. Overconfidence persists even when users know AI errs 50% of the time, indicating cognitive surrender rather than rational reliance. Metacognitive stability revealed as critical: empirical testing across 11 frontier models shows 8 collapse under adversarial pressure (30.2pp accuracy drops), with only Anthropic's Constitutional AI showing near-immunity—suggesting alignment-specific training is prerequisite for trustworthy reasoning, not achievable through standard RLHF. Enterprise adoption gaps widened: <20% of AI pilots reach production due to missing trust infrastructure (audit trails, explainability, liability frameworks); yet production deployments in bounded domains (insurance claims, pharma R&D) show 50%+ efficiency gains when governance frameworks are load-bearing. Real-world interaction data (5.5M instances) reveals expert-domain performance plateaued at 14-16% dissatisfaction despite scale; raw models generate confident false theories when given deliberately flawed premises. Most critical finding: brief AI assistance (10 minutes) systematically impairs independent problem-solving through cognitive offloading, reducing retention and analytical skepticism—unintended consequence suggesting tool design actively erodes reasoning autonomy. Governance literature consolidated: AI system abandonment driven primarily by organizational dynamics and resource constraints, not ethics concerns; clinical AI remains confined to pilots not due to model limitations but institutional capacity gaps. The field's consensus strengthened: decision-support reliability requires orchestration of governance, human oversight, design discipline, and alignment-specific training—capability alone is necessary but insufficient."
    },
    {
      "period": "2026-Late May",
      "text": "Enterprise deployment momentum accelerated with defined governance patterns. SAP Sapphire 2026 unveiled 224 specialized Claude agents handling autonomous decisions across finance, HR, supply chain, and procurement for hundreds of thousands of enterprise customers globally—largest-scale production deployment of autonomous decision-making systems. CLR-voyance clinical reasoning system demonstrated production-scale maturity: 6+ months in hospital operation, 84.91% accuracy with physician-validated outcome rubrics, drafting thousands of inpatient notes—proving structured decision-support can reach institutional deployment. Yet critical limitations emerged: meta-analysis of 5 RCTs (12,657 participants) found AI clinical decision support produces only small, marginal improvements in diagnostic accuracy; pooled effect size narrowly above zero; strongest signal in radiology, weakest in complex reasoning domains. Governance failure patterns crystallized: survey of 650 enterprise leaders found 78% ran AI agent pilots but only 14% scaled to production; 90% of agent codebases fail EU AI Act compliance scans; governance architecture (audit trails, escalation paths, ownership documentation) emerged as primary blocker, not model performance. Personal effectiveness frameworks showed concrete signals: executive decision-support (persistent context, standing priorities, weekly accountability reviews) demonstrated production viability; structured reasoning patterns (Clarify-Then-Act, Plan-Then-Execute, Human-in-the-Loop gates) identified enterprise reliability requirements. Research confirmed psychological decision-making frameworks improve AI reasoning: TU Berlin study showed Recognition-Primed Decision-Making and Data-Frame Theory prompts boost healthcare advice accuracy 13% → 30%, establishing that human decision-making structures significantly enhance AI decision quality. Institutional barriers remained decisive: ICML paper documented that technical viability alone insufficient for scaling; decision-support systems require institutional alignment across approvals, oversight capacity, fiscal sustainability, and regulatory readiness. The field consolidated around a fundamental insight: decision-support reliability depends primarily on governance, human oversight, and institutional capacity—not model capability alone."
    },
    {
      "period": "2026-Jun",
      "text": "Consensus solidified around organizational readiness as the binding constraint. Thomson Reuters survey (1,816 professionals): 91% report AI value shortfall despite 74% regular use; 90% demand reasoning that can be explained and defended. NBER study of ~6,000 executives: ~90% report zero productivity impact; PwC CEO survey: 56% experienced no revenue/cost improvement, only 12% significant benefits—a $2.5T \"validation gap\" between investment and outcomes. But concrete deployments prove bounded decision-support works: Cox Communications achieved 7x first-year ROI on multi-agent sales decisioning (15,000 employees); Kai eliminated 99.5% false positives in security triage (2.5M findings). Nature Medicine RCT (9,600 patients, 16 sites, Kenya): AI-assisted clinical decision support improved documentation quality but failed to reduce treatment failure (2.2% vs 2.0%, P=0.13)—emblematic outcome-documentation gap. The deepest barrier is organizational: only 15% of enterprises can leverage 82% of AI-ready individuals; 55% of professionals cite lack of structured workflows as adoption barrier. Harvard AI Institute director (HBR): \"core constraint is organizational ability to make decision-making processes explicit.\" Mechanistic research exposures: Stroop task (PNAS Nexus) shows models collapse under cognitive load (91%→1% accuracy); reasoning models spend MORE tokens on failed tasks (arXiv:2606.26502, Cohen's d 1.47–3.13), inverse of human behavior; frontier models commit 6-12% safety violations on critical tasks (Arachne, 81,000-person survey). Reasoning capability advances documented (OpenAI disproving 80-year-old math conjecture, TTE-Flash, Recursive Language Models), yet hybrid architectures combining LLM planners with deterministic execution engines prove necessary for reliable analytics (trade-off: reasoning depth vs. factual grounding). Albabtain's organizational design study: AI augments or substitutes judgment based on design choices (transparency, override friction, performance metrics)—augmentation requires low-friction overrides, not AI compliance metrics. 72% of CEOs expect AI to support human-directed decisions (not autonomous), with trust/security concerns (31%) and data quality (34%) cited as barriers. Government-scale adoption signals: California partnership provides Claude across all state agencies (policy deliberation, Medicaid workflows, cyber triage). The field's 2026 consensus: decision-support reliability requires governance, explicit workflow design, and human judgment integration—not model accuracy alone."
    },
    {
      "period": "2026-Jul",
      "text": "The validation gap deepened with harder numbers. NBER (~6,000 executives) and PwC CEO survey together document 90% reporting zero productivity impact and 56% seeing no revenue or cost improvement — a $2.5T blind spot. Against that backdrop, two production deployments documented genuine bounded ROI: Cox Communications 7x first-year return on multi-agent sales decisioning and Kai's 99.5% false-positive elimination in security triage. Architectural research confirmed models cannot self-regulate reasoning effort: reasoning tokens scale inversely with task success (Cohen's d 1.47–3.13), and the Stroop task collapse (91%→1% accuracy under cognitive load) is now confirmed as architectural rather than tunable. Organizational design research (Albabtain) established that augmentation vs. substitution of human judgment is determined by governance choices — low-friction overrides and judgment-rewarding metrics — not model capability. Mid-July evidence added both operational tooling and governance boundaries: a three-country RCT of 249 physicians found AI improved clinical reasoning (Kenya +18%) when paired with proper human-AI design, while peer-reviewed ACL research found frontier reasoning models achieve under 25% instruction-following compliance during their own reasoning traces — a control gap limiting governed deployment. Structured decision frameworks moved from instructional content toward operational tooling: an open-source Claude Code plugin implemented multi-agent debate to surface reasoning blind spots, and a 39-framework cognitive-technique catalog shipped via the Claude Code marketplace. Legal commentary reinforced that decision accountability cannot be delegated to AI, and a 1,000-dealmaker survey found 62% consider human-only decisions indefensible yet only 22% delegate final decisions to AI — underscoring the governance boundary."
    },
    {
      "period": "2026-Aug",
      "text": "New evidence sharpened the accountability-capability gap: a real-world benchmark found frontier reasoning models fail to beat market baseline on decision-making under uncertainty, chain-of-thought explanations were shown to systematically misrepresent actual model computation, and a peer-reviewed study found AI advice cut human accuracy from 27% to 9% while nearly tripling confidence — while an HBR field study documented frontline workers defending AI decisions they neither created nor understood. A 758-consultant BCG field experiment quantified the capability boundary precisely (+12.2% productivity/+40% quality inside the AI's competence frontier vs. a 19pp correctness drop outside it), and multi-source analysis (S&P, Gartner, MIT, BCG, IBM) confirmed 42% of companies abandon AI initiatives due to organizational rather than technical barriers. Mid-to-late August deployments revealed both operational maturity and architectural constraints: ambient clinical decision-support reached 1,744 U.S. hospitals with measurable burnout reduction (51.9%→38.8%) yet only 10% have formal governance, while Maven demonstrates 5,000+ daily targeting decisions in military operations constrained by organizational diffusion barriers. Research confirmed three binding failure modes—position bias renders decisions order-sensitive across all tested models; interface design that helps experts actively misleads novices in explainability; constraint-preservation failures (Gemini 30+ cases) show models violate explicit logical rules—indicating architectural limits rather than trainable gaps. Widespread personal adoption (55% of U.S. workers report AI time-savings, yet 70% use AI for high-stakes decisions with 50% reporting overreliance) paired with documented harms (5M+ UK citizens reporting financial/health losses) establishes adoption outpacing governance maturity as the binding constraint on tier advancement. A US Census Bureau survey independently corroborated the adoption baseline (55% weekly worker use, 31% reporting 1-2 hour time savings), while a 9-LLM study confirmed AI decision quality is sensitive to option-presentation order, and 81% of physicians now use AI professionally even as FDA narrowed oversight and shifted liability onto clinicians—reinforcing that governance gaps, not model capability, remain the binding constraint. Late-August evidence added labor-tribunal and framework signals: Australia's Fair Work Commission condemned a dismissed worker's reliance on AI as a \"quasi-legal advisor\" and will mandate AI-use disclosure from October 2026; a University of Bayreuth experiment (n=3,700) found AI financial advice raised decision accuracy from 64% to 87% when aligned but crashed it to 34% when misaligned, with conflict-of-interest disclosure offering little protection; and new drafting-versus-deciding frameworks (a GA Claude Skill for structured reflection, and a Salesforce-consulting staged-autonomy model achieving 50%+ case deflection) continued formalizing where AI assistance should stop short of the final decision."
    },
    {
      "period": "2026-Sep",
      "text": "Judgment-erosion and governance-rollback evidence converged. BCG's C-suite survey (70 leaders) found 50% already observe judgment/problem-solving skill decay and 90% report employees have stopped checking AI work; a Harvard field experiment (228 senior reviewers) documented automation bias overwhelming expert judgment, with explanations paradoxically worsening rather than correcting the bias. Rollback evidence hardened: a meta-analysis of MIT/RAND/S&P Global/Gartner studies found 95% of AI pilots report zero ROI and 42% abandon before production; AvePoint's survey of 3,235 leaders found 87% delayed deployments over governance risk (40.7% canceled GenAI rollouts, up from 31.7% in 2025); a separate 2,527-enterprise survey found 74% rolled back deployed AI agents, with \"mature guardrails\" firms rolling back at an even higher 81%. PACT's benchmark of 22 models across 12 regulated domains found compliance violations escalate from 6-10% at baseline to 65% under workplace pressure, and a Polish Supreme Administrative Court case documented AI hallucination (fabricated propositions attached to real case numbers) creating professional liability. Against this backdrop, Meta's CORAL framework demonstrated closed-loop decision-optimization at production scale across two billion-user platforms, and Salesforce embedded 37 prebuilt sales decision-support skills inside Claude for pilot customers, routing actions through business-rule governance — showing capable, governed systems remain the exception rather than the norm. Further evidence pointed the same way: Taste-Bench found frontier models answer only 59.7% of decision-fork questions correctly regardless of reasoning budget, and a SOCPAC case saw an AI-fabricated nuclear claim come within minutes of triggering a ship boarding. Scale deployments continued (Xylem 15,000 employees; One New Zealand's 50+ agents), alongside a five-dimension accountability framework for AI briefings and a Microsoft review of 120+ overreliance papers."
    }
  ],
  "historyFallback": false,
  "lastUpdated": "2026-09-27",
  "domain": {
    "id": "personal-effectiveness",
    "label": "Personal Effectiveness",
    "icon": "✨"
  },
  "url": "https://www.thestateofplay.ai/practice/decision-support-and-reasoning-frameworks",
  "license": "CC BY 4.0",
  "licenseUrl": "https://creativecommons.org/licenses/by/4.0/",
  "generatedAt": "2026-10-01"
}