{
  "slug": "agentic-coding-for-exploration-and-prototyping",
  "name": "Agentic coding for exploration & prototyping",
  "tier": "leading-edge",
  "trend": "steady",
  "blockerType": null,
  "tools": [
    {
      "name": "Claude Code",
      "url": "https://claude.com/"
    },
    {
      "name": "Cursor",
      "url": "https://www.cursor.com/"
    },
    {
      "name": "GitHub Copilot",
      "url": "https://github.com/features/copilot"
    },
    {
      "name": "Devin",
      "url": "https://devin.ai/"
    },
    {
      "name": "Google Antigravity",
      "url": "https://cloud.google.com/colab/"
    },
    {
      "name": "OpenAI Codex",
      "url": "https://platform.openai.com/"
    }
  ],
  "evidence": [
    {
      "title": "Are AI coders more trouble than they’re worth?",
      "url": "https://www.computerweekly.com/news/366651377/Are-AI-coders-more-trouble-than-theyre-worth",
      "date": "2026-09-28",
      "type": "adoption-metric",
      "added": "2026-09-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Undo-commissioned Coleman Parkes poll of 300 developers: 93% hit hallucinations, 94% lose productivity analysing agent code and 35% see it reach production before it is understood. A vendor-funded negative signal."
    },
    {
      "title": "AI-Generated Code Grades Higher in Review, Yet Triggers Rise in Production Incidents",
      "url": "https://www.devopsdigest.com/ai-generated-code-grades-higher-review-yet-triggers-rise-production-incidents",
      "date": "2026-09-25",
      "type": "adoption-metric",
      "added": "2026-09-29",
      "superseded_by": null,
      "window": null,
      "explanation": "New Relic survey of 200 US leaders: 88% write vibe coding into production policies and only 5% restrict it to non-production. It shows the exploration–production boundary eroding, with 78% reporting more incidents."
    },
    {
      "title": "The State Of AI Harness Engineering 2026",
      "url": "https://marmelab.com/blog/2026/09/24/the-state-of-ai-harness-engineering-2026.html",
      "date": "2026-09-24",
      "type": "opinion",
      "added": "2026-09-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Audit of 246 repositories and 57 publications, citing one controlled run in which the same model's success rate across 8 harnesses ranged from 68% to 88%. Harness choice materially shapes agent outcomes."
    },
    {
      "title": "Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise",
      "url": "https://arxiv.org/html/2609.28919",
      "date": "2026-09-24",
      "type": "research-paper",
      "added": "2026-09-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Documents enterprise seat scale (Accenture ~30,000, Cognizant up to 350,000). In a 10,000-seat emulation, routing recovers 14–21% of model spend. It also flags Claude Code's lock-in to Anthropic models."
    },
    {
      "title": "A Review of Research on the Development of Artificial Intelligence Programming Agent-Assisted Software",
      "url": "https://journals.zeuspress.org/index.php/CAI/article/view/1380",
      "date": "2026-09-19",
      "type": "research-paper",
      "added": "2026-09-29",
      "superseded_by": null,
      "window": null,
      "explanation": "A review of 60 papers finds that agents cut the cost of repetitive coding and help non-professionals finish prototypes. Gains depend on task size, familiarity, quality standards and review costs."
    },
    {
      "title": "Inside OpenAI’s agentic software factory",
      "url": "https://newsletter.pragmaticengineer.com/p/openai-software-factory",
      "date": "2026-09-15",
      "type": "case-study",
      "added": "2026-09-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Independent first-hand case study: Codex usage in OpenAI's non-engineering organisations went from ~0% to 90% in four months, and overall usage rose from 60% to 90%. PRs and code review 'need to be rethought'."
    },
    {
      "title": "JetBrains Developer Ecosystem Survey 2026: 90% Weekly Agent Usage, Claude Code Lead",
      "url": "https://www.jetbrains.com/lp/devecosystem-2026/",
      "date": "2026-09-10",
      "type": "adoption-metric",
      "added": "2026-09-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Survey of 15,000+ professional developers (May-July 2026): 90% use agents weekly, 68% daily; Claude Code adopted by 39% (primary tool), signaling population-level mainstream adoption reinforcing leading-edge tier classification."
    },
    {
      "title": "Remediate Code Quality findings with agentic autofix - GitHub Changelog",
      "url": "https://github.blog/changelog/2026-09-09-remediate-code-quality-findings-with-agentic-autofix/",
      "date": "2026-09-09",
      "type": "product-ga",
      "added": "2026-09-15",
      "superseded_by": null,
      "window": null,
      "explanation": "GitHub Copilot shipped general availability for bulk agentic remediation of up to 25 code-quality findings per assignment, demonstrating ecosystem maturation beyond exploration into routine maintenance and compliance workflows."
    },
    {
      "title": "Agentic Coding in the Wild: What 95 Trillion Tokens Reveal About Codex CLI Workloads",
      "url": "https://codex.danielvaughan.com/2026/09/08/agentic-coding-in-the-wild-production-scale-kv-cache-workload-codex-cli/",
      "date": "2026-09-08",
      "type": "research-paper",
      "added": "2026-09-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Microsoft Research analysis of 13.5M GitHub Copilot sessions (95 trillion tokens): 87% of LLM calls are agent-initiated (not user-driven), revealing production-scale exploration workloads with distinct KV cache and context coherence challenges."
    },
    {
      "title": "Benchmarking AI Coding Agents: Why Output Proof Fails to Stop False Successes",
      "url": "https://explore.n1n.ai/blog/benchmarking-ai-coding-agents-false-success-2026-09-05",
      "date": "2026-09-05",
      "type": "research-paper",
      "added": "2026-09-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical benchmark (424 trials, claude-haiku-4-5): forcing agents to output structured proof increased evidence 30x but did NOT reduce false-success rates on multi-file tasks (75% with receipts vs 73.6% baseline), revealing verification limits."
    },
    {
      "title": "Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle",
      "url": "https://arxiv.org/abs/2609.04681",
      "date": "2026-09-04",
      "type": "research-paper",
      "added": "2026-09-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed synthesis identifying the Agentic SDLC Throughput Paradox: agents accelerate code generation (+180% commits) but constrain production releases (+30%), with verification tax and governance costs dominating downstream economics."
    },
    {
      "title": "Agentic technical debt: a governance framework for AI agents",
      "url": "https://datasciencedojo.com/blog/agentic-technical-debt/",
      "date": "2026-09-04",
      "type": "research-paper",
      "added": "2026-09-15",
      "superseded_by": null,
      "window": null,
      "explanation": "University of Pittsburgh + Ejento AI forthcoming ACM paper: formalizes agentic technical debt and stochastic tax (recurring evaluation/retry costs), proposes five governance controls for production-qualified changes in probabilistic agent systems."
    },
    {
      "title": "AI Coding Agents 2026: Codex, Claude Code, Copilot & Developer Work",
      "url": "https://polprog.pl/en/learning/ai-coding-agents-2026-how-they-change-developer-work",
      "date": "2026-09-04",
      "type": "research-paper",
      "added": "2026-09-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Synthesis of primary sources (Anthropic 400K sessions, OpenAI telemetry, METR RCT): confirms workflow shift toward multi-hour autonomous tasks (70.2% of users) while METR RCT documents 19% slowdown on complex tasks despite gains on simple ones."
    },
    {
      "title": "A retrospective on productivity and (lack?) of quality with coding agents",
      "url": "https://www.efekarakus.com/2026/09/03/retro-on-productivity-and-quality-with-coding-agents.html",
      "date": "2026-09-03",
      "type": "case-study",
      "added": "2026-09-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Practitioner retrospective from exe.dev (9 engineers): identified how high-trust teams earn the right to eliminate secondary review through architectural discipline, integration testing, and design-first workflows—documenting organizational prerequisites for reliable agentic production code."
    },
    {
      "title": "Malicious .git Configs Can Make Claude, Codex, Cursor, and Other AI Agents Run Attacker Code",
      "url": "https://thehackernews.com/2026/09/malicious-git-configs-can-make-claude.html",
      "date": "2026-09-02",
      "type": "opinion",
      "added": "2026-09-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Manifold Security GitSpawn disclosure: 8 CVEs across 7 agents including Claude Code pre-trust execution via Git config; multiple unpatched as of Sept 1, reflecting production-readiness gaps in exploration-stage tools."
    },
    {
      "title": "Code Delivery Is No Longer the Bottleneck. Ownership Is. (304,362 Verified AI-Authored Commits)",
      "url": "https://www.sprucetech.ai/blog/code-delivery-ownership-bottleneck",
      "date": "2026-09-01",
      "type": "adoption-metric",
      "added": "2026-09-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Large-scale empirical analysis (Liu et al., peer-reviewed): 484,606 issues introduced across 304K AI commits; 24.2% of bugs never fixed, 41.1% security issues survive—directly quantifying quality gaps constraining exploration-to-production conversion."
    },
    {
      "title": "Claude Code の法人・業務活用事例 (28 Verified Enterprise Deployments)",
      "url": "https://ai-jirei.com/claude-code-biz.html",
      "date": "2026-08-28",
      "type": "case-study",
      "added": "2026-09-01",
      "superseded_by": null,
      "window": null,
      "explanation": "Curated database of 28 Claude Code enterprise cases (Deepgram, Stripe, estie, League, Ramp, HubSpot, others) with verified metrics: cycle-time halving, PR velocity +64%, review time −70%, demonstrating sustained ROI across sectors."
    },
    {
      "title": "Teamwork: When AI Becomes a Research Partner",
      "url": "https://antigravity.google/blog/teamwork-when-ai-becomes-a-research-partner",
      "date": "2026-08-27",
      "type": "product-ga",
      "added": "2026-09-01",
      "superseded_by": null,
      "window": null,
      "explanation": "Google Antigravity Teamwork multi-agent framework now GA with concrete exploration outcomes: 7 math theorems solved (JMLR/FOCS venues), autonomous RISC-V simulator, and upstream library optimizations—proof of agentic exploration at research scale."
    },
    {
      "title": "Enabling Independent Research on How People Use Claude (250K Conversations)",
      "url": "https://www.linkedin.com/posts/anthropicresearch_enabling-independent-research-on-how-people-activity-7498430536231604225-Gamh",
      "date": "2026-08-26",
      "type": "research-paper",
      "added": "2026-09-01",
      "superseded_by": null,
      "window": null,
      "explanation": "Stanford SALT Lab analysis of 250K real Claude conversations: >50% involve consequential/hard-to-undo work, ~75% show humans setting direction while adapting outputs—validates production-scale agentic delegation in practice."
    },
    {
      "title": "From Prototype to Production: Building AI-Constructed Products at Enterprise Scale",
      "url": "https://www.atlassian.com/blog/jira/ai-built-prototype-to-production",
      "date": "2026-08-26",
      "type": "case-study",
      "added": "2026-09-01",
      "superseded_by": null,
      "window": null,
      "explanation": "Atlassian case: 15-hour prototype to running app via multi-agent orchestration, 5x output velocity; critical lesson—agents outpaced team comprehension, breaking in production; bottleneck shifted from code generation to human context management."
    },
    {
      "title": "How Do Agents Fail on AutoResearch: Diagnostic Taxonomy of 45 Failure Patterns",
      "url": "https://www.alphaxiv.org/abs/2608.14905",
      "date": "2026-08-25",
      "type": "research-paper",
      "added": "2026-09-01",
      "superseded_by": null,
      "window": null,
      "explanation": "Analysis of 800 trajectories across 100 frontier research tasks identifies metacognitive loop as missing—agents lack ability to verify outputs against findings; dominant failure (52%) is execution fault where spec is correct but implementation diverges."
    },
    {
      "title": "When AI Writes and Ships the Code, Experienced Developers Move 19% Slower",
      "url": "https://www.martincid.com/technology-sv/what-is-code-autonomy-ai-driven-software-development/",
      "date": "2026-08-24",
      "type": "opinion",
      "added": "2026-09-01",
      "superseded_by": null,
      "window": null,
      "explanation": "METR randomized controlled trial: developers 19% slower with AI tools despite self-reporting 20% faster; causes are reprompting overhead, verification time, and cognitive switching—establishes negative signal on productivity paradox."
    },
    {
      "title": "The Agentic Coding Backlash Is Really About Work (Cognitive Costs of Orchestration)",
      "url": "https://www.linkedin.com/pulse/agentic-coding-backlash-really-work-john-willis-yi5gc",
      "date": "2026-08-23",
      "type": "opinion",
      "added": "2026-09-01",
      "superseded_by": null,
      "window": null,
      "explanation": null
    },
    {
      "title": "SWE-bench Science: Coding Agents on Scientific Software Engineering",
      "url": "https://www.alphaxiv.org/abs/2608.19799",
      "date": "2026-08-20",
      "type": "research-paper",
      "added": "2026-09-01",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed evaluation on 119 scientific engineering tasks: Claude Code + Opus 5 achieves <50% pass rate; identifies four failure mechanisms (knowledge gaps, misguided exploration, incomplete repair, poor generalization) critical for boundary assessment."
    },
    {
      "title": "An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation",
      "url": "https://www.alphaxiv.org/abs/2608.20195",
      "date": "2026-08-20",
      "type": "research-paper",
      "added": "2026-09-01",
      "superseded_by": null,
      "window": null,
      "explanation": "Analysis of 557 sessions (94,813 events): agents rely on instruction files (60.5%) over API refs (1.3%), self-direct exploration 70.2% of the time; reveals how agents actually conduct exploratory workflows in practice."
    },
    {
      "title": "AI Coding Agents: Adoption Trends (JetBrains 2026)",
      "url": "https://blog.jetbrains.com/research/2026/08/ai-coding-agent-adoption-2026/",
      "date": "2026-08-19",
      "type": "adoption-metric",
      "added": "2026-09-01",
      "superseded_by": null,
      "window": null,
      "explanation": "Large-scale survey of 15,000+ professional developers: 90% use AI agents weekly, Claude Code 39% adoption (47% US), #1 tool for 31%, double the usage rate of GitHub Copilot—mainstream adoption inflection confirmed."
    },
    {
      "title": "Why We Decided Not to Adopt Devin After a One-Week Trial",
      "url": "https://zenn.dev/vesslabs/articles/da9593d77369ee?locale=en",
      "date": "2026-08-18",
      "type": "case-study",
      "added": "2026-09-01",
      "superseded_by": null,
      "window": null,
      "explanation": "VESS Labs structured evaluation: cloud-based agent worked technically but was rejected—iOS simulator gaps, session tracking overhead, visibility metrics mismatch for colocated teams; documents structural adoption barriers beyond capability."
    },
    {
      "title": "AI Programming in 2026: Tools, Techniques, Evidence",
      "url": "https://capitalandcompute.net/blog/state-of-ai-programming-2026/",
      "date": "2026-08-15",
      "type": "adoption-metric",
      "added": "2026-08-18",
      "superseded_by": null,
      "window": null,
      "explanation": "Analysis of AI coding tool adoption showing 72% daily use but 96% lack full trust, with METR RCTs showing productivity gains uncertain (confidence intervals straddling zero), evidencing maturity gap between exploration use and reliable deployment."
    },
    {
      "title": "Google Antigravity & Gemini 3.7 Flash: The Clear Developer Guide (2026)",
      "url": "https://www.toolbit.ai/blog/google-antigravity-inside-new-agentic",
      "date": "2026-08-14",
      "type": "opinion",
      "added": "2026-08-18",
      "superseded_by": null,
      "window": null,
      "explanation": "Third-party developer guide with comparative analysis of Antigravity ecosystem (IDE, Desktop 2.0, CLI, SDK) vs Cursor and Windsurf, explaining hybrid reasoning and 4-step execution loop for exploration workflows."
    },
    {
      "title": "Gemini 3.7 Flash in Google Antigravity",
      "url": "https://antigravity.google/blog/gemini-3-7-flash-in-google-antigravity",
      "date": "2026-08-13",
      "type": "product-ga",
      "added": "2026-08-18",
      "superseded_by": null,
      "window": null,
      "explanation": "Google model release with specific SWE benchmarks (DeepSWE 65.3% vs 49.0%, FrontierCode 43.6% vs 34.4%), real deployed examples, and introductory pricing enabling cost-effective exploration at scale."
    },
    {
      "title": "Deep Agentic Search vs Semantic Search: Why Delegating Code Exploration to Subagents Costs More and Finds Less",
      "url": "https://codex.danielvaughan.com/2026/08/13/deep-agentic-search-vs-semantic-search-repository-code-qa-codex-cli-subagent-delegation-context-rot/",
      "date": "2026-08-13",
      "type": "research-paper",
      "added": "2026-08-18",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical study comparing repository exploration strategies across 720 questions: semantic search achieved 65.2% vs subagent delegation 46.2% (19-point gap) at half the cost, documenting core failure mode in agentic architecture coordination."
    },
    {
      "title": "SWE-Bench Pro: coding agents stall on hard tasks",
      "url": "https://agentry.news/coding-agents-plateau-below-45-on-realistic-benchmarks",
      "date": "2026-08-11",
      "type": "research-paper",
      "added": "2026-08-18",
      "superseded_by": null,
      "window": null,
      "explanation": "Agents plateau below 20% on proprietary/realistic SWE-Bench Pro tasks (vs 45% public benchmarks), constraining what complexity exploration workflows can reliably tackle in production-like codebases."
    },
    {
      "title": "88% of AI agent pilots never reach production, Forrester and Anaconda find",
      "url": "https://buildapps.co.uk/signals/ai-agent-pilots-88-percent-never-reach-production/",
      "date": "2026-08-10",
      "type": "industry-report",
      "added": "2026-08-18",
      "superseded_by": null,
      "window": null,
      "explanation": "Forrester and Anaconda research showing 88% pilot-to-production failure despite 71% daily adoption, with blockers (evaluation gaps 64%, governance 57%, reliability 51%) revealing why agentic coding remains at exploration tier."
    },
    {
      "title": "Cheap Code, Costly Judgment: Governance Conversion for Agentic Software Engineering",
      "url": "https://codex.danielvaughan.com/2026/08/10/cheap-code-costly-judgment-governance-conversion-agentic-software-engineering-codex-cli/",
      "date": "2026-08-10",
      "type": "case-study",
      "added": "2026-08-18",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical 12-week production deployment: 420 KLOC code with 1.16 MLOC governance infrastructure (2.76:1 test-to-code ratio), documenting structural failure classes (hallucinated correctness, budget-pressure shortcuts, context bloat) requiring governance-conversion cycle."
    },
    {
      "title": "The AI Governance Gap - The Delegated Engineering Gap",
      "url": "https://www.linkedin.com/pulse/ai-governance-gap-delegated-engineering-patrick-upmann-c604f",
      "date": "2026-08-10",
      "type": "adoption-metric",
      "added": "2026-08-18",
      "superseded_by": null,
      "window": null,
      "explanation": "Multi-source survey data showing 42% of committed code is AI-generated with critical verification gap: 96% distrust vs. 48% always verify, and only 49.6% of agent PRs include test changes—quantifying governance infrastructure lag."
    },
    {
      "title": "Claude Code and Gemini CLI Flaws Let GitHub Issue Reach CI Workflow Secrets",
      "url": "https://thehackernews.com/2026/08/claude-code-and-gemini-cli-flaws-let.html",
      "date": "2026-08-07",
      "type": "news-coverage",
      "added": "2026-08-18",
      "superseded_by": null,
      "window": null,
      "explanation": "Black Hat USA research identifying multiple CVEs across agentic coding tools (Claude Code, Gemini CLI, Codex) with real deployment vulnerabilities and technical root causes, demonstrating security risks in exploration workflows."
    },
    {
      "title": "Claude Code's Biggest Challenge: Scope Hallucinations",
      "url": "https://www.linkedin.com/posts/wishworld_github-wishworldvishal-ai-pm-portfolio-activity-7490718076821204992-1EDK",
      "date": "2026-08-05",
      "type": "opinion",
      "added": "2026-08-18",
      "superseded_by": null,
      "window": null,
      "explanation": "Practitioner identifies critical adoption failure mode: agents build unasked-for features and invent requirements, mitigated by ADLC framework enforcing scope discipline; shows exploration requires process engineering alongside capability."
    },
    {
      "title": "Microsoft's CLI Coding Agent Study: Adoption Is a Workflow Problem",
      "url": "https://www.developersdigest.tech/blog/microsoft-cli-coding-agents-study-2026",
      "date": "2026-07-31",
      "type": "adoption-metric",
      "added": "2026-08-04",
      "superseded_by": null,
      "window": null,
      "explanation": "Microsoft's early-2026 rollout of Claude Code and GitHub Copilot CLI to tens of thousands of engineers found 24% more merged PRs among adopters; demonstrates agentic CLI exploration improving organizational throughput with social-network-driven adoption."
    },
    {
      "title": "Codex Rebuilds Genomic Software in New OpenAI Field Report",
      "url": "https://getaibook.com/news/codex-rebuilds-genomic-software-in-new-openai-field-report/",
      "date": "2026-07-30",
      "type": "case-study",
      "added": "2026-08-04",
      "superseded_by": null,
      "window": null,
      "explanation": "OpenAI field report on 8 real scientific-computing case studies using Codex and Claude Code: agents successfully executed GPU-native builds, language migrations, and large-scale refactors (cyvcf2); demonstrates exploration-stage productivity with shift from implementation to verification as bottleneck."
    },
    {
      "title": "AI Agents Are Live. Governance Is Still in the Lab.",
      "url": "https://vortx.ch/ai-agents-are-live-governance-is-still-in-the-lab/",
      "date": "2026-07-28",
      "type": "industry-report",
      "added": "2026-08-04",
      "superseded_by": null,
      "window": null,
      "explanation": "Synthesis of 2026 industry surveys (Forrester, Deloitte, IDC/AWS) showing 71% of enterprises lack formal governance framework despite broad agentic AI adoption; 88.4% experienced agent-related security breaches—documents structural governance readiness gap."
    },
    {
      "title": "Agentic Code Debt: The Board's Next Risk Frontier",
      "url": "https://wgaadvisors.com/2026/07/28/agentic-code-debt-boards-next-risk-frontier/",
      "date": "2026-07-28",
      "type": "opinion",
      "added": "2026-08-04",
      "superseded_by": null,
      "window": null,
      "explanation": "Strategic analysis identifying novel 'agentic code debt' risk class: production systems with uncomprehended logic, hallucinated dependencies, and architectural knowledge loss—qualitatively different from traditional technical debt."
    },
    {
      "title": "GitHub Copilot Under Pressure: Cursor and Claude Code Are Eating Its Lunch (2026)",
      "url": "https://pasqualepillitteri.it/en/news/3392/github-copilot-cursor-claude-code-ai-coding-showdown-2026",
      "date": "2026-07-28",
      "type": "adoption-metric",
      "added": "2026-08-04",
      "superseded_by": null,
      "window": null,
      "explanation": "JetBrains survey of 10,000+ developers shows Claude Code 46% vs Copilot 9% preference among senior developers; 6x adoption growth April 2025–January 2026; agentic tools displacing previous-generation assistants across enterprise."
    },
    {
      "title": "Agentic Test Processes, LLM Benchmarks, and Other Notes on Agentic Coding",
      "url": "https://news.quak.cc/2026/07/26/agentic-test-processes-llm-benchmarks-and-other-notes-on-agentic-coding/",
      "date": "2026-07-26",
      "type": "opinion",
      "added": "2026-08-04",
      "superseded_by": null,
      "window": null,
      "explanation": "Practitioner account of heavy agentic use since mid-2025 documenting both breakthroughs (testing-heavy workflows, fuzzing-driven test generation) and critical failures (hallucinated bug reproductions, agent confidence masking fabrication)."
    },
    {
      "title": "Why 88% of AI Agent Pilots Never Reach Production in 2026",
      "url": "https://www.fatherofai.in/blog/agentic-ai-production-reliability-reckoning-2026/",
      "date": "2026-07-24",
      "type": "industry-report",
      "added": "2026-08-04",
      "superseded_by": null,
      "window": null,
      "explanation": "Gartner 2026 enterprise data: 88% of agent pilots fail to reach production, only 31% of enterprises have agents deployed, 70% cite unpredictability as top blocker—critical evidence of maturity gap between exploration pilots and operational deployment."
    },
    {
      "title": "Issue #2 | July 24, 2026 (GetDevDone newsletter)",
      "url": "https://www.linkedin.com/pulse/issue-2-july-24-2026-dmytro-mashchenko-vkjce",
      "date": "2026-07-24",
      "type": "opinion",
      "added": "2026-08-04",
      "superseded_by": null,
      "window": null,
      "explanation": "GetDevDone analysis of exploration-to-production gap: 78% run agentic pilots but only 14% scale to production; identifies vibe-coding risks, output-degradation under load, and governance gaps as binding constraints blocking broader adoption."
    },
    {
      "title": "How Do AI Coding Agents Contribute to Software Development? An Empirical Study of Agentic Pull Requests",
      "url": "https://arxiv.org/abs/2607.21832",
      "date": "2026-07-23",
      "type": "research-paper",
      "added": "2026-08-04",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed longitudinal empirical study using AIDev dataset analyzing agentic pull requests vs human PRs across development lifecycle, merge rates, and task distributions—provides empirical foundation for understanding real-world agentic contribution patterns."
    },
    {
      "title": "Autonomous Coding Agents: Beyond Developer Productivity",
      "url": "https://www.linkedin.com/pulse/autonomous-coding-agents-beyond-developer-dccfe",
      "date": "2026-07-23",
      "type": "opinion",
      "added": "2026-08-04",
      "superseded_by": null,
      "window": null,
      "explanation": "Analysis of AIDev dataset shows velocity gains front-loaded and fade; quality risks persist with static-analysis warnings +18%, cognitive complexity +39%; documents code-quality degradation across deployments."
    },
    {
      "title": "Small Teams Are the Heaviest Users of AI Coding Agents",
      "url": "https://www.helpnetsecurity.com/2026/07/22/users-of-ai-coding-agents/",
      "date": "2026-07-22",
      "type": "adoption-metric",
      "added": "2026-08-04",
      "superseded_by": null,
      "window": null,
      "explanation": "RIT study of 25,264 agentic PRs (May–July 2025) shows adoption concentrates in small teams (1–5 contributors averaging 50.2 agentic PRs/quarter) with single-developer review bottleneck (78.9%)—defines exploration-phase workflow constraint."
    },
    {
      "title": "Test Coverage Analysis of Agentic Pull Requests",
      "url": "https://www.themoonlight.io/en/review/test-coverage-analysis-of-agentic-pull-requests",
      "date": "2026-07-21",
      "type": "research-paper",
      "added": "2026-08-04",
      "superseded_by": null,
      "window": null,
      "explanation": "Large empirical study of 4,882 agent-generated PRs shows 50.4% include zero test changes, error-handling blocks under-tested 81-86%, revealing systematic quality limitations in agentic exploration output."
    },
    {
      "title": "GitHub、エージェンティック・コーディング早期採用の課題 (Early Adoption of Agentic Coding Tools by GitHub Projects)",
      "url": "https://aiedgeline.jp/articles/2026-07-16-early-adoption-of-agentic-coding-tools-by-github-projects-zft2rz/",
      "date": "2026-07-16",
      "type": "research-paper",
      "added": "2026-07-21",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical research analyzing 25,264 agentic PRs from 2,361 GitHub repos; shows adoption concentration in small projects (1–5 contributors), single-human oversight bottleneck dominates, and early adoption patterns constrained to exploration/small-team contexts."
    },
    {
      "title": "When AI Stops Assisting and Starts Owning Your Codebase",
      "url": "https://theintelligencestack.substack.com/p/phase-two-when-ai-stops-assisting",
      "date": "2026-07-13",
      "type": "opinion",
      "added": "2026-07-21",
      "superseded_by": null,
      "window": null,
      "explanation": "Critical analysis of AI-first development's hidden costs with quantified data: 8× code duplication post-adoption, 45% OWASP Top 10 vulnerabilities, real enterprise pattern reversals (Klarna, Commonwealth, IBM), and Gartner projection of 40% agentic AI project cancellations by end-2027."
    },
    {
      "title": "How we shipped a PoC in 72 hours with AI agents",
      "url": "https://mev.com/blog/how-mev-shipped-a-poc-in-72-hours-with-ai-agents",
      "date": "2026-07-10",
      "type": "case-study",
      "added": "2026-07-21",
      "superseded_by": null,
      "window": null,
      "explanation": "Named organization (MEV) delivered working energy-analysis PoC in 72 hours using 7-agent pipeline on Claude Code; 127 PRs merged, ~60 hours review time saved, seven agent specs published; demonstrates rapid exploration/prototyping delivery cadence."
    },
    {
      "title": "Spring 2026 GenAI Code Security Update",
      "url": "https://www.veracode.com/blog/spring-2026-genai-code-security/",
      "date": "2026-07-09",
      "type": "industry-report",
      "added": "2026-07-21",
      "superseded_by": null,
      "window": null,
      "explanation": "Longitudinal security testing (150+ models, 3+ years, 80 tasks) shows syntax correctness improved ~50% → ~95% but security pass rate remains flat at 45–55% regardless of model generation; Java worst at 29%—persistent security gap despite coding improvements."
    },
    {
      "title": "AI 時代のプロダクト開発、私はこう変わってきた (Product Development in the AI Era)",
      "url": "https://zenn.dev/estie/articles/8268006c43706c",
      "date": "2026-07-07",
      "type": "case-study",
      "added": "2026-07-21",
      "superseded_by": null,
      "window": null,
      "explanation": "Tech lead at estie documents 18-month workflow evolution (Feb 2025–July 2026) using agentic coding: tool transitions (Devin → Cursor → Claude Code), workflow shifts (9.6-day PR turnaround to 2.4 days), current state uses 3–4 parallel Claude Code sessions for exploration and rapid development."
    },
    {
      "title": "Claude Code vs Cursor vs GitHub Copilot in 2026 - Jacar",
      "url": "https://jacar.es/en/claude-code-vs-cursor-vs-github-copilot-in-2026-a-comparison-with-measured-tasks/",
      "date": "2026-07-07",
      "type": "case-study",
      "added": "2026-07-21",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical testing of three agentic tools on five real platform-team tasks (80k LOC Go+TypeScript); measured time-to-merge and quality: Claude Code excels on multi-file refactors and security review; Cursor on interactive workflows; shows task-specific tool strengths."
    },
    {
      "title": "GPT-5.6 Can Code Autonomously — But AI-Generated Code Has 2.7x More Defects. Here's What QA Teams Must Do Now.",
      "url": "https://launchweld.com/blog/2026-07-07/gpt-5-6-agentic-coding-defect-rate-qa-response/",
      "date": "2026-07-07",
      "type": "product-ga",
      "added": "2026-07-21",
      "superseded_by": null,
      "window": null,
      "explanation": "OpenAI GPT-5.6 Sol launch enables full autonomous agentic coding (repo inspection, code generation, test creation, deployment without human-in-the-loop), representing major capability milestone for the practice; dual signal: capability advancement and quality concerns (2.7x defect rate vs. human code)."
    },
    {
      "title": "The Shift to Agentic AI: Evidence from Codex",
      "url": "https://arxiv.org/abs/2606.26959",
      "date": "2026-06-25",
      "type": "adoption-metric",
      "added": "2026-07-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed analysis of 5x user growth on OpenAI Codex, 26.6% using multi-step skills, and task complexity escalation (70%+ >1 hour by May 2026) across occupations, confirming workflow sophistication beyond single-task use."
    },
    {
      "title": "Claude Code Speeds Prototyping, But Not Shipping",
      "url": "https://www.linkedin.com/posts/priyanshsinghal_aiengineering-buildinpublic-genai-activity-7475034911784136705-OyVB",
      "date": "2026-06-23",
      "type": "opinion",
      "added": "2026-07-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Practitioner distinguishes exploration speed from shipping speed: 18-hour RAG pipeline prototype, 4-day debugging for hallucinations and token limits; net loss for delivery speed; frames agentic coding as exploration accelerator, not shipping accelerator."
    },
    {
      "title": "How Long Does Idea to Production Actually Take with Claude Code?",
      "url": "https://www.buildthisnow.com/fr/blog/real-examples/idea-to-production-time-data",
      "date": "2026-06-22",
      "type": "case-study",
      "added": "2026-07-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Realistic timelines for SaaS development: prototype 4-72 hours, MVP 1-4 weeks, production 4-8 weeks; OnboardingHub example (55 days, 727 commits); documents the 'last 20% tax' on polish, security, edge cases."
    },
    {
      "title": "Agentic Coding in Mid-2026: What Changed and How I Actually Use It - DEV Community",
      "url": "https://dev.to/pavelespitia/agentic-coding-in-mid-2026-what-changed-and-how-i-actually-use-it-2kl4",
      "date": "2026-06-19",
      "type": "case-study",
      "added": "2026-07-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Practitioner documents workflow across four real projects (refactoring, testing, module work) with specific technique guidance: full specification outperforms iteration; xhigh effort efficient for multi-step work; 95% accuracy still insufficient for security-sensitive code."
    },
    {
      "title": "The Commonplace — 2026-06-15 • Buttondown",
      "url": "https://buttondown.com/workforcefutures/archive/the-commonplace-2026-06-15/",
      "date": "2026-06-15",
      "type": "opinion",
      "added": "2026-07-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Research digest synthesizing preregistered RCT evidence: human-in-the-loop with gated autonomy reduces critical failures from 72% to 16%; unconstrained multi-agent systems underperform; process design steer outcomes more than capability."
    },
    {
      "title": "Anthropic's 2026 Agentic Coding Report: The Delegation Gap and Eight Trends Reshaping Software Development",
      "url": "https://faq.com.tw/en/developer-tools/2026-06-11-anthropic-2026-agentic-coding-report-en/",
      "date": "2026-06-11",
      "type": "industry-report",
      "added": "2026-07-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Identifies the 'delegation gap' (60% AI usage, 0-20% full autonomy), 8 structural trends shaping agentic workflows, and case studies (Rakuten 7+ hours, CRED 89% adoption) showing exploration-mode gains."
    },
    {
      "title": "How Claude Code is used in practice - Anthropic",
      "url": "https://www.anthropic.com/research/claude-code-expertise",
      "date": "2026-06-11",
      "type": "product-ga",
      "added": "2026-07-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Official research on 400k Claude Code sessions showing division of labor (70% planning, 80% execution decisions by Claude), task distribution across occupations, and success rates confirming exploration-stage adoption."
    },
    {
      "title": "The Missing Layer in Agentic AI: Why Evaluation Is the Next Enterprise Platform",
      "url": "https://www.newabdullah.com/posts/state-of-the-art-agentic-ai-evaluation-end-to-end/",
      "date": "2026-06-09",
      "type": "opinion",
      "added": "2026-07-07",
      "superseded_by": null,
      "window": null,
      "explanation": "Technical synthesis of frontier-lab agentic evaluation research identifying evaluation maturity as lagging deployment—enterprises can observe activity and outcomes but cannot reliably grade agent behavior, establishing governance as binding constraint."
    },
    {
      "title": "Gartner Hype Cycle for Agentic AI: 17% Deployed, 42% Plan Within 12 Months",
      "url": "https://1password.com/blog/gartner-agentic-ai-governance/",
      "date": "2026-06-05",
      "type": "industry-report",
      "added": "2026-06-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Gartner April 2026: only 17% deployed but 42% plan within 12 months (most aggressive adoption curve); fully autonomous agents not ready for enterprise; human oversight essential — validates governance and readiness gaps."
    },
    {
      "title": "Anthropic 2026 Agentic Coding Trends Report: Multi-File Editing Surge and Deployment Acceleration",
      "url": "https://note.com/snake_dragon/n/n216df10517e3?hl=en",
      "date": "2026-06-03",
      "type": "adoption-metric",
      "added": "2026-06-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Multi-file editing 34%→78%, session duration 4→23min with 47 avg tool calls; Rakuten processed 12.5M LOC in 7-hour unattended run — confirms exploration-mode adoption and capability maturation."
    },
    {
      "title": "Scout: GitHub's Exploratory Research Agent for Token Anomaly Investigation",
      "url": "https://github.github.com/gh-aw/blog/2026-06-02-agent-of-the-day/",
      "date": "2026-06-02",
      "type": "case-study",
      "added": "2026-06-09",
      "superseded_by": null,
      "window": null,
      "explanation": "GitHub deployed Scout agent for exploratory investigation; identified token consumption spike via 8 turns of reasoning in 8.1 minutes across 61 network requests — demonstrates exploration-task autonomy at scale."
    },
    {
      "title": "The Great Agentic Hangover: Production Failures and the Cost of Absent Guardrails",
      "url": "https://futurium.ec.europa.eu/pt/apply-ai-alliance/community-content/great-agentic-hangover",
      "date": "2026-06-02",
      "type": "opinion",
      "added": "2026-06-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Critical assessment of Q2 2026 agentic failures: PocketOS database wipe in <10 seconds, token billing DDoS, API exfiltration — documents why exploration/prototyping requires guardrails and why autonomy remains constrained."
    },
    {
      "title": "Vercel AI Gateway Production Metrics: 30% Agent-Driven Deployments",
      "url": "https://www.originbrief.app/en/reports/developer-tools-platforms/2026-06-01/weekly",
      "date": "2026-06-01",
      "type": "adoption-metric",
      "added": "2026-06-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Real infrastructure metrics from 200K+ teams: 30% of deployments initiated by agentic tools (1000% growth in 6 months), Claude Code 75% market share — demonstrates infrastructure-scale adoption."
    },
    {
      "title": "Cursor Developer Habits Report: Velocity Surge Coupled with Governance Challenge (36.3% Unreviewed Changes)",
      "url": "https://mnemehq.com/insights/cursor-developer-habits-report-governance-infrastructure/",
      "date": "2026-05-30",
      "type": "adoption-metric",
      "added": "2026-06-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Cursor telemetry: 8.6K LOC/week (2.4x increase), but agent changes committed without diff review jumped 7%→36.3% in 5 months — signals productivity gains coupled with governance infrastructure gaps."
    },
    {
      "title": "OpenAI Workspace Agents GA: Enterprise Validation with Case Studies",
      "url": "https://www.bighatgroup.com/blog/codex-weekly-2026-05-29/",
      "date": "2026-05-29",
      "type": "product-ga",
      "added": "2026-06-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Workspace Agents reached GA (May 21) with admin governance. Virgin Atlantic/Endava case studies show production deployment of agentic workflows—signals ecosystem readiness for organizational adoption."
    },
    {
      "title": "What Breaks When LLMs Code: 547 Safety Failures Across Deployed Tools",
      "url": "https://arxiv.org/abs/2605.30777v1",
      "date": "2026-05-29",
      "type": "research-paper",
      "added": "2026-06-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Systematic study of 547 operational safety failures across 16K+ GitHub issues; 59% high/critical severity—reveals real failure modes (constraint violations, destructive ops) that benchmarks miss, critical for tier assessment."
    },
    {
      "title": "GitHub Research Shows AI Coding Agents No Longer Niche: 2026 Development Rewriting Processes",
      "url": "https://masonailab.com/insights/github-ai-coding-agent-adoption-research-2026/",
      "date": "2026-05-27",
      "type": "research-paper",
      "added": "2026-06-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed arXiv analysis of 932K agent-authored GitHub PRs across five tools: effectiveness varies by task (82% docs, 66% features) — strongest empirical evidence of exploration-task success rates."
    },
    {
      "title": "Stack Overflow Pulse: Agentic Adoption Nearly Doubled to 59%, Barriers Remain",
      "url": "https://stackoverflow.blog/2026/05/27/agents-on-a-leash-agentic-ai-remains-mostly-monitored-at-work/",
      "date": "2026-05-27",
      "type": "adoption-metric",
      "added": "2026-06-09",
      "superseded_by": null,
      "window": null,
      "explanation": "1,100-developer survey: 59% adoption (up from 31%), but 60% block unapproved changes and 68% prefer single-agent—strong validation that exploration/prototyping under human oversight is primary use case."
    },
    {
      "title": "Anthropic's 2026 Agentic Coding Report: 8 Trends Now",
      "url": "https://byteiota.com/anthropics-2026-agentic-coding-report-8-trends-now/",
      "date": "2026-05-23",
      "type": "adoption-metric",
      "added": "2026-05-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Enterprise deployments: Rakuten 79% delivery-time reduction, TELUS 500k hours saved, Zapier 89% adoption; context engineering replaces prompt engineering as adoption lever."
    },
    {
      "title": "The autonomy trap: What AI agents and vibe coding are actually doing to production systems",
      "url": "https://kz.kursiv.media/en/opinions/the-autonomy-trap-what-ai-agents-and-vibe-coding-are-actually-doing-to-production-systems/",
      "date": "2026-05-22",
      "type": "opinion",
      "added": "2026-05-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Critical assessment: six incident postmortems show boundary-less agents cause cascading failures; chain probability (85% per-step = 20% end-to-end over 10 steps) explains silent autonomy failures."
    },
    {
      "title": "Agentic Coding in 2026: A Practical Guide for Big Code | Sourcegraph",
      "url": "https://sourcegraph.com/blog/agentic-coding",
      "date": "2026-05-21",
      "type": "opinion",
      "added": "2026-05-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Practitioner architecture guide: 80% problem—agents miss cross-cutting dependencies outside context window; context gathering determines 80% of output quality at scale."
    },
    {
      "title": "Week 20 · May 11–15, 2026 - Claude Code Docs",
      "url": "https://code.claude.com/docs/en/whats-new/2026-w20",
      "date": "2026-05-19",
      "type": "product-ga",
      "added": "2026-05-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Agent View and /goal features enable autonomous multi-turn agentic workflows for exploration; developers set completion conditions and Claude works autonomously across turns."
    },
    {
      "title": "Researchers Disclose Multiple Security Flaws in Anthropic's Claude",
      "url": "https://letsdatascience.com/news/researchers-disclose-multiple-security-flaws-in-anthropics-c-08a9a4c4",
      "date": "2026-05-15",
      "type": "news-coverage",
      "added": "2026-05-26",
      "superseded_by": null,
      "window": null,
      "explanation": "May 2026 CVE disclosures (CVE-2025-59536, CVE-2026-21852): untrusted repo configs trigger RCE and API key theft; TrustFall class issues enable silent MCP privilege escalation."
    },
    {
      "title": "Claude Code #1 Tool: MCP Enterprise Roadmap 2026 - Springvanta",
      "url": "https://springvanta.com/blog/claude-code-number-one-mcp-enterprise-roadmap-2026",
      "date": "2026-05-14",
      "type": "industry-report",
      "added": "2026-05-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Market leadership confirmed: 71% of 900+ engineers surveyed prefer Claude Code first; 75% among startups; fastest adoption among tools with years of head start."
    },
    {
      "title": "Four silent failures landed on the day Claude Code v2.1.141 shipped its silent-failure fixes",
      "url": "https://gist.github.com/yurukusa/959f64e23a0ed53b464270f3c284924f",
      "date": "2026-05-14",
      "type": "opinion",
      "added": "2026-05-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Systemic claim/reality divergence: VSCode streaming stalls, Edit tool silently truncates, worktrees accumulate 36GB orphaned data, MCP servers leave 3.1GB resident RAM uncleaned."
    },
    {
      "title": "Claude Updates by Anthropic - May 2026 - Releasebot",
      "url": "https://releasebot.io/updates/anthropic/claude",
      "date": "2026-05-11",
      "type": "product-ga",
      "added": "2026-05-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Claude Platform on AWS now GA with managed agents, code execution, MCP connector—first-party infrastructure for enterprise agentic exploration at scale."
    },
    {
      "title": "Anthropic Claude Code Quality Drop: What Actually Happened",
      "url": "https://devtoolpicks.com/blog/anthropic-claude-code-quality-fix-postmortem-2026",
      "date": "2026-05-11",
      "type": "case-study",
      "added": "2026-05-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Detailed postmortem audit of 6,852 sessions documenting measurable quality degradation (March–April 2026) and recovery; provides evidence of adoption scale and reliability gaps requiring governance."
    },
    {
      "title": "Claude Code Is Killing Software Engineering (2026)",
      "url": "https://www.devflokers.com/blog/claude-code-destroying-software-engineering-2026",
      "date": "2026-05-07",
      "type": "opinion",
      "added": "2026-05-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Critical assessment with named sources (AMD Sr Director, TrustedSec CEO): specific failure metrics (47% quality drop, 52% vulnerability rate) and failure modes (laziness, incomplete reasoning)."
    },
    {
      "title": "The Productivity Impact of Coding Agents: The Real Truth - HCL GUVI",
      "url": "https://www.guvi.in/blog/the-productivity-impact-of-coding-agents/",
      "date": "2026-05-06",
      "type": "adoption-metric",
      "added": "2026-05-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Large-scale adoption (84% developer use, 41% of code authored by AI) with wide variance: 3.6 hrs/week median savings but 41% see 'little effect' and 19% report slowdown—signals practice maturity with limits."
    },
    {
      "title": "Mise en Place for Agentic Coding: Deliberate Preparation as Context Engineering Methodology",
      "url": "https://arxiv.org/abs/2605.05400",
      "date": "2026-05-06",
      "type": "research-paper",
      "added": "2026-05-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Systematic preparation methodology for agentic coding: contextual grounding, collaborative specification, task decomposition applied to rapid parallel exploration in hackathon environment."
    },
    {
      "title": "AI Coding Agents: Market Developments, Risks and Developer Takeup",
      "url": "https://acquinox.capital/insights/gen-ai-and-ai-agents/ai-coding-agents-market-developments-risks-and-developer-takeup",
      "date": "2026-05-06",
      "type": "adoption-metric",
      "added": "2026-05-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Investment analyst assessment: Claude Code $1B+ ARR, GitHub Copilot 20M users, 55.8% faster completion, quality trade-offs—signals market maturity at significant scale."
    },
    {
      "title": "Higher usage limits for Claude and a compute deal with SpaceX",
      "url": "https://www.anthropic.com/news/higher-limits-spacex",
      "date": "2026-05-05",
      "type": "product-ga",
      "added": "2026-05-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Doubled Claude Code rate limits and removed peak-hour throttling (May 2026), backed by 300+ MW SpaceX compute: removes mid-session interruption constraint that blocked multi-file exploration tasks."
    },
    {
      "title": "After Pushback, Amazon Rolls Out Claude Code, Codex to All Employees",
      "url": "https://www.businessinsider.com/amazon-claude-code-codex-all-employees-after-pushback-2026-5",
      "date": "2026-05-04",
      "type": "adoption-metric",
      "added": "2026-05-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Named enterprise (Amazon) formally deploys Claude Code company-wide May 2026; signals exploration-stage maturity at FAANG scale with both external tools and internal Kiro infrastructure."
    },
    {
      "title": "Agentic Coding Is Not a Trap: I Answered the Viral HN Post With My Own Production Logs",
      "url": "https://dev.to/jtorchia/agentic-coding-is-not-a-trap-i-answered-the-viral-hn-post-with-my-own-production-logs-33d9",
      "date": "2026-05-04",
      "type": "case-study",
      "added": "2026-05-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Practitioner's 6-week production usage analysis: agent effectiveness dependent on specification clarity; validates exploration/prototyping patterns with quantified productivity metrics."
    },
    {
      "title": "Your coding agent is under-specified",
      "url": "https://hsaghir.com/blog/2026-05-02-under-specified-coding-agent/",
      "date": "2026-05-02",
      "type": "opinion",
      "added": "2026-05-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Fundamental structural limitation: agents on under-specified natural language prompts vs. precise specs; Alibaba SWE-CI found 75%+ of agents showed accelerating regression despite passing each task."
    },
    {
      "title": "Vibe Coding vs. Agentic Engineering: The Dual Evolution of Software Development in 2026",
      "url": "https://solafide.ca/blog/2026-05-vibe-coding-vs-agentic-engineering",
      "date": "2026-05-01",
      "type": "opinion",
      "added": "2026-05-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Balanced analysis of vibe coding as legitimate exploration paradigm with adoption (Karpathy Feb 2025 framing, COED 'Word of the Year'), workflow reality, and limitations (1.7x bugs, 30% review tax)."
    },
    {
      "title": "The 17% Comprehension Gap: Anthropic's Own Study Says AI Coding Tools Are Quietly Deskilling Your Engineers",
      "url": "https://getburnrate.io/blog/ai-coding-skill-atrophy-comprehension-debt",
      "date": "2026-04-29",
      "type": "adoption-metric",
      "added": "2026-05-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Critical assessment: hidden comprehension debt and production failures in high-AI-adoption teams with specific cost multipliers (26-32x subscription bill) from verified Anthropic RCT data."
    },
    {
      "title": "GPT-5.5 Launches: 82.7% Terminal-Bench, Agentic Coding Era",
      "url": "https://aiautomationglobal.com/blog/openai-gpt-5-5-agentic-coding-terminal-bench-2026",
      "date": "2026-04-24",
      "type": "product-ga",
      "added": "2026-04-28",
      "superseded_by": null,
      "window": null,
      "explanation": "OpenAI GPT-5.5 GA achieving 82.7% Terminal-Bench SOTA, 88.7% SWE-bench, with internal evidence of autonomous 20-hour engineering tasks in single runs and 60% hallucination reduction—crossing agentic coding into production default."
    },
    {
      "title": "Anthropic Admits 3 Bugs Killed Claude Code For 50 Days",
      "url": "https://theplanettools.ai/blog/anthropic-claude-code-quality-fix-three-bugs-postmortem-april-2026",
      "date": "2026-04-23",
      "type": "case-study",
      "added": "2026-04-28",
      "superseded_by": null,
      "window": null,
      "explanation": "Official Anthropic postmortem: three production bugs (reasoning-effort flip, cache bug, verbosity constraint) degraded Claude Code March–April 2026, causing cost spikes ($42K) and context loss; reveals QA gaps despite multiple testing gates during peak adoption."
    },
    {
      "title": "AI Agentic Coding Reliability: 2026 Guide - Tech Daily Care",
      "url": "https://techdailycare.com/tech-3/",
      "date": "2026-04-23",
      "type": "adoption-metric",
      "added": "2026-04-28",
      "superseded_by": null,
      "window": null,
      "explanation": "MSR 2026 empirical study of 11,771 real PRs documents Claude Opus 4.7 at 87.6% SWE-Bench yet only 24% autonomous completion on complex tasks, with 70-90% failure as complexity rises—dual signal for tier classification."
    },
    {
      "title": "275M GitHub Commits, Real API Bills, Tax-Filing Agents - Buttondown",
      "url": "https://buttondown.com/theheartbeat/archive/20260421/",
      "date": "2026-04-21",
      "type": "adoption-metric",
      "added": "2026-04-28",
      "superseded_by": null,
      "window": null,
      "explanation": "Population-level adoption signal: 275M weekly AI-authored commits on GitHub signal agentic coding past novelty into infrastructure; bottleneck shifted from agent output to human review/orchestration pipeline."
    },
    {
      "title": "Agentic Coding Trends 2026 | Anthropic Report Key Insights",
      "url": "https://www.libertify.com/interactive-library/agentic-coding-trends-2026-anthropic-report/",
      "date": "2026-04-18",
      "type": "industry-report",
      "added": "2026-04-28",
      "superseded_by": null,
      "window": null,
      "explanation": "Anthropic 2026 industry report: developers use AI in 60% of work but fully delegate only 0-20%, TELUS saved 500K+ hours, 27% of AI work represents new (unattempted) tasks—establishes agentic coding as standard business practice with measurable ROI."
    },
    {
      "title": "Introducing Claude Opus 4.7 - Anthropic",
      "url": "https://www.anthropic.com/news/claude-opus-4-7",
      "date": "2026-04-16",
      "type": "product-ga",
      "added": "2026-04-28",
      "superseded_by": null,
      "window": null,
      "explanation": "Claude Opus 4.7 GA showing 10-point SWE-Bench Pro jump (64.3%), CursorBench 70%, and customer testimonials from Devin, Hex, Notion confirming long-running exploration task capability with new xhigh effort and task budgets."
    },
    {
      "title": "Week 15 · April 6–10, 2026 - Claude Code Docs",
      "url": "https://code.claude.com/docs/en/whats-new/2026-w15",
      "date": "2026-04-13",
      "type": "product-ga",
      "added": "2026-04-14",
      "superseded_by": null,
      "window": null,
      "explanation": "Official release notes documenting major feature additions for agentic workflows: Ultraplan cloud-based planning, Monitor tool for reactive agents, /autofix-pr integration, and /team-onboarding automation."
    },
    {
      "title": "Claude Code Review 2026: The Tool That Flipped the Dev Market in 8 Months",
      "url": "https://neuriflux.com/en/blog/claude-code-review-2026",
      "date": "2026-04-10",
      "type": "adoption-metric",
      "added": "2026-04-14",
      "superseded_by": null,
      "window": null,
      "explanation": "Independent review documenting Claude Code's 46% 'most loved' rating (vs. Copilot 9%, Cursor 19%), 95% first-try correctness, 67% code quality wins, SWE-bench 80.8%—market leadership achieved in 8 months."
    },
    {
      "title": "Is Claude Code getting worse? March 2026 data | TokenCost",
      "url": "https://tokencost.app/blog/claude-code-getting-worse-april-2026",
      "date": "2026-04-08",
      "type": "case-study",
      "added": "2026-04-14",
      "superseded_by": null,
      "window": null,
      "explanation": "AMD Sr Director Stella Laurenzo's forensic analysis of 6,852 sessions (234,760 tool calls): thinking depth -67%, reads per edit -70%, full-file rewrites +127%, monthly costs +122x—documents reliability regression March 8–April 2026."
    },
    {
      "title": "Beyond Human-Readable: Rethinking Software Engineering Conventions for the Agentic Development Era",
      "url": "https://papers.cool/arxiv/2604.07502",
      "date": "2026-04-08",
      "type": "research-paper",
      "added": "2026-04-14",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed research on codebase conventions for agentic navigation: semantic density optimization reveals compression paradox (67% cost increase despite 17% token reduction), informs exploration-stage architecture."
    },
    {
      "title": "Agentic Coding for Humanists",
      "url": "https://computationalhistory.substack.com/p/agentic-coding-for-humanists",
      "date": "2026-04-02",
      "type": "case-study",
      "added": "2026-04-14",
      "superseded_by": null,
      "window": null,
      "explanation": "Cameron Blevins (computational historian) successfully built interactive data visualization with Claude Code and Gemini without coding—demonstrates accessibility of agentic exploration for domain experts."
    },
    {
      "title": "Agentic Coding in Production: The Q1 2026 Landscape | Zylos",
      "url": "https://zylos.ai/research/2026-04-02-agentic-coding-production-q1-2026-landscape",
      "date": "2026-04-02",
      "type": "industry-report",
      "added": "2026-04-14",
      "superseded_by": null,
      "window": null,
      "explanation": "Q1 2026 landscape analysis marking inflection point: autonomous agentic coding moved from experimental to mainstream. Three drivers: capability thresholds (Opus 4.5/GPT-5.1), tooling maturity (IDE integration), governance consolidation."
    },
    {
      "title": "Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time",
      "url": "https://papers.cool/arxiv/2604.00917",
      "date": "2026-04-01",
      "type": "research-paper",
      "added": "2026-04-14",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed study of 110K OSS pull requests: agent activity increasing but agent-generated code shows higher churn rates and longer survival times vs. human code—quality/maintenance trade-off documented."
    },
    {
      "title": "Claude Code Now Writes 4% of All GitHub Commits - and Growing 8% Weekly",
      "url": "https://megaoneai.com/blog/claude-code-github-commits-2026/",
      "date": "2026-03-31",
      "type": "adoption-metric",
      "added": "2026-04-14",
      "superseded_by": null,
      "window": null,
      "explanation": "SemiAnalysis report: 4% of public GitHub commits (135K/day, 1M+ repos), 8% week-over-week growth projecting 20%+ by end of 2026. Caveat: 90% in low-star repos (throwaway/exploratory projects)."
    },
    {
      "title": "Agentic AI is here — and it is changing how scientific software is built and used",
      "url": "https://blog.sintef.com/digital-en/agentic-ai-is-here-and-it-is-changing-how-scientific-software-is-built-and-used/",
      "date": "2026-03-27",
      "type": "case-study",
      "added": "2026-03-31",
      "superseded_by": null,
      "window": null,
      "explanation": "Norwegian research institute deployed agentic coding to refactor 500K-LOC reservoir simulator (OPM Flow), completing multi-file class restructuring and GPU pattern changes autonomously with 99.9% accuracy."
    },
    {
      "title": "AI agents are getting more capable, but reliability is lagging",
      "url": "https://fortune.com/2026/03/24/ai-agents-are-getting-more-capable-but-reliability-is-lagging-narayanan-kapoor/",
      "date": "2026-03-24",
      "type": "industry-report",
      "added": "2026-03-31",
      "superseded_by": null,
      "window": null,
      "explanation": "Princeton study evaluating Claude Opus 4.5, GPT-5.2, Gemini 3 Pro: reliability improved at half the accuracy improvement rate; critical gaps in calibration (52%) and safety (25%) signal adoption requires human oversight."
    },
    {
      "title": "Anthropic 2026 Agentic Coding Trends Report Explained | 8 Trends, TELUS & Rakuten Case Studies",
      "url": "https://www.paperclipped.de/en/blog/anthropic-agentic-coding-trends-report-2026/",
      "date": "2026-03-22",
      "type": "industry-report",
      "added": "2026-03-31",
      "superseded_by": null,
      "window": null,
      "explanation": "TELUS saved 500K+ hours across 57,000 employees with 13,000 custom AI solutions; Rakuten reduced feature delivery time 79% (24 days → 5 days) using exploration-mode agentic coding with human oversight."
    },
    {
      "title": "Coding Agents are Effective Long-Context Processors",
      "url": "https://arxiv.org/abs/2603.20432",
      "date": "2026-03-20",
      "type": "research-paper",
      "added": "2026-03-31",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed empirical study: coding agents outperform state-of-the-art by 17.3% on long-context reasoning and RAG tasks via native tool proficiency and file-system navigation."
    },
    {
      "title": "How We Stopped Our AI Agents From Getting Dumber Mid-Session",
      "url": "https://www.monkeyrun.com/blog/how-we-stopped-our-ai-agents-from-getting-dumber",
      "date": "2026-03-07",
      "type": "case-study",
      "added": "2026-03-31",
      "superseded_by": null,
      "window": null,
      "explanation": "MonkeyRun team quantified context-rot degradation patterns and documented four architectural solutions (fresh subagents, atomic tasks, upfront design phase, token budgeting) maintaining quality across extended exploration work."
    },
    {
      "title": "AI Tooling for Software Engineers in 2026",
      "url": "https://newsletter.pragmaticengineer.com/p/ai-tooling-2026",
      "date": "2026-03-03",
      "type": "adoption-metric",
      "added": "2026-03-31",
      "superseded_by": null,
      "window": null,
      "explanation": "Survey of 900+ professional software engineers: 55% regularly use AI agents, 63.5% adoption among senior engineers, Claude Code now #1 agentic coding tool (achieved in 8 months since launch)."
    },
    {
      "title": "Devin, the AI Engineer: Review, Testing & Limitations in 2026 | Idlen",
      "url": "https://www.idlen.io/blog/devin-ai-engineer-review-limits-2026/",
      "date": "2026-03-03",
      "type": "case-study",
      "added": "2026-03-31",
      "superseded_by": null,
      "window": null,
      "explanation": "Independent real-world testing across 5 production codebases shows task-specific success rates: well-defined bug fixes 78%, test writing 82%, refactoring 45%, architecture 15%—documents persistent limitations of fully autonomous agents."
    },
    {
      "title": "The train has left the station: Agentic AI and the future of social science research",
      "url": "https://www.brookings.edu/articles/the-train-has-left-the-station-agentic-ai-and-the-future-of-social-science-research/",
      "date": "2026-03-03",
      "type": "case-study",
      "added": "2026-03-31",
      "superseded_by": null,
      "window": null,
      "explanation": "Brookings researchers used Claude Code to transform minimal implementations into full R packages in one day, produce 20-page analysis in under an hour, and build pilot study infrastructure iteratively."
    },
    {
      "title": "Only 11% of Companies Use AI Agents in Production - imarch.dev",
      "url": "https://imarch.dev/en/blog/agentic-reality-check-deloitte-2026/",
      "date": "2026-03-02",
      "type": "industry-report",
      "added": "2026-03-31",
      "superseded_by": null,
      "window": null,
      "explanation": "Deloitte's 2025 'Agentic Reality Check' identifies three systemic adoption barriers (legacy systems, data architecture, governance), names four successful production case studies (Toyota, HPE, Dell, Moderna), projects 40% of agentic projects will fail by 2027."
    },
    {
      "title": "Devin and Claude Code in SRE Practice (DevinとClaude Code、SREの現場で使い倒してみた件)",
      "url": "https://speakerdeck.com/karia/devintoclaude-code-srenoxian-chang-deshi-idao-sitemitajian",
      "date": "2026-02-28",
      "type": "case-study",
      "added": "2026-03-31",
      "superseded_by": null,
      "window": null,
      "explanation": "Filmarks/TSUMIKI SRE team documented 6+ production exploration use cases: CI/CD setup, Git workflow automation, infrastructure-as-code generation, achieving autonomous first commits and PR creation."
    },
    {
      "title": "最近更新 - Devin Docs",
      "url": "https://docs.devin.ai/zh/release-notes/overview",
      "date": "2026-02-27",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Devin 2.2 general availability: 3x faster startup, new lifecycle-aware UI, desktop end-to-end testing support, enhanced Slack/Linear integrations; signals continued platform maturation for autonomous exploration workflows."
    },
    {
      "title": "AI Agents 2026: Overhyped, 40% Will Fail | byteiota",
      "url": "https://byteiota.com/ai-agents-2026-overhyped-40-will-fail/",
      "date": "2026-02-26",
      "type": "industry-report",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Gartner predicts 40% of agentic AI projects will be canceled by 2027; analysis of 847 deployments shows 76% fail in production; developers spend 67% more time debugging AI-generated code."
    },
    {
      "title": "DevIn AI Review 2026: Is the $500/Month AI Developer Worth It?",
      "url": "https://runaicode.ai/devin-ai-review-2026-is-the-500-month-ai-developer-worth-it/",
      "date": "2026-02-26",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Senior developer 3-month review: Devin strong on repetitive tasks (70-80% quality) but struggles with architecture, ambiguous requirements, and complex debugging; exploration gains limited by autonomy ceiling."
    },
    {
      "title": "Claude Code Changed How We Ship Software — Our First 90 Days",
      "url": "https://www.codercops.com/blog/claude-code-changed-how-we-ship-software",
      "date": "2026-02-22",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "CODERCOPS 90-day production deployment: 45% sprint velocity increase, 50% faster PR merges, 56% faster bug resolution, 28-point test coverage gain; highlights exploration-to-sustained productivity workflow."
    },
    {
      "title": "Sola Fide - Anthropic's 2026 Agentic Coding Trends Report",
      "url": "https://solafide.ca/blog/anthropic-2026-agentic-coding-trends-reshaping-software-development",
      "date": "2026-02-10",
      "type": "industry-report",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Rakuten reduced time-to-market by 79% (24→5 days) using Claude Code; TELUS shipped code 30% faster with 500k+ total hours saved; demonstrates sustained ROI from exploration-stage agentic coding."
    },
    {
      "title": "TXSL #6: The State of AI Agents in 2026: Incompetent and Dangerous",
      "url": "https://txsling.com/2026/02/01/txsl-6-the-state-of-ai-agents-in-2026-incompetent-and-dangerous/",
      "date": "2026-02-01",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Carnegie Mellon study: AI agents complete only 8.6-30% of simulated workplace tasks; Anthropic research finds models may act maliciously under threat; critical reality check on autonomy claims."
    },
    {
      "title": "Are bugs and incidents inevitable with AI coding agents?",
      "url": "https://stackoverflow.blog/2026/01/28/are-bugs-and-incidents-inevitable-with-ai-coding-agents/",
      "date": "2026-01-28",
      "type": "adoption-metric",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "GitHub analysis of 470 repos finds AI-generated code creates 1.7x more bugs than humans, with 75% higher logic errors, 1.5-2x more security issues, indicating quality barriers to adoption."
    },
    {
      "title": "Agentic Much? Adoption of Coding Agents on GitHub",
      "url": "https://arxiv.org/abs/2601.18341",
      "date": "2026-01-26",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Large-scale empirical study of 129,134 GitHub projects finds 15.85%-22.60% adoption of coding agents, with broad usage across project maturity, major organizations (Microsoft 18.57%), and diverse languages."
    },
    {
      "title": "Wilson Lin on FastRender: a browser built by thousands of parallel agents",
      "url": "https://simonwillison.net/2026/Jan/23/fastrender/",
      "date": "2026-01-23",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Cursor research experiment: 2,000 parallel agents built a functional browser in one week, making ~30,000 commits, demonstrating exploration-stage agentic coding at massive scale."
    },
    {
      "title": "Claude Code just got updated with one of the most-requested user features",
      "url": "https://novalogiq.com/2026/01/16/claude-code-just-got-updated-with-one-of-the-most-requested-user-features/",
      "date": "2026-01-16",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Claude Code MCP Tool Search feature reduces context usage by 85% (from ~134k to ~5k tokens) and improves accuracy from 49% to 74% for Opus 4, addressing bloat and enabling richer tool sets in agentic workflows."
    },
    {
      "title": "Infosys to use AI coder Devin across company",
      "url": "https://www.indiatoday.in/technology/news/story/infosys-to-use-ai-coder-devin-across-company-sparks-fear-of-job-loss-for-freshers-and-junior-developers-2849088-2026-01-09",
      "date": "2026-01-09",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Major IT services firm Infosys deployed Devin company-wide for COBOL migrations and JCP modernization, reporting material productivity gains from six months of experimentation."
    },
    {
      "title": "Spec-Driven Development: Agentic Coding at FAANG Scale",
      "url": "https://bretthamlin.com/briefing/2026-01/2026-01-09-spec-driven-development-agentic-coding-at-faang-sc/",
      "date": "2026-01-09",
      "type": "conference-talk",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Amazon Kiro agentic IDE presentation on spec-driven development replacing vibe coding, using EARS format and MCP server integration to improve code quality and reliability at scale."
    },
    {
      "title": "Agentic Coding in 2025",
      "url": "https://ethanswan.com/feed/2025/12/24/agentic-coding-in-2025/",
      "date": "2025-12-24",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Practitioner year-end analysis of Claude Code and Cursor adoption; identifies persistent weaknesses (lack of proactive refactoring, architectural vision), categorizes developer attitudes by adoption and skepticism."
    },
    {
      "title": "57% Have AI Agents in Production. Only 6% Are Ready",
      "url": "https://codingwithroby.substack.com/p/57-have-ai-agents-in-production-only",
      "date": "2025-12-23",
      "type": "adoption-metric",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Survey data: 57% of companies deploy agents; only 6% have required infrastructure; 60% of multi-agent systems fail; cites JPMorgan 95% faster research retrieval, revealing adoption-readiness gap."
    },
    {
      "title": "Devin: The Autonomous Engineer (Or Is It?) | MMNTM",
      "url": "https://www.mmntm.net/articles/devin-deep-dive",
      "date": "2025-12-14",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Critical analysis of Devin architecture (13.86% unassisted SWE-bench, 15% real-world success); documents happy path (migrations, linting, tests) and failure modes (design, ambiguity, business context)."
    },
    {
      "title": "Claude Moves to the Darkside: What a Rogue Coding Agent Could Do Inside Your Org",
      "url": "https://zenity.io/blog/current-events/claude-moves-to-the-darkside-what-a-rogue-coding-agent-could-do-inside-your-org",
      "date": "2025-11-15",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Security incident: GTG-1002 weaponized Claude Code for cyberattack across 30+ orgs, 80% autonomous execution via prompt engineering; critical failure signal for governance and trustworthiness barriers."
    },
    {
      "title": "Introduction to agentic coding",
      "url": "https://www.claude.com/blog/introduction-to-agentic-coding",
      "date": "2025-10-30",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Anthropic vendor blog formally introduces agentic coding concept, promotes Claude Code capabilities (autonomy, whole-codebase understanding, multi-step execution), signaling market positioning and tool maturity."
    },
    {
      "title": "Claude Code on the Web: How UK Businesses Can Deploy Code Anywhere",
      "url": "https://www.grow-fast.co.uk/blog/claude-code-web-deploy-anywhere-uk-businesses-october-2025",
      "date": "2025-10-21",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Claude Code web launch removes local-development constraint, enabling exploration coding from any browser/device with cloud-hosted execution; platform expansion signal for exploration-stage adoption."
    },
    {
      "title": "Rebuilding Devin for Claude Sonnet 4.5: Lessons and Challenges",
      "url": "https://cognition.ai/blog/devin-sonnet-4-5-lessons-and-challenges",
      "date": "2025-09-29",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "Devin rebuilt for Sonnet 4.5 showing 2× speed improvement and 12% better Junior Developer eval scores; technical detail on model behavior changes enabling agentic improvements."
    },
    {
      "title": "The Productivity Paradox of AI Coding Assistants",
      "url": "https://www.cerbos.dev/blog/productivity-paradox-of-ai-coding-assistants",
      "date": "2025-09-12",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "Internal team case study showing split adoption; cites METR slowdown data and Stack Overflow survey (only 33% trust accuracy, 66% frustrated by 'almost right' code), revealing adoption–trust gap."
    },
    {
      "title": "Claude in enterprise: case studies of successful AI deployments",
      "url": "https://www.datastudios.org/post/claude-in-enterprise-case-studies-of-successful-ai-deployments",
      "date": "2025-09-02",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "Enterprise Claude Code deployments (TELUS, Brex, Snowflake, CRED) showing measurable ROI: 500k+ hours saved, 75% auto-processing, 2× developer velocity, demonstrating production-grade adoption."
    },
    {
      "title": "AI Coding Agents Are Infiltrating the Corporate World",
      "url": "https://www.businessinsider.com/ai-coding-agents-adoption-top-tools-2025-8",
      "date": "2025-08-05",
      "type": "adoption-metric",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "Jellyfish survey (400+ companies): agentic AI adoption jumped 50% to 82% (Dec 2024–May 2025); AI-powered code reviews rose 39% to 76%, signaling rapid enterprise adoption of exploration tools."
    },
    {
      "title": "Measuring the Impact of Early-2025 AI on Experienced Developers",
      "url": "https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/",
      "date": "2025-07-10",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "Randomized controlled trial (16 experienced developers, 246 real issues) finding AI tools made developers 19% slower despite subjective belief in productivity, contradicting claimed benefits."
    },
    {
      "title": "What I learned trying seven coding agents",
      "url": "https://www.understandingai.org/p/what-i-learned-trying-seven-coding",
      "date": "2025-06-27",
      "type": "news-coverage",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Comparative testing of seven coding agents (Bolt, Replit, Lovable, Windsurf, Codex, Cursor, Claude Code) on real prototyping task; Claude Code performed best with less trial-and-error than alternatives."
    },
    {
      "title": "Claude Code issue #2423: Session-Destroying Failures from Compaction Timeouts",
      "url": "https://github.com/anthropics/claude-code/issues/2423",
      "date": "2025-06-21",
      "type": "news-coverage",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Critical bug report: Claude Code sessions destroyed by compaction lockups and tool_use 400 crashes on Windows, documenting stability issues and session recovery gaps in exploration tool."
    },
    {
      "title": "Coding agents have crossed a chasm",
      "url": "https://blog.singleton.io/posts/2025-06-14-coding-agents-cross-a-chasm/",
      "date": "2025-06-14",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Practitioner analysis documenting personal shift to relying on Claude Code and Codex for small tasks and debugging, with workflow integration and specific productivity gains in exploration mode."
    },
    {
      "title": "Iterate and refine with the Cursor Agent",
      "url": "https://sylverstudios.dev/blog/2025/04/30/ai-for-refinement.html",
      "date": "2025-04-30",
      "type": "news-coverage",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Practitioner workflow demonstrating Cursor Agent Mode for UI refinement and test polishing using visual comparison (screenshots vs. design), showing iterative exploration-mode strengths."
    },
    {
      "title": "AI is Still a Long Way From Directly Replacing Programmers",
      "url": "https://markpelf.com/2721/ai-isnt-ready-to-directly-replace-programmers-april-2025/",
      "date": "2025-04-16",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Critical assessment documenting debugging failures, hallucinations, multi-file task limitations, and developer resistance; AI tools save time but fall far short of autonomy for serious work."
    },
    {
      "title": "Cognition AI launches revamped Devin 2.0",
      "url": "https://siliconangle.com/2025/04/03/cognition-ai-launches-revamped-coding-assistant-devin-2-0-much-lower-starting-price/",
      "date": "2025-04-03",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Devin 2.0 launch with pricing drop (from $500/month to $20/month starting tier), addressing adoption barriers and expanding market accessibility for exploration-mode agentic coding."
    },
    {
      "title": "Claude Code saved us 97% of the work — then failed utterly",
      "url": "https://www.thoughtworks.com/en-gb/insights/blog/generative-ai/claude-code-codeconcise-experiment",
      "date": "2025-03-10",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "Thoughtworks real-world case: Claude Code added language support to CodeConcise in half a day (97% effort savings) but exhibited unreliable performance in subsequent attempts, confirming exploration-mode brittleness."
    },
    {
      "title": "Anthropic's Claude Code tool had a bug that 'bricked' some systems",
      "url": "https://techcrunch.com/2025/03/06/anthropics-claude-code-tool-had-a-bug-that-bricked-some-systems/",
      "date": "2025-03-06",
      "type": "news-coverage",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "Claude Code auto-update bug corrupted system files and bricked workstations, signaling tool maturity gaps and security risks in early-stage agentic coding products."
    },
    {
      "title": "10. Test-Case Driven...",
      "url": "https://benhouston3d.com/blog/agentic-coding-best-practices",
      "date": "2025-03-05",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "Practitioner guidance from six weeks building with mycoder.ai: adapting codebases for AI (flat structures, minimal indirection, compile-time validation) improves both AI and human developer efficiency in exploration workflows."
    },
    {
      "title": "The world's 'first AI software engineer' isn't living up to expectations",
      "url": "https://www.itpro.com/software/development/the-worlds-first-ai-software-engineer-isnt-living-up-to-expectations-cognition-ais-devin-assistant-was-touted-as-a-game-changer-for-developers-but-so-far-its-fumbling-tasks-and-struggling-to-compete-with-human-workers",
      "date": "2025-02-04",
      "type": "news-coverage",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "Independent evaluation (Answer.AI): Devin achieved 15% success rate (3/20 tasks), with failures including hallucinated features and technical dead-ends, demonstrating persistent autonomy limitations."
    },
    {
      "title": "A Look at Three Major Failure Patterns: Spatial Mismatch, Temporal Forgetfulness, and Reinventing the Wheel",
      "url": "https://yage.ai/agentic-memory-en.html",
      "date": "2025-01-01",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "Critical analysis of exploration-stage agentic tools (Cursor, WindSurf) failing on larger codebases due to context window limits, identifying spatial mismatch, temporal forgetfulness, and code duplication as core failure patterns."
    },
    {
      "title": "Real-World Case Studies | Vibe Coding Guide - Master AI ...",
      "url": "https://moinsen-dev.github.io/claude_code_vibe_coding_guide/examples/real-world-examples",
      "date": "2025-01-01",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "TaskFlow SaaS case study: Claude Code built MVP in 3 weeks (1 week ahead) with 95% test coverage and 1000+ concurrent user capacity, demonstrating exploration-to-production success with human-in-the-loop oversight."
    },
    {
      "title": "Devin is now generally available",
      "url": "https://cognition.ai/blog/devin-generally-available",
      "date": "2024-12-10",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Product GA announcement: Devin AI engineer now available at $500/month with Slack/IDE integration; includes concrete use cases (frontend bugs, draft PRs, refactoring) and open source deployment examples."
    },
    {
      "title": "The 70% problem: Hard truths about AI-assisted coding",
      "url": "https://addyo.substack.com/p/the-70-problem-hard-truths-about",
      "date": "2024-12-04",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Practitioner analysis by Google engineer identifying '70% problem': non-engineers hit wall after initial progress; quality not improved despite productivity gains; best use as prototyping accelerator, not democratization tool."
    },
    {
      "title": "Independent Coding Agents Aren't Ready",
      "url": "https://www.chrismdp.com/independent-coding-agents-arent-ready/",
      "date": "2024-11-01",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Infrastructure and security assessment: deploying Claude Code in secure environments requires substantial expertise; AI agents with broad permissions pose unpredictable security risks, indicating deployment maturity gaps."
    },
    {
      "title": "Claude 3.5 Sonnet on GitHub Copilot",
      "url": "https://www.anthropic.com/news/github-copilot",
      "date": "2024-10-25",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Major vendor integration: Claude 3.5 Sonnet rolling out on GitHub Copilot to 100M+ developers; 93.7% HumanEval score demonstrates ecosystem maturity and accessibility for exploration coding workflows."
    },
    {
      "title": "How SailPoint uses Anthropic's Claude on Amazon Bedrock to automatically generate TypeScript code for SaaS connectors",
      "url": "https://aws.amazon.com/blogs/machine-learning/how-sailpoint-uses-anthropics-claude-on-amazon-bedrock-to-automatically-generate-typescript-code-for-saas-connectors/",
      "date": "2024-10-16",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Enterprise deployment: SailPoint uses Claude on Amazon Bedrock to automatically generate TypeScript code for SaaS connectors, with specific technical details on API integration and pagination handling."
    },
    {
      "title": "DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models",
      "url": "https://arxiv.org/abs/2410.07331",
      "date": "2024-10-09",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Peer-reviewed benchmark (EMNLP 2024) showing current LLMs achieve only 30.5% accuracy on agent-based data science code generation, documenting persistent capability gaps in agentic coding."
    },
    {
      "title": "Businesses thought AI could write code. Now they're playing whack-a-mole",
      "url": "https://techtonicshifts.blog/2024/09/17/businesses-thought-ai-could-write-code-now-theyre-playing-whack-a-mole-with-outages/",
      "date": "2024-09-17",
      "type": "news-coverage",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "News coverage of AI coding tool downsides: ChatGPT 65.2% accuracy, GitHub Copilot 46.3%, Amazon CodeWhisperer 31.1%; 38% of developers report inaccuracies and over half face security issues from generated code."
    },
    {
      "title": "Earning Agentic (and LangChain) Complexity - ISE Developer Blog",
      "url": "https://devblogs.microsoft.com/ise/earning-agentic-complexity/",
      "date": "2024-09-12",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "Microsoft ISE team warning that agentic frameworks are brittle, hard to debug, and risky for production; customers preferring explicit chained designs over multi-agent patterns, highlighting deployment barriers."
    },
    {
      "title": "Devin September '24 Product Update - Cognition",
      "url": "https://cognition.ai/blog/sept-24-product-update",
      "date": "2024-09-05",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "Devin product update reporting 80% reduction in task completion time, MultiDevin for parallel delegation, and VPC deployment for enterprise use, demonstrating continued tool evolution in Q3 2024."
    },
    {
      "title": "CursorLens: Open-source analytics dashboard for Cursor IDE",
      "url": "https://github.com/HamedMP/CursorLens",
      "date": "2024-08-16",
      "type": "significant-repo",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "Open-source tool with 386 stars for monitoring and analytics of AI-assisted coding sessions in Cursor IDE, showing ecosystem maturation and real-world usage instrumentation for exploration workflows."
    },
    {
      "title": "Claude Engineer - Build with Sonnet 3.5",
      "url": "https://llmindset.co.uk/posts/2024/07/claude-engineer-build-sonnet3-5/",
      "date": "2024-07-23",
      "type": "tutorial",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "Hands-on tutorial on Claude Engineer 2.0 (1200 lines of Python, 25% prompts) demonstrating iterative code modification and self-improvement for exploration tasks, with documented successes and challenges."
    },
    {
      "title": "claude-artifacts-react: Deploy React code from Claude Artifacts",
      "url": "https://github.com/risonsimon/claude-artifacts-react",
      "date": "2024-07-02",
      "type": "significant-repo",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "Open-source tool with 49 stars enabling one-click deployment of React code from Claude Artifacts to Vercel/Cloudflare, demonstrating community tooling for Claude-based exploration and prototyping."
    },
    {
      "title": "How AppMap Navie solved the SWE bench AI coding challenge",
      "url": "https://dev.to/appmap/how-appmap-navie-solved-the-swe-bench-ai-coding-challenge-20an",
      "date": "2024-06-25",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "AppMap Navie case study achieving 14.6% on SWE-Bench with semi-agentic architecture (RAG + planning + code editing), demonstrating cost-efficient alternative to full agentic approaches."
    },
    {
      "title": "Using AI-Based Coding Assistants in Practice: State of Affairs, Perceptions, and Ways Forward",
      "url": "https://arxiv.org/abs/2406.07765v2",
      "date": "2024-06-11",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "Survey of 481 programmers documenting adoption of AI assistants across development stages, showing test generation and documentation as preferred use cases, with trust and context limits as barriers."
    },
    {
      "title": "Stack Overflow's 2024 Developer Survey: AI adoption and trust gap",
      "url": "https://stackoverflow.co/company/press/archive/stack-overflow-2024-developer-survey-gap-between-ai-use-trust/",
      "date": "2024-05-29",
      "type": "adoption-metric",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "Stack Overflow survey of 65,000+ developers: 76% use or plan to use AI tools, 81% report productivity gains, but only 43% trust accuracy, showing rapid adoption with persistent skepticism."
    },
    {
      "title": "SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities",
      "url": "https://arxiv.org/html/2511.05459v3",
      "date": "2024-05-22",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "Comprehensive benchmark (SWE-Compass) evaluating agentic coding across 8 task types, 10 languages, and 2,000 real GitHub issues, revealing persistent gaps and hierarchies in LLM capability."
    },
    {
      "title": "The Pulse #90: Devin reversing ambitious claims",
      "url": "https://newsletter.pragmaticengineer.com/p/the-pulse-90",
      "date": "2024-04-18",
      "type": "news-coverage",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "Critical analysis exposing Devin's marketing overstatement; open-source alternatives (SWE-agent 12.3%, AutoCodeRover ~16%) emerging as competitive, signaling market correction and rapid commoditization."
    },
    {
      "title": "AI Research Roundup 24.04.05: SWE-Agent benchmarking",
      "url": "https://patmcguinness.substack.com/p/ai-research-roundup-240405",
      "date": "2024-04-05",
      "type": "news-coverage",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "SWE-Agent (open-source) achieving 12.3% on SWE-Bench with novel Agent-Computer Interface; base Claude Opus scores 11.07%, highlighting rapid ecosystem commoditization and accessible alternatives."
    },
    {
      "title": "Is the \"AI developer\" a threat to jobs – or a marketing stunt?",
      "url": "https://newsletter.pragmaticengineer.com/p/is-the-ai-developera-threat-to-jobs",
      "date": "2024-03-19",
      "type": "news-coverage",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Pragmatic Engineer analysis of Devin's 13.86% SWE-Bench performance and agentic coding tool claims, highlighting skepticism about autonomous capabilities versus marketing narratives."
    },
    {
      "title": "A look at (the demo of) Devin, the AI-powered software engineer",
      "url": "https://www.fikisipi.com/blog/why-control-matters-in-production-ai-agents",
      "date": "2024-03-12",
      "type": "news-coverage",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Critical analysis of Devin's demo capabilities and benchmark scoring, noting skepticism about cherry-picked demonstrations and long distance to practical deployment."
    },
    {
      "title": "If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents",
      "url": "https://openreview.net/forum?id=8dmNOD9hbq",
      "date": "2024-03-11",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "ICLR 2024 workshop survey analyzing how code training enables LLMs to serve as intelligent agents, providing academic foundation for understanding agentic coding capabilities."
    },
    {
      "title": "claude-3-opus-code.md",
      "url": "https://gist.github.com/simonw/2002e2b56a97053bd9302a34e0b83074",
      "date": "2024-03-07",
      "type": "tutorial",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Practical example of Claude 3 Opus autonomously modifying JavaScript code for file uploads and image conversion, demonstrating agentic capabilities in exploration tasks."
    },
    {
      "title": "Code Aesthetics with Agentic Reward Feedback",
      "url": "https://arxiv.org/html/2510.23272v1",
      "date": "2024-03-06",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Microsoft Research multi-agent system for evaluating and improving code aesthetics, with AesCoder-4B surpassing GPT-4o performance on visual code generation benchmarks."
    },
    {
      "title": "Practical Considerations for Agentic LLM Systems",
      "url": "https://arxiv.org/html/2412.04093v1",
      "date": "2024-03-06",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "University of Edinburgh research bridging academic and industry perspectives on building robust LLM agents, covering planning, memory, tools, and control flow for real-world deployment."
    }
  ],
  "tierHistory": [
    {
      "tier": "research",
      "from": "2024-03-01",
      "to": "2024-07-01"
    },
    {
      "tier": "bleeding-edge",
      "from": "2024-07-01",
      "to": "2026-01-01"
    },
    {
      "tier": "leading-edge",
      "from": "2026-01-01",
      "to": null
    }
  ],
  "trendHistory": [
    {
      "trend": "steady",
      "blockerType": null,
      "from": "2026-09-26",
      "to": null
    }
  ],
  "description": "AI agents that autonomously write, run, and iterate on code for proofs of concept and exploratory development. Includes tools like Claude Code, Cursor agent mode, and Devin used for throwaway prototypes and spikes; distinct from production agentic coding which requires CI/CD integration and review workflows.",
  "overview": "Agentic coding for exploration and prototyping hands an AI agent a loose brief and lets it write, run and revise code until a spike or proof of concept stands up, trading durability for speed of learning. It is a leading-edge practice and steady: developers now use it routinely, and teams at forward-leaning organisations report genuine gains, but only inside deliberate guardrails. What holds it back is the boundary it depends on. Throwaway code keeps leaking into production before anyone understands it, agents claim successes they have not earned, and security flaws in the tools themselves remain open. Until independent analysts judge it safe to roll out and the line between exploration and production holds, its readiness stays contested rather than settled.",
  "currentLandscape": "Agent use is now routine among professional developers. JetBrains' Developer Ecosystem Survey 2026 reports 90% weekly agent usage, with Claude Code leading. Enterprise deployments have moved from pilots of a few hundred seats to tens of thousands of seats. An arXiv study of enterprise harness governance notes Accenture training about 30,000 professionals on Claude and Claude Code. The same study notes Cognizant rolling out Claude for up to 350,000 employees. Ai-jirei separately catalogues 28 verified enterprise deployments of Claude Code.\n\nOpenAI offers the deepest internal case. The Pragmatic Engineer reports that its finance, recruitment and legal teams went from roughly 0% to 90% Codex usage in four months. Overall usage rose from 60% to 90% between April and May, which OpenAI attributes to better handling of long-running tasks. Its Perf Factory monitors production and launches Codex agents to fix performance issues. Leaders say PRs and code review need to be rethought. Dependence is now complete: minor Codex outages trigger immediate internal alerts.\n\nPrototyping is where agents deliver most reliably. MEV reports shipping a proof of concept in 72 hours with AI agents. A structured review of 60 papers finds that programming agents cut the cost of repetitive coding and help users without professional experience to complete prototypes. The same review finds that efficiency gains depend on task size, project familiarity, quality standards and the cost of manual review. Atlassian's prototype-to-production account shows where the effort returns: agent output outpaced team comprehension and needed production hardening.\n\nThe line between throwaway and production code is eroding. New Relic's 2026 State of AI Coding survey of 200 US technology decision-makers finds that 88% have written vibe coding into formal production policies. Only 5% restrict it to non-production use, and none ban it. New Relic calls the resulting gap \"agent debt\": 94% rate AI-generated code higher quality at review. Yet 78% report more incidents after deployment. Some 82% had at least one production failure tied to AI-generated code in six months.\n\nDebugging agent output has become a major cost. A Coleman Parkes poll of 300 UK and US developers, commissioned by the debugging vendor Undo, finds that 93% have hit AI hallucinations that led to an incorrect diagnosis. In the same poll, 94% lose productivity analysing AI-generated code at least monthly. Some 35% say such code reaches production before teams understand what it does. Respondents average 16.9 hours a week debugging, roughly twice the time they spend writing code.\n\nMeasured productivity gains remain contested. One randomised trial found experienced developers 19% slower when using AI tools. LaunchWeld reports that AI-generated code carries 2.7x more defects. SpruceTech's analysis of 304,362 verified AI-authored commits concludes that code delivery is no longer the bottleneck. In its analysis, ownership of the resulting code is.\n\nCapability ceilings appear on harder and more specialised work. Agentry reports that coding agents plateau below 45% on realistic benchmarks such as SWE-Bench Pro. SWE-bench Science finds that agents struggle with scientific software engineering and miss domain-specific constraints. A study of how agents use documentation finds they lean on instruction files rather than reading API references systematically. Harness choice also matters: Marmelab's audit of 246 repositories cites a run of one model through 8 harnesses on 25 tasks, in which success ranged from 68% to 88%.\n\nCost control is becoming a governance problem. An arXiv study that emulates a 10,000-seat enterprise finds that cache-safe model routing recovers 14-21% of model spend, or $3.3M-$5.0M a year at Anthropic's list prices. The same study flags lock-in, because Claude Code can only call Anthropic models. Separate arXiv work on the agentic development lifecycle identifies verification, retries and escalation as recurring costs that often exceed token spend.\n\nSecurity exposure grows with local autonomy. The Hacker News reports that malicious .git configs can make Claude, Codex, Cursor and other agents run attacker code. It separately reports flaws in Claude Code and Gemini CLI that let a GitHub issue reach CI workflow secrets.\n\nOrganisational readiness is the remaining blocker. VESS Labs decided against adopting Devin after a one-week trial. Forrester and Anaconda find that 88% of AI agent pilots never reach production. Broader adoption beyond bounded exploration depends on three things: verification infrastructure, clear ownership of agent output, and security controls that match the autonomy agents already have.",
  "history": "- **2024-Q1:** Research papers on agentic coding (code aesthetics evaluation, code-LLM agency theory) published alongside product launches (Devin). Tools demonstrated 13-15% autonomous code task completion. Practitioner adoption accelerated experimentally; skepticism about production readiness grew in parallel.\n\n- **2024-Q2:** Market reality correction: open-source tools (SWE-Agent, AutoCodeRover) matched or exceeded Devin's benchmarks; critical analysis exposed marketing overstatement. Academic rigor increased with SWE-Compass and other evaluation frameworks revealing persistent capability gaps. Developer surveys showed 76% adoption but only 43% accuracy trust. Semi-agentic architectures (AppMap Navie, Claude Code) emerged as practical exploration-stage solutions with cost advantages.\n\n- **2024-Q3:** Tool ecosystem matured: Devin reported 80% speed improvements, Cursor gained community analytics tooling, Claude Artifacts spawned deployment facilitators. Quality metrics hardened—ChatGPT 65.2% accuracy, Copilot 46.3%, CodeWhisperer 31.1%; surveys documented 38% developer accuracy concerns and 50%+ company security issues. Microsoft ISE team warned agentic frameworks were brittle and unsuitable for production. Consensus shifted: agentic coding viable for exploration when constrained by human oversight and guardrails; risk management became the defining concern.\n\n- **2024-Q4:** Product ecosystem matured: Claude 3.5 reached 100M+ developers on GitHub Copilot; SailPoint deployed Claude for TypeScript generation on Amazon Bedrock; Devin reached GA at $500/month. Reality-check evidence consolidated: DA-Code showed 30.5% accuracy on data science tasks; Google engineer Addy Osmani documented the \"70% problem\" (70% quick, 30% wall); infrastructure assessments detailed security and deployment risks. Practitioner consensus hardened: agentic coding effective for exploration within guardrails (tight loops, human oversight, clear scope), but brittleness and accuracy gaps remained unresolved. The practice matured from research to pragmatism, not to solved-problem.\n\n- **2025-Q1:** Tool maturity concerns emerged: Claude Code auto-update bug bricked systems; independent evaluations (Answer.AI) showed Devin at 15% success rate (3/20 tasks); technical analysis identified core failure patterns (spatial mismatch, temporal forgetfulness, code duplication). Yet constrained workflows succeeded: Thoughtworks saved 97% effort with Claude Code when adding language support (then faced reliability issues); TaskFlow SaaS built production MVP in 3 weeks with 95% test coverage. Practitioner guidance emphasized codebase adaptation and human oversight as prerequisites. The paradox hardened: agentic coding produced real results within guardrails but remained brittle at scale.\n\n- **2025-Q2:** Practitioner adoption accelerated; Devin 2.0 launched with pricing drop (to $20/month) addressing cost barriers. Comparative testing (seven agents on prototyping task) ranked Claude Code first but showed all tools required trial-and-error iteration. Tool stability issues persisted: Claude Code exhibited instruction-following drift and session-destroying compaction bugs on Windows. Critical assessment documented why autonomy remained unreachable: debugging failures, hallucinations, multi-file task limitations. Market inflection: messaging shifted from \"autonomous engineers\" to rapid iteration within guardrails. Agentic exploration now positioned as time-saving accelerator proven to work under tight human supervision, but tool maturity gaps remained blockers for enterprise adoption.\n\n- **2025-Q3:** Enterprise adoption reached critical mass (82% across 400+ companies, up from 50% in Dec 2024). Named deployments with measurable ROI emerged (TELUS 500k+ hours saved, Brex 75% auto-processing, CRED 2× velocity). Yet independent RCT (METR) contradicted productivity claims, finding experienced developers 19% slower with AI tools despite subjective belief in gains. Tool evolution continued (Cursor 1.4 improvements, Devin Sonnet rebuild with 2× speed, 12% better evals). Adoption–trust gap widened: only 33% trusted accuracy; 66% frustrated by \"almost right\" code. Landscape shifted from capability debate to ROI uncertainty and operational reliability concerns.\n\n- **2025-Q4:** Platform expansion and consolidation: Claude Code reached web availability (October), enabling exploration from any browser; Devin reached 25% of Cognition's internal PRs (December). Yet credibility crisis emerged: state-sponsored threat actor weaponized Claude Code in 80%+ automated cyberattack across 30+ organizations (November, Zenity disclosure). Adoption-readiness gap hardened: 57% of companies in production but only 6% with required infrastructure; 60% of multi-agent systems failing. Practitioner sentiment bifurcated into four camps (vibecoders, sweet spot, dubious, artisans) based on adoption–skepticism matrix. Consensus shifted: exploration-mode agentic coding works under constraints (small tasks, tight loops, human oversight) but security governance and readiness gaps remain critical blockers.\n\n- **2026-Jan:** Population-level adoption reached empirical confirmation: 15.85%-22.60% of 129k GitHub projects employ agentic tools (arxiv); Microsoft shows 18.57% adoption within enterprise. Tool ecosystem matured with infrastructure improvements: Claude Code introduced MCP Tool Search reducing context from 134k to 5k tokens (85% reduction) and improving accuracy from 49% to 74%; Amazon released Kiro agentic IDE promoting spec-driven development over vibe coding. Enterprise deployment accelerated: Infosys rolled out Devin company-wide for COBOL migration and modernization (January), reporting material productivity gains. Credibility tension hardened: parallel GitHub analysis of 470 repos documented AI code creating 1.7x more bugs than humans, with 75% higher logic errors and 1.5-2x security issues, directly contradicting productivity claims. Practitioner exploration and prototyping workflows mainstream but quality concerns prevent broader production adoption.\n\n- **2026-Feb:** Real-world deployment metrics emerged validating exploration-stage ROI: CODERCOPS reported 90-day production gains (45% sprint velocity, 50% faster PR merges, 56% faster bug resolution), while Rakuten and TELUS achieved 79% time-to-market reduction and 30% faster shipping via Claude Code. Platform evolution continued: Devin 2.2 released with 3x faster startup and desktop testing support. Yet credibility crisis deepened: Carnegie Mellon research showed agents complete only 8.6-30% of workplace tasks; Gartner predicted 40% of agentic AI projects will be canceled by 2027; analysis of 847 deployments showed 76% fail in production. The window captured the field's fundamental tension: named orgs achieved measurable exploration gains within guardrails, while independent research documented persistent autonomy failures and governance gaps. Practitioner sentiment remained bifurcated: exploration-stage agentic coding proven viable for constrained tasks but unreliable for autonomous operation.\n\n- **2026-Mar:** Population-level adoption confirmed at 55% of professional engineers regularly using agents, with Claude Code reaching #1 status in just 8 months; task-specific success rates documented: bug fixes 78%, test writing 82%, refactoring 45%, architecture 15%. Named deployments validated the constrained-use model: SINTEF completed autonomous multi-file refactoring of a 500K-LOC reservoir simulator with 99.9% accuracy; TELUS and Rakuten case studies (Anthropic trends report) reconfirmed 79% delivery-time reduction and 500K+ hours saved under human oversight; Brookings researchers used Claude Code to build full R packages and 20-page analyses in hours. A Princeton study found reliability improvements lagging capability growth at half the rate, with critical calibration gaps (52%); context-rot patterns were quantified with architectural countermeasures emerging (fresh subagents, atomic tasks, token budgeting). Deloitte's 2025 reality check confirmed 11% production deployment and projected 40% project failure by 2027. The defining tension held: exploration-mode agents deliver real productivity within guardrails, but autonomy gaps and reliability lags prevent broader deployment.\n\n- **2026-Apr (Early):** Market inflection confirmed: Claude Code achieved 46% 'most loved' ranking (vs. Copilot 9%, Cursor 19%), writing 4% of all GitHub commits (135K/day across 1M+ repos, 8% week-over-week growth). JetBrains survey shows 41% market share vs. Copilot's 38%, with 91% satisfaction and 54 NPS. Yet critical reliability degradation documented: AMD Sr Director's forensic analysis of 6,852 sessions revealed thinking depth collapsed 67-75% (late February–March), reads per edit dropped 70%, full-file rewrites doubled, monthly Bedrock costs exploded 122x ($345→$42K)—traced to Opus 4.6 adaptive thinking misconfiguration. Peer-reviewed research on 110K OSS PRs confirms agent activity scaling but documents higher code churn vs. human-authored code. Leading-edge capability (Ultraplan cloud planning, Monitor tool for reactive agents) shipped alongside hidden reliability regression, exemplifying the practice's defining tension: exploration-mode agentic coding delivers measurable productivity gains within guardrails, but tool maturity gaps (thinking depth, cost control, code quality) create operational hazards that constrain adoption to disciplined teams with oversight. New accessibility signal: domain experts (computational historians) successfully prototype without coding, expanding addressable population beyond software engineers.\n\n- **2026-Apr (Mid-Late):** Competitive capability inflection and reliability crisis converged. OpenAI GPT-5.5 released April 23 (82.7% Terminal-Bench SOTA, 88.7% SWE-bench, 20-hour autonomous engineering runs), crossing agentic coding from research into production default—yet same window revealed critical tool maturity gaps. Claude Opus 4.7 released April 16 with major benchmarks (SWE-Bench Pro 64.3%, CursorBench 70%), new xhigh reasoning effort tier, and task budgets for cost management. Yet April 23 postmortem documented three production bugs had degraded Claude Code quality March 4–April 20: reasoning-effort default flipped (34 days undetected), cache bug caused context loss and repetition, verbosity constraint cut coding-eval 3%—user costs spiked $345→$42K/month. Real-world empirical study (MSR 2026) on 11,771 PRs found Claude Opus 4.7 leads at 87.6% SWE-Bench but only 24% autonomous completion on complex tasks (70-90% failure as complexity rises), establishing critical gap between benchmark performance and real-world autonomy. Population-level adoption signal solidified: 275M weekly AI-authored commits on GitHub (agentic baseline past novelty into infrastructure), with bottleneck shifted from agent output velocity to human review and orchestration pipeline capacity. Anthropic 2026 industry report documented TELUS 500K+ hour savings, Augment Code 4-to-2-week project compression, yet showed developers use AI in 60% of work but fully delegate only 0-20% (collaborative model, not autonomous). The period crystallized the practice's maturity: capability crossed into production-ready (GPT-5.5 architectural reasoning, Opus 4.7 reasoning control), deployment reached infrastructure scale (275M commits/week), yet reliability gaps (tool bugs, benchmark-reality gap, limited autonomy on complex tasks) remain binding constraints on broader adoption beyond disciplined teams with strong oversight cultures.\n\n- **2026-May:** Infrastructure expansion and enterprise normalization continued. Anthropic doubled Claude Code rate limits and removed peak-hour throttling (May 6), backed by a 300+ MW SpaceX compute deal; Claude Platform on AWS reached GA (May 11) with managed agents, code execution, and MCP integration. Amazon formally deployed Claude Code company-wide (May 4) and investment analysts valued it at $1B+ ARR alongside GitHub Copilot at 20M users. Anthropic's 2026 agentic coding report corroborated enterprise gains at named scale (Rakuten 79% delivery-time reduction, TELUS 500K hours saved, Zapier 89% adoption) while quantifying the cascade-failure risk: 85% per-step reliability yields only 20% end-to-end success over a 10-step workflow. Claude Code Agent View and /goal features shipped (Week 20), enabling developers to set completion conditions and run multi-turn autonomous exploration workflows unattended. Security vulnerabilities surfaced: CVE-2025-59536 and CVE-2026-21852 expose RCE and API key theft via untrusted repository configs in Claude Code. Adoption paradoxes deepened: HCL GUVI documented 84% developer adoption (41% of code AI-authored) but with median savings of 3.6 hrs/week and 19% of developers reporting slowdowns. A \"mise en place\" methodology—deliberate context engineering before agent runs—emerged as the key preparation pattern for reliable exploration workflows.\n\n- **2026-Jun:** Adoption metrics crossed new thresholds: Stack Overflow survey shows 59% of developers using agentic tools (up from 31%), Vercel AI Gateway reports 30% of all deployments initiated by agentic tools (1,000% growth in six months, Claude Code at 75% share), and a peer-reviewed study of 932K agent-authored PRs finds adoption in new projects more than doubled versus the prior cohort. Anthropic's updated trends data shows multi-file editing surged from 34% to 78% of sessions with Rakuten processing 12.5M LOC in a 7-hour unattended run. GitHub's Scout agent (exploratory token-anomaly investigation, 61 network requests in 8 minutes) provides a named enterprise proof-point for exploration-task autonomy. The governance gap sharpened: Cursor telemetry shows unreviewed agent changes jumping from 7% to 36.3% of merges in five months, and a systematic study of 547 operational safety failures (59% high/critical severity) documents the real failure modes that constrained exploration still misses.\n\n- **2026-Jul:** Peer-reviewed Codex analysis confirmed 5x user growth with 26.6% employing multi-step skills and sessions extending to 70%+ over one hour — the clearest large-scale signal of workflow sophistication. Anthropic's 400K-session study quantified the exploration division of labor (humans 70% of planning, agents 80% of execution), while practitioner evidence sharpened the exploration-vs-shipping boundary: one case documented an 18-hour RAG prototype followed by four days of debugging, confirming agentic coding as an exploration accelerator but not a delivery shortcut. A large-scale GitHub study (25,264 agentic PRs, 2,361 repos) confirmed adoption concentrates in small teams (1-5 contributors) under single-human oversight, while named case studies validated exploration-stage velocity — MEV shipped a PoC in 72 hours via a 7-agent pipeline, and estie's 18-month workflow evolution cut PR turnaround from 9.6 to 2.4 days. Veracode's longitudinal testing found syntax correctness climbing to ~95% while security pass rates stayed flat at 45-55% regardless of model generation, and OpenAI's GPT-5.6 Sol launch extended autonomous coding capability while carrying a 2.7x higher defect rate than human code — reinforcing the exploration-vs-production quality gap.\n\n- **2026-Aug:** Adoption metrics kept climbing — Microsoft's CLI agent study found 24% more merged PRs among adopters and a JetBrains survey showed Claude Code preferred 46% to Copilot's 9% among senior developers (6x growth since April 2025) — while a wave of empirical studies sharpened the exploration-vs-production gap: Gartner found 88% of agent pilots never reach production, an AIDev PR study documented static-analysis warnings up 18% and cognitive complexity up 39%, and a 4,882-PR test-coverage analysis found 50.4% of agentic PRs ship with zero test changes. Analysts introduced \"agentic code debt\" as a named governance risk class for exploration-stage output that accumulates uncomprehended logic and hallucinated dependencies. Google shipped Antigravity with Gemini 3.7 Flash as a new agentic-IDE entrant, while SWE-Bench Pro data showed coding agents plateauing below 45% on realistic tasks and a subagent-delegation study found deep agentic search costs more and finds less than semantic search for repository exploration. A CI/CD security flaw let GitHub issues reach Claude Code and Gemini CLI workflow secrets, reinforcing governance-gap commentary around \"scope hallucination\" as agents exceed their intended exploration boundaries.\n\n- **2026-Sep:** Large-sample evidence reinforced both sides of the exploration debate. A JetBrains survey of 15,000+ developers found 90% weekly agent usage with Claude Code the #1 tool for 31% (double GitHub Copilot's rate), and Stanford SALT Lab's analysis of 250,000 real Claude conversations found humans setting direction while adapting outputs in ~75% of cases, with over half involving consequential or hard-to-undo work. Google's Antigravity Teamwork framework reached GA with concrete exploration wins (7 solved math theorems, an autonomous RISC-V simulator), and Atlassian documented a 15-hour prototype-to-production build via multi-agent orchestration — though the same case showed agents outpacing team comprehension and breaking in production, shifting the bottleneck from code generation to human review. Countervailing evidence persisted: a METR randomized controlled trial found experienced developers 19% slower with AI tools despite self-reporting 20% faster, VESS Labs published a structured rejection of Devin after a one-week trial citing session-tracking and visibility gaps, and new research diagnosed 45 distinct agent failure patterns in autonomous research tasks, with a missing \"metacognitive loop\" as the dominant root cause. Peer-reviewed research formalizing production deployment boundaries emerged. Happy Bhati's synthesis (arxiv:2609.04681) identified the Agentic SDLC Throughput Paradox: agents accelerate code generation (+180% commits) but constrain production releases (+30%) due to verification, review, and security bottlenecks; introduced Production-Qualified Change (PQC) units and Verification Tax concept (recurring evaluation/retry/escalation costs). University of Pittsburgh + Ejento AI developed governance framework for agentic technical debt and stochastic tax management, proposing five control layers (golden-set evaluation, tool schema contracts, model gateways, graduated autonomy, workflow redesign). Large-scale empirical studies quantified code-quality gaps: Liu et al. (peer-reviewed, 304K AI-authored commits across 6,275 repos) found 484,606 introduced issues with 24.2% never fixed, 41.1% security issues persisting to latest revision; n1n.ai study (424 trials, multi-file tasks) showed forcing structured proof output increased evidence 30x but did NOT reduce false-success rates (75% with receipts vs 73.6% baseline). Security governance urgency escalated: Manifold disclosed GitSpawn (8 CVEs across Claude Code, Cursor, Codex, Gemini CLI enabling pre-trust execution), with multiple unpatched as of Sept 1. GitHub shipped Copilot agentic autofix GA (bulk remediation of 25 findings per assignment), signaling ecosystem maturation beyond exploration into routine maintenance. Microsoft Research analysis of 13.5M Copilot sessions (95 trillion tokens) revealed 87% of LLM calls are agent-initiated, establishing production-scale exploration workload patterns. A synthesis of Anthropic (400K sessions), OpenAI telemetry, and METR RCT data confirmed 70.2% of users now run multi-hour autonomous tasks even as the RCT found experienced developers 19% slower on complex work despite gains on simple ones, and a 9-engineer exe.dev retrospective documented the organizational prerequisites (architectural discipline, integration testing, design-first workflows) that let high-trust teams safely eliminate secondary review. The month crystallized the field's core tension: exploration-stage agents are now mainstream (90% weekly adoption, $2B+ ARR at scale), capability is proven in bounded scopes, yet production deployment requires investment in verification infrastructure, security hardening, and governance models that are still emerging. Fresh evidence added nuance: a Coleman Parkes poll found 93% of developers hit hallucinations and 35% report agent code reaching production before being understood, while New Relic found 88% of leaders permit vibe coding in production policies (78% seeing more incidents). OpenAI's internal Codex adoption rose from ~0% to 90% in non-engineering teams within four months, and an audit of 246 repositories found harness choice alone swinging one model's success rate from 68% to 88%.",
  "historyEntries": [
    {
      "period": "2024-Q1",
      "text": "Research papers on agentic coding (code aesthetics evaluation, code-LLM agency theory) published alongside product launches (Devin). Tools demonstrated 13-15% autonomous code task completion. Practitioner adoption accelerated experimentally; skepticism about production readiness grew in parallel."
    },
    {
      "period": "2024-Q2",
      "text": "Market reality correction: open-source tools (SWE-Agent, AutoCodeRover) matched or exceeded Devin's benchmarks; critical analysis exposed marketing overstatement. Academic rigor increased with SWE-Compass and other evaluation frameworks revealing persistent capability gaps. Developer surveys showed 76% adoption but only 43% accuracy trust. Semi-agentic architectures (AppMap Navie, Claude Code) emerged as practical exploration-stage solutions with cost advantages."
    },
    {
      "period": "2024-Q3",
      "text": "Tool ecosystem matured: Devin reported 80% speed improvements, Cursor gained community analytics tooling, Claude Artifacts spawned deployment facilitators. Quality metrics hardened—ChatGPT 65.2% accuracy, Copilot 46.3%, CodeWhisperer 31.1%; surveys documented 38% developer accuracy concerns and 50%+ company security issues. Microsoft ISE team warned agentic frameworks were brittle and unsuitable for production. Consensus shifted: agentic coding viable for exploration when constrained by human oversight and guardrails; risk management became the defining concern."
    },
    {
      "period": "2024-Q4",
      "text": "Product ecosystem matured: Claude 3.5 reached 100M+ developers on GitHub Copilot; SailPoint deployed Claude for TypeScript generation on Amazon Bedrock; Devin reached GA at $500/month. Reality-check evidence consolidated: DA-Code showed 30.5% accuracy on data science tasks; Google engineer Addy Osmani documented the \"70% problem\" (70% quick, 30% wall); infrastructure assessments detailed security and deployment risks. Practitioner consensus hardened: agentic coding effective for exploration within guardrails (tight loops, human oversight, clear scope), but brittleness and accuracy gaps remained unresolved. The practice matured from research to pragmatism, not to solved-problem."
    },
    {
      "period": "2025-Q1",
      "text": "Tool maturity concerns emerged: Claude Code auto-update bug bricked systems; independent evaluations (Answer.AI) showed Devin at 15% success rate (3/20 tasks); technical analysis identified core failure patterns (spatial mismatch, temporal forgetfulness, code duplication). Yet constrained workflows succeeded: Thoughtworks saved 97% effort with Claude Code when adding language support (then faced reliability issues); TaskFlow SaaS built production MVP in 3 weeks with 95% test coverage. Practitioner guidance emphasized codebase adaptation and human oversight as prerequisites. The paradox hardened: agentic coding produced real results within guardrails but remained brittle at scale."
    },
    {
      "period": "2025-Q2",
      "text": "Practitioner adoption accelerated; Devin 2.0 launched with pricing drop (to $20/month) addressing cost barriers. Comparative testing (seven agents on prototyping task) ranked Claude Code first but showed all tools required trial-and-error iteration. Tool stability issues persisted: Claude Code exhibited instruction-following drift and session-destroying compaction bugs on Windows. Critical assessment documented why autonomy remained unreachable: debugging failures, hallucinations, multi-file task limitations. Market inflection: messaging shifted from \"autonomous engineers\" to rapid iteration within guardrails. Agentic exploration now positioned as time-saving accelerator proven to work under tight human supervision, but tool maturity gaps remained blockers for enterprise adoption."
    },
    {
      "period": "2025-Q3",
      "text": "Enterprise adoption reached critical mass (82% across 400+ companies, up from 50% in Dec 2024). Named deployments with measurable ROI emerged (TELUS 500k+ hours saved, Brex 75% auto-processing, CRED 2× velocity). Yet independent RCT (METR) contradicted productivity claims, finding experienced developers 19% slower with AI tools despite subjective belief in gains. Tool evolution continued (Cursor 1.4 improvements, Devin Sonnet rebuild with 2× speed, 12% better evals). Adoption–trust gap widened: only 33% trusted accuracy; 66% frustrated by \"almost right\" code. Landscape shifted from capability debate to ROI uncertainty and operational reliability concerns."
    },
    {
      "period": "2025-Q4",
      "text": "Platform expansion and consolidation: Claude Code reached web availability (October), enabling exploration from any browser; Devin reached 25% of Cognition's internal PRs (December). Yet credibility crisis emerged: state-sponsored threat actor weaponized Claude Code in 80%+ automated cyberattack across 30+ organizations (November, Zenity disclosure). Adoption-readiness gap hardened: 57% of companies in production but only 6% with required infrastructure; 60% of multi-agent systems failing. Practitioner sentiment bifurcated into four camps (vibecoders, sweet spot, dubious, artisans) based on adoption–skepticism matrix. Consensus shifted: exploration-mode agentic coding works under constraints (small tasks, tight loops, human oversight) but security governance and readiness gaps remain critical blockers."
    },
    {
      "period": "2026-Jan",
      "text": "Population-level adoption reached empirical confirmation: 15.85%-22.60% of 129k GitHub projects employ agentic tools (arxiv); Microsoft shows 18.57% adoption within enterprise. Tool ecosystem matured with infrastructure improvements: Claude Code introduced MCP Tool Search reducing context from 134k to 5k tokens (85% reduction) and improving accuracy from 49% to 74%; Amazon released Kiro agentic IDE promoting spec-driven development over vibe coding. Enterprise deployment accelerated: Infosys rolled out Devin company-wide for COBOL migration and modernization (January), reporting material productivity gains. Credibility tension hardened: parallel GitHub analysis of 470 repos documented AI code creating 1.7x more bugs than humans, with 75% higher logic errors and 1.5-2x security issues, directly contradicting productivity claims. Practitioner exploration and prototyping workflows mainstream but quality concerns prevent broader production adoption."
    },
    {
      "period": "2026-Feb",
      "text": "Real-world deployment metrics emerged validating exploration-stage ROI: CODERCOPS reported 90-day production gains (45% sprint velocity, 50% faster PR merges, 56% faster bug resolution), while Rakuten and TELUS achieved 79% time-to-market reduction and 30% faster shipping via Claude Code. Platform evolution continued: Devin 2.2 released with 3x faster startup and desktop testing support. Yet credibility crisis deepened: Carnegie Mellon research showed agents complete only 8.6-30% of workplace tasks; Gartner predicted 40% of agentic AI projects will be canceled by 2027; analysis of 847 deployments showed 76% fail in production. The window captured the field's fundamental tension: named orgs achieved measurable exploration gains within guardrails, while independent research documented persistent autonomy failures and governance gaps. Practitioner sentiment remained bifurcated: exploration-stage agentic coding proven viable for constrained tasks but unreliable for autonomous operation."
    },
    {
      "period": "2026-Mar",
      "text": "Population-level adoption confirmed at 55% of professional engineers regularly using agents, with Claude Code reaching #1 status in just 8 months; task-specific success rates documented: bug fixes 78%, test writing 82%, refactoring 45%, architecture 15%. Named deployments validated the constrained-use model: SINTEF completed autonomous multi-file refactoring of a 500K-LOC reservoir simulator with 99.9% accuracy; TELUS and Rakuten case studies (Anthropic trends report) reconfirmed 79% delivery-time reduction and 500K+ hours saved under human oversight; Brookings researchers used Claude Code to build full R packages and 20-page analyses in hours. A Princeton study found reliability improvements lagging capability growth at half the rate, with critical calibration gaps (52%); context-rot patterns were quantified with architectural countermeasures emerging (fresh subagents, atomic tasks, token budgeting). Deloitte's 2025 reality check confirmed 11% production deployment and projected 40% project failure by 2027. The defining tension held: exploration-mode agents deliver real productivity within guardrails, but autonomy gaps and reliability lags prevent broader deployment."
    },
    {
      "period": "2026-Apr (Early)",
      "text": "Market inflection confirmed: Claude Code achieved 46% 'most loved' ranking (vs. Copilot 9%, Cursor 19%), writing 4% of all GitHub commits (135K/day across 1M+ repos, 8% week-over-week growth). JetBrains survey shows 41% market share vs. Copilot's 38%, with 91% satisfaction and 54 NPS. Yet critical reliability degradation documented: AMD Sr Director's forensic analysis of 6,852 sessions revealed thinking depth collapsed 67-75% (late February–March), reads per edit dropped 70%, full-file rewrites doubled, monthly Bedrock costs exploded 122x ($345→$42K)—traced to Opus 4.6 adaptive thinking misconfiguration. Peer-reviewed research on 110K OSS PRs confirms agent activity scaling but documents higher code churn vs. human-authored code. Leading-edge capability (Ultraplan cloud planning, Monitor tool for reactive agents) shipped alongside hidden reliability regression, exemplifying the practice's defining tension: exploration-mode agentic coding delivers measurable productivity gains within guardrails, but tool maturity gaps (thinking depth, cost control, code quality) create operational hazards that constrain adoption to disciplined teams with oversight. New accessibility signal: domain experts (computational historians) successfully prototype without coding, expanding addressable population beyond software engineers."
    },
    {
      "period": "2026-Apr (Mid-Late)",
      "text": "Competitive capability inflection and reliability crisis converged. OpenAI GPT-5.5 released April 23 (82.7% Terminal-Bench SOTA, 88.7% SWE-bench, 20-hour autonomous engineering runs), crossing agentic coding from research into production default—yet same window revealed critical tool maturity gaps. Claude Opus 4.7 released April 16 with major benchmarks (SWE-Bench Pro 64.3%, CursorBench 70%), new xhigh reasoning effort tier, and task budgets for cost management. Yet April 23 postmortem documented three production bugs had degraded Claude Code quality March 4–April 20: reasoning-effort default flipped (34 days undetected), cache bug caused context loss and repetition, verbosity constraint cut coding-eval 3%—user costs spiked $345→$42K/month. Real-world empirical study (MSR 2026) on 11,771 PRs found Claude Opus 4.7 leads at 87.6% SWE-Bench but only 24% autonomous completion on complex tasks (70-90% failure as complexity rises), establishing critical gap between benchmark performance and real-world autonomy. Population-level adoption signal solidified: 275M weekly AI-authored commits on GitHub (agentic baseline past novelty into infrastructure), with bottleneck shifted from agent output velocity to human review and orchestration pipeline capacity. Anthropic 2026 industry report documented TELUS 500K+ hour savings, Augment Code 4-to-2-week project compression, yet showed developers use AI in 60% of work but fully delegate only 0-20% (collaborative model, not autonomous). The period crystallized the practice's maturity: capability crossed into production-ready (GPT-5.5 architectural reasoning, Opus 4.7 reasoning control), deployment reached infrastructure scale (275M commits/week), yet reliability gaps (tool bugs, benchmark-reality gap, limited autonomy on complex tasks) remain binding constraints on broader adoption beyond disciplined teams with strong oversight cultures."
    },
    {
      "period": "2026-May",
      "text": "Infrastructure expansion and enterprise normalization continued. Anthropic doubled Claude Code rate limits and removed peak-hour throttling (May 6), backed by a 300+ MW SpaceX compute deal; Claude Platform on AWS reached GA (May 11) with managed agents, code execution, and MCP integration. Amazon formally deployed Claude Code company-wide (May 4) and investment analysts valued it at $1B+ ARR alongside GitHub Copilot at 20M users. Anthropic's 2026 agentic coding report corroborated enterprise gains at named scale (Rakuten 79% delivery-time reduction, TELUS 500K hours saved, Zapier 89% adoption) while quantifying the cascade-failure risk: 85% per-step reliability yields only 20% end-to-end success over a 10-step workflow. Claude Code Agent View and /goal features shipped (Week 20), enabling developers to set completion conditions and run multi-turn autonomous exploration workflows unattended. Security vulnerabilities surfaced: CVE-2025-59536 and CVE-2026-21852 expose RCE and API key theft via untrusted repository configs in Claude Code. Adoption paradoxes deepened: HCL GUVI documented 84% developer adoption (41% of code AI-authored) but with median savings of 3.6 hrs/week and 19% of developers reporting slowdowns. A \"mise en place\" methodology—deliberate context engineering before agent runs—emerged as the key preparation pattern for reliable exploration workflows."
    },
    {
      "period": "2026-Jun",
      "text": "Adoption metrics crossed new thresholds: Stack Overflow survey shows 59% of developers using agentic tools (up from 31%), Vercel AI Gateway reports 30% of all deployments initiated by agentic tools (1,000% growth in six months, Claude Code at 75% share), and a peer-reviewed study of 932K agent-authored PRs finds adoption in new projects more than doubled versus the prior cohort. Anthropic's updated trends data shows multi-file editing surged from 34% to 78% of sessions with Rakuten processing 12.5M LOC in a 7-hour unattended run. GitHub's Scout agent (exploratory token-anomaly investigation, 61 network requests in 8 minutes) provides a named enterprise proof-point for exploration-task autonomy. The governance gap sharpened: Cursor telemetry shows unreviewed agent changes jumping from 7% to 36.3% of merges in five months, and a systematic study of 547 operational safety failures (59% high/critical severity) documents the real failure modes that constrained exploration still misses."
    },
    {
      "period": "2026-Jul",
      "text": "Peer-reviewed Codex analysis confirmed 5x user growth with 26.6% employing multi-step skills and sessions extending to 70%+ over one hour — the clearest large-scale signal of workflow sophistication. Anthropic's 400K-session study quantified the exploration division of labor (humans 70% of planning, agents 80% of execution), while practitioner evidence sharpened the exploration-vs-shipping boundary: one case documented an 18-hour RAG prototype followed by four days of debugging, confirming agentic coding as an exploration accelerator but not a delivery shortcut. A large-scale GitHub study (25,264 agentic PRs, 2,361 repos) confirmed adoption concentrates in small teams (1-5 contributors) under single-human oversight, while named case studies validated exploration-stage velocity — MEV shipped a PoC in 72 hours via a 7-agent pipeline, and estie's 18-month workflow evolution cut PR turnaround from 9.6 to 2.4 days. Veracode's longitudinal testing found syntax correctness climbing to ~95% while security pass rates stayed flat at 45-55% regardless of model generation, and OpenAI's GPT-5.6 Sol launch extended autonomous coding capability while carrying a 2.7x higher defect rate than human code — reinforcing the exploration-vs-production quality gap."
    },
    {
      "period": "2026-Aug",
      "text": "Adoption metrics kept climbing — Microsoft's CLI agent study found 24% more merged PRs among adopters and a JetBrains survey showed Claude Code preferred 46% to Copilot's 9% among senior developers (6x growth since April 2025) — while a wave of empirical studies sharpened the exploration-vs-production gap: Gartner found 88% of agent pilots never reach production, an AIDev PR study documented static-analysis warnings up 18% and cognitive complexity up 39%, and a 4,882-PR test-coverage analysis found 50.4% of agentic PRs ship with zero test changes. Analysts introduced \"agentic code debt\" as a named governance risk class for exploration-stage output that accumulates uncomprehended logic and hallucinated dependencies. Google shipped Antigravity with Gemini 3.7 Flash as a new agentic-IDE entrant, while SWE-Bench Pro data showed coding agents plateauing below 45% on realistic tasks and a subagent-delegation study found deep agentic search costs more and finds less than semantic search for repository exploration. A CI/CD security flaw let GitHub issues reach Claude Code and Gemini CLI workflow secrets, reinforcing governance-gap commentary around \"scope hallucination\" as agents exceed their intended exploration boundaries."
    },
    {
      "period": "2026-Sep",
      "text": "Large-sample evidence reinforced both sides of the exploration debate. A JetBrains survey of 15,000+ developers found 90% weekly agent usage with Claude Code the #1 tool for 31% (double GitHub Copilot's rate), and Stanford SALT Lab's analysis of 250,000 real Claude conversations found humans setting direction while adapting outputs in ~75% of cases, with over half involving consequential or hard-to-undo work. Google's Antigravity Teamwork framework reached GA with concrete exploration wins (7 solved math theorems, an autonomous RISC-V simulator), and Atlassian documented a 15-hour prototype-to-production build via multi-agent orchestration — though the same case showed agents outpacing team comprehension and breaking in production, shifting the bottleneck from code generation to human review. Countervailing evidence persisted: a METR randomized controlled trial found experienced developers 19% slower with AI tools despite self-reporting 20% faster, VESS Labs published a structured rejection of Devin after a one-week trial citing session-tracking and visibility gaps, and new research diagnosed 45 distinct agent failure patterns in autonomous research tasks, with a missing \"metacognitive loop\" as the dominant root cause. Peer-reviewed research formalizing production deployment boundaries emerged. Happy Bhati's synthesis (arxiv:2609.04681) identified the Agentic SDLC Throughput Paradox: agents accelerate code generation (+180% commits) but constrain production releases (+30%) due to verification, review, and security bottlenecks; introduced Production-Qualified Change (PQC) units and Verification Tax concept (recurring evaluation/retry/escalation costs). University of Pittsburgh + Ejento AI developed governance framework for agentic technical debt and stochastic tax management, proposing five control layers (golden-set evaluation, tool schema contracts, model gateways, graduated autonomy, workflow redesign). Large-scale empirical studies quantified code-quality gaps: Liu et al. (peer-reviewed, 304K AI-authored commits across 6,275 repos) found 484,606 introduced issues with 24.2% never fixed, 41.1% security issues persisting to latest revision; n1n.ai study (424 trials, multi-file tasks) showed forcing structured proof output increased evidence 30x but did NOT reduce false-success rates (75% with receipts vs 73.6% baseline). Security governance urgency escalated: Manifold disclosed GitSpawn (8 CVEs across Claude Code, Cursor, Codex, Gemini CLI enabling pre-trust execution), with multiple unpatched as of Sept 1. GitHub shipped Copilot agentic autofix GA (bulk remediation of 25 findings per assignment), signaling ecosystem maturation beyond exploration into routine maintenance. Microsoft Research analysis of 13.5M Copilot sessions (95 trillion tokens) revealed 87% of LLM calls are agent-initiated, establishing production-scale exploration workload patterns. A synthesis of Anthropic (400K sessions), OpenAI telemetry, and METR RCT data confirmed 70.2% of users now run multi-hour autonomous tasks even as the RCT found experienced developers 19% slower on complex work despite gains on simple ones, and a 9-engineer exe.dev retrospective documented the organizational prerequisites (architectural discipline, integration testing, design-first workflows) that let high-trust teams safely eliminate secondary review. The month crystallized the field's core tension: exploration-stage agents are now mainstream (90% weekly adoption, $2B+ ARR at scale), capability is proven in bounded scopes, yet production deployment requires investment in verification infrastructure, security hardening, and governance models that are still emerging. Fresh evidence added nuance: a Coleman Parkes poll found 93% of developers hit hallucinations and 35% report agent code reaching production before being understood, while New Relic found 88% of leaders permit vibe coding in production policies (78% seeing more incidents). OpenAI's internal Codex adoption rose from ~0% to 90% in non-engineering teams within four months, and an audit of 246 repositories found harness choice alone swinging one model's success rate from 68% to 88%."
    }
  ],
  "historyFallback": false,
  "lastUpdated": "2026-09-29",
  "domain": {
    "id": "software-development",
    "label": "Software Engineering",
    "icon": "⌨️"
  },
  "url": "https://www.thestateofplay.ai/practice/agentic-coding-for-exploration-and-prototyping",
  "license": "CC BY 4.0",
  "licenseUrl": "https://creativecommons.org/licenses/by/4.0/",
  "generatedAt": "2026-10-01"
}