{
  "id": "software-development",
  "label": "Software Engineering",
  "description": "AI across the development lifecycle — writing, reviewing, testing, and shipping code. Code completion is established and IDE-native; agentic coding and AI-driven CI/CD are advancing fast but half the domain remains bleeding-edge. The widest maturity spread of any domain: a few practices are table stakes while many are still experimental.",
  "icon": "⌨️",
  "filters": [
    "building"
  ],
  "hasSummary": true,
  "hasExecSummary": true,
  "practiceCount": 24,
  "evidenceCount": 4075,
  "practices": [
    {
      "slug": "adversarial-test-generation",
      "name": "Adversarial test generation",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI using reinforcement learning or adversarial techniques to generate edge-case and fault-finding test scenarios. Includes fuzz testing augmented with LLMs and RL-based test case evolution; distinct from standard test generation which aims for coverage rather than fault discovery.",
      "evidenceCount": 141
    },
    {
      "slug": "agentic-coding-for-exploration-and-prototyping",
      "name": "Agentic coding for exploration & prototyping",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI agents that autonomously write, run, and iterate on code for proofs of concept and exploratory development. Includes tools like Claude Code, Cursor agent mode, and Devin used for throwaway prototypes and spikes; distinct from production agentic coding which requires CI/CD integration and review workflows.",
      "evidenceCount": 176
    },
    {
      "slug": "agentic-coding-for-production-integration",
      "name": "Agentic coding for production integration",
      "tier": "bleeding-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI agents completing production coding tasks within supervised workflows, including PR creation and review cycles. Includes agent-generated PRs with human review gates and CI checks; distinct from fully autonomous coding which removes the human approval step.",
      "evidenceCount": 156
    },
    {
      "slug": "agentic-coding-with-full-autonomy",
      "name": "Agentic coding with full autonomy",
      "tier": "bleeding-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI agents independently completing development tasks end-to-end with minimal human oversight or intervention. Includes autonomous issue-to-merge workflows and self-directed multi-file changes; distinct from supervised production integration which retains human review gates.",
      "evidenceCount": 144
    },
    {
      "slug": "ai-assisted-code-review-with-auto-approve",
      "name": "AI-assisted code review with auto-approve",
      "tier": "bleeding-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that autonomously approves and merges code changes meeting defined quality and safety thresholds. Includes automated merge for low-risk changes like dependency bumps; distinct from suggestion-mode review which always requires human sign-off.",
      "evidenceCount": 161
    },
    {
      "slug": "ai-assisted-code-review-with-suggestions",
      "name": "AI-assisted code review with suggestions",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that reviews pull requests and annotates code with improvement suggestions for human reviewers to accept or reject. Includes PR review bots and automated code quality comments; distinct from auto-approve which removes the human decision step.",
      "evidenceCount": 170
    },
    {
      "slug": "ai-assisted-test-generation",
      "name": "AI-assisted test generation",
      "tier": "bleeding-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that generates unit, integration, or end-to-end tests from source code, requirements documents, or API specifications. Includes tools generating test suites from implementations, PRDs, and OpenAPI specs; distinct from adversarial test generation which targets fault discovery rather than coverage.",
      "evidenceCount": 165
    },
    {
      "slug": "api-and-schema-generation-from-natural-language",
      "name": "API & schema generation from natural language",
      "tier": "bleeding-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI generating API endpoints, database schemas, or data models from natural language descriptions of requirements. Includes REST/GraphQL API scaffolding and database schema design; distinct from infrastructure-as-code which targets deployment resources rather than application interfaces.",
      "evidenceCount": 172
    },
    {
      "slug": "architecture-documentation-and-specification-writing",
      "name": "Architecture documentation & specification writing",
      "tier": "good-practice",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that generates architecture diagrams, system design documents, and technical specifications from codebases and requirements. Includes C4 diagram generation and design doc drafting; distinct from code documentation which targets inline and API-level references.",
      "evidenceCount": 157
    },
    {
      "slug": "chat-based-code-assistance-and-debugging",
      "name": "Chat-based code assistance & debugging",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "Conversational AI that answers coding questions, explains errors, and helps debug issues in a chat interface. Includes IDE chat panels, web-based coding assistants, and error explanation tools; distinct from inline autocomplete which operates without explicit prompting.",
      "evidenceCount": 198
    },
    {
      "slug": "cicd-and-infrastructure-as-code-generation",
      "name": "CI/CD & infrastructure-as-code generation",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that generates, configures, or optimises CI/CD pipelines and infrastructure-as-code definitions for faster, safer deployments. Includes pipeline YAML generation, build optimisation, and cloud resource templating; distinct from deployment risk assessment which evaluates changes rather than generating configurations.",
      "evidenceCount": 173
    },
    {
      "slug": "code-and-api-documentation-generation",
      "name": "Code & API documentation generation",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that generates inline code documentation, API references, commit messages, and changelogs from source code and change history. Includes docstring generation and OpenAPI doc creation; distinct from architecture documentation which produces system-level design documents.",
      "evidenceCount": 177
    },
    {
      "slug": "code-refactoring-and-technical-debt-management",
      "name": "Code refactoring & technical debt management",
      "tier": "bleeding-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that identifies refactoring opportunities, surfaces technical debt, and suggests prioritised improvements for code quality and maintainability. Includes dead code detection, complexity reduction, and maintenance cost estimation; distinct from code review which evaluates new changes rather than existing code.",
      "evidenceCount": 172
    },
    {
      "slug": "code-search-and-codebase-qanda",
      "name": "Code search & codebase Q&A",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI-powered semantic search and question answering across large codebases, going beyond keyword matching. Includes tools that answer questions about architecture, dependencies, and usage patterns; distinct from documentation generation which produces static artefacts.",
      "evidenceCount": 166
    },
    {
      "slug": "copilot-style-inline-code-autocomplete",
      "name": "Copilot-style inline code autocomplete",
      "tier": "established",
      "trend": "steady",
      "blockerType": null,
      "description": "AI-powered real-time code suggestions appearing inline as developers type, predicting next tokens or lines. Includes IDE-integrated completion tools like GitHub Copilot and Tabnine; distinct from chat-based assistance which involves conversational interaction rather than inline prediction.",
      "evidenceCount": 202
    },
    {
      "slug": "dependency-management-and-cross-repository-impact-analysis",
      "name": "Dependency management & cross-repository impact analysis",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that manages dependencies, remediates vulnerabilities, and analyses the impact of changes across multiple repositories. Includes automated dependency updates and cross-repo change impact prediction; distinct from security code review which examines code logic rather than dependency graphs.",
      "evidenceCount": 170
    },
    {
      "slug": "deployment-risk-assessment-and-rollout-management",
      "name": "Deployment risk assessment & rollout management",
      "tier": "established",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that evaluates deployment risk, recommends rollback strategies, and manages feature flag rollouts to reduce release incidents. Includes change impact prediction and progressive delivery analysis; distinct from CI/CD generation which creates pipeline configurations.",
      "evidenceCount": 168
    },
    {
      "slug": "legacy-code-analysis-and-migration",
      "name": "Legacy code analysis & migration",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that analyses legacy systems to document behaviour, identify dependencies, and assist migration to modern platforms. Includes COBOL-to-Java migration and mainframe modernisation; distinct from code refactoring which improves existing code within its current platform.",
      "evidenceCount": 164
    },
    {
      "slug": "multi-agent-development-pipelines",
      "name": "Multi-agent development pipelines",
      "tier": "bleeding-edge",
      "trend": "accelerating",
      "blockerType": null,
      "description": "Multiple AI agents collaborating across development tasks such as planning, coding, reviewing, and testing in coordinated workflows. Includes orchestrated agent teams with specialised roles; distinct from single-agent agentic coding which uses one agent across the lifecycle.",
      "evidenceCount": 157
    },
    {
      "slug": "natural-language-to-code-for-non-developers",
      "name": "Natural language to code for non-developers",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "Tools enabling non-technical users to generate functional code or automations from plain language descriptions. Includes no-code/low-code AI builders and spreadsheet-to-app tools; distinct from chat-based code assistance which targets developers.",
      "evidenceCount": 174
    },
    {
      "slug": "performance-optimisation-and-query-tuning",
      "name": "Performance optimisation & query tuning",
      "tier": "good-practice",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that identifies application performance bottlenecks, optimises database queries, and recommends resource improvements. Includes query plan analysis and runtime profiling; distinct from code refactoring which targets code quality rather than runtime performance.",
      "evidenceCount": 200
    },
    {
      "slug": "security-focused-code-review",
      "name": "Security-focused code review",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI augmenting static and dynamic application security testing to identify vulnerabilities in code before deployment. Includes LLM-augmented SAST/DAST tools and AI-powered vulnerability explanation; distinct from general code review which focuses on quality rather than security.",
      "evidenceCount": 176
    },
    {
      "slug": "test-coverage-analysis-and-gap-identification",
      "name": "Test coverage analysis & gap identification",
      "tier": "good-practice",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that analyses test suites to identify untested paths, missing edge cases, and coverage blind spots. Includes intelligent coverage gap analysis beyond line-count metrics; distinct from test generation which creates tests rather than analysing existing ones.",
      "evidenceCount": 154
    },
    {
      "slug": "visual-regression-testing-and-self-healing-test-maintenance",
      "name": "Visual regression testing & self-healing test maintenance",
      "tier": "good-practice",
      "trend": "steady",
      "blockerType": null,
      "description": "AI-powered visual comparison of UI across builds and automatic repair of test scripts when application changes break existing tests. Includes intelligent screenshot diffing and selector auto-repair; distinct from test generation which creates new tests rather than maintaining existing ones.",
      "evidenceCount": 182
    }
  ],
  "summary": "## Where AI Stands in Software Engineering\n\nSoftware engineering has the widest maturity spread of any domain AI touches, and the gap is getting wider. At one end, inline completion has become plumbing. GitHub's rebuilt Copilot suggestion model served more than 61 billion successful requests in 90 days, but interest has moved on. Copilot's share of Stack Overflow's survey fell from 67% to 51% as Cursor and Claude Code posted the fastest first-year debuts on record. Autocomplete-only tools keep 31% of users after 90 days, against 78% for agentic tools. At the other end, agents now write production systems. GitHub says agents wrote most of the 800,000-plus lines of Rust in its rewritten Copilot runtime, landed across 128 pull requests by mainly one developer in a few months. Linear's telemetry across 47,900 workspaces shows agents authoring about half of all issues and team pull-request throughput tripling from 21 to 65 a week. It also shows total developer time going up, not down.\n\nThat last detail is the arc of the domain. Producing code is no longer the constraint; checking it is. McKinsey data reported by CIO Dive shows investment in agentic software development growing more than twelvefold from 2025 to 2026. Yet only a quarter of companies see meaningful acceleration, and productivity fell in 30% after adoption. Qodo's survey of 800 developers and leaders found 89% of organisations had suffered an AI-related production incident, and only 3.7% of leaders think their processes are adequate. A study of 6,774 merged agent pull requests found they attract follow-up fixes at 1.62 times the odds of human ones, and 61.4% had no recorded human review. GitClear finds refactoring has shrunk from 21% of changed lines in 2022 to 3.8%. This explains why the practices on firmest ground are the ones that constrain rather than generate. Feature flags and progressive delivery, architecture specifications that agents execute against, and query tuning validated against EXPLAIN plans are mature because they put limits on what gets shipped.\n\nWhere momentum is real, the work is bounded and deterministic. Adyen's OpenRewrite rollout produced 4,000 automated merge requests with a 70% merge rate. DoorDash's four-agent pipeline produced usable pull requests for 45 of 50 sampled stale feature flags at about $4.79 each. Duolingo lets a risk bot approve only the lowest-risk tenth of its changes. Where the work is open-ended, it stalls. Omdia finds only 10% of application-development leaders using fully autonomous AI. IDC finds only 29% of production agents interact with one another, and just 7% of enterprises have advanced multi-agent orchestration. On whole-repository migrations, only 5.4% of agent runs pass every evaluation stage. One further feature sets this domain apart: it generates its own telemetry at enormous scale. The gap between how productive AI feels and how productive it measures is therefore better documented here than anywhere else.\n\n## What's New, 2026-09-15 to 2026-09-29\n\nThis fortnight brought no change in standing, only more evidence behind the verification problem. The McKinsey, Qodo and New Relic surveys arrived together. New Relic found 94% of US technology leaders rate AI code higher at review, yet 78% report more incidents once it ships. The merge decision is where the practical detail sharpened. Duolingo's pull-request risk bot now auto-approves 10% of changes across roughly 300 engineers, cutting median merge time from 18 hours to 12. Finance, audited and infrastructure repositories are explicitly excluded. Corridor, which made human review optional in July, showed that the choice of model alone can move false approvals from 42.9% (GPT-5.5) to 5.0% (GPT-5.6 Terra). On disputed changes, Grok-4.6 wrongly approved 81.7%. Harness found 76% of respondents believe they could disable a misbehaving agent within 15 minutes, but only 33% have a kill switch and 19% an automatic release gate. BCG reports that nearly two-thirds of AI front-runner firms have swapped upfront sign-off for staged rollouts. Even autocomplete produced an instructive result. GitHub's unified suggestion model showed no gain in acceptance, and the biggest improvement came from a client-side fix. A separate study found developers choosing between five AI suggestions picked the most and least secure at the same rate.\n\nSecurity and the non-developer edge produced the sharpest news. UpGuard found 16,326 Supabase databases with publicly readable tables. AI-assisted development now creates more than 60% of new Supabase databases, and row-level security was on by default only for tables made in the dashboard, not those created by agents in SQL. Plugin4Shell, a zero-click remote-code-execution flaw, hit Codex, Claude Code, Gemini CLI and Copilot. Endor Labs found Opus 5.5 produced secure code on only 33.5% of tasks. On adversarial testing, NIST's CAISI scored Z.ai's GLM-5.3 at 40.4% on SEC-Bench Pro against 90.2% for the leading US model. Scale AI reported automated red-teaming broke a client's system in 3% of attempts, against 68% of sessions for human testers. Microsoft relaunched its Copilot app on 25 September with a Code builder for non-technical staff, billed by consumption and published to a governed runtime. Fewer than 7% of more than 450 million commercial Microsoft 365 seats carry Copilot licences. On legacy systems, Deloitte India and Cast formed an Asia-Pacific modernisation alliance on 23 September, while Ensono's survey showed most enterprises extending legacy estates in place rather than exiting them.\n\n## Key Tensions\n\n- **Throughput outruns verification capacity.** Linear's data shows weekly pull requests tripling while total developer time rises, because review and coordination absorb the gain. CircleCI's CTO reports main-branch success at a five-year low of 70.8%, and Qodo finds engineering leaders name reviewing AI code as their main delivery constraint. Every productivity gain upstream becomes a queue downstream.\n\n- **Confidence at review, incidents in production.** New Relic's respondents rate AI code highly at review and still see more incidents after it ships. A study of 11,429 reviews found approval rates for AI code rising from 30.5% to 36.6% as reviewers got used to it. Harness shows the same pattern in operations: most believe they can stop a bad agent quickly, and a third have the switch to do it.\n\n- **Bounded determinism beats open-ended autonomy.** Agents succeed when they call real engines or work to narrow scopes: JetBrains Rider's refactoring skill cut agent task time by 83%, and Adyen's recipe-driven merge requests land 70% of the time. Scale breaks them. RepoMod-Bench records accuracy falling from 91.3% to 15.3% as codebases grow, and only 5.4% of agent runs on whole-repository migrations pass every stage.\n\n- **The toolchain has become the attack surface.** Plugin4Shell and the earlier GitSpawn disclosures show coding agents themselves carrying zero-click and pre-trust execution flaws. Platform defaults built for humans leave tables exposed when agents create them, as the Supabase findings show. Agents that install packages directly also bypass Dependabot's new release cooldown, the ecosystem's main defence against freshly published malicious versions.\n\n- **Token economics now decide deployments.** In OpenChamber's analysis of 66,320 Reddit complaints, token cost rose from 9.1% to 13.7% of posts while buggy code fell from 13.1% to 9.6%. OpenChamber competes with several of the tools it counted. Earlier this year Uber exhausted its AI coding budget and Microsoft withdrew Claude Code from most engineers. Research on a simulated 10,000-seat enterprise finds cache-safe model routing recovers only 14–21% of spend, so cost control is a governance problem rather than a tuning one.\n\n## Top 10 Evidence Items\n\n1. **Migrating the GitHub Copilot runtime to Rust, using Copilot** (case-study) — The headline case for agents writing production-scale code, with review depth left unquantified — exactly the confidence gap the briefing flags. https://github.blog/ai-and-ml/generative-ai/migrating-the-github-copilot-runtime-to-rust-using-copilot/\n2. **Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests** (research-paper) — Hard evidence that agent PRs draw more follow-up fixes and skip review far more often, grounding the 'checking it is the constraint' claim. https://arxiv.org/html/2609.26847\n3. **AI Code Generation Scaled. Verification Didn't.** (industry-report) — Direct quantification of the verification gap: near-universal incidents against almost no leaders trusting their own processes. https://futurumgroup.com/insights/ai-code-generation-scaled-verification-didnt/\n4. **Teaching Engineers, Trusting AI: How Education Enabled Autonomous Code Review** (conference-talk) — Shows bounded, exclusion-scoped autonomy in merge decisions actually working at production scale, contrasting with open-ended failures elsewhere. https://www.infoq.com/presentations/duolingo-ai-literacy-code-review/\n5. **How Changing Models Cut False PR Approvals by 80%** (case-study) — Demonstrates that model choice alone swings false-approval rates by an order of magnitude, sharpening the merge-decision tension. https://www.corridor.dev/blog/how-changing-models-cut-false-pr-approvals\n6. **16,326 Supabase Databases Had Public Data: How to Fix Missing RLS** (news-coverage) — Shows the toolchain's attack surface concretely: agent defaults bypassing security protections humans would have set. https://windowsforum.com/news/16-326-supabase-databases-had-public-data-how-to-fix-missing-rls.446355/\n7. **A zero-click RCE flaw in AI coding agents could have exposed enterprise systems** (news-coverage) — A zero-click RCE across nearly every major coding agent turns the abstract 'toolchain as attack surface' point into a dated, concrete event. https://www.infoworld.com/article/4223907/a-zero-click-rce-flaw-in-ai-coding-agents-could-have-exposed-enterprise-systems.html\n8. **DoorDash Uses Multi-Agent LLM System to Remove 60,000 Stale Feature Flags** (news-coverage) — The clearest working example of bounded, deterministic agent scope succeeding where open-ended autonomy stalls. https://www.webpronews.com/doordash-uses-multi-agent-llm-system-to-remove-60000-stale-feature-flags/\n9. **Harness report finds enterprise AI agent confidence outpaces actual governance controls** (news-coverage) — Captures the confidence-versus-control gap directly: most believe they could stop a bad agent, few actually can. https://digitalisationworld.com/news/23784-harness-report-finds-enterprise-ai-agent-confidence-outpaces-actual-governance-controls\n10. **Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise** (research-paper) — Shows that even with routing, most of the cost problem persists, reframing token economics as a governance rather than tuning problem. https://arxiv.org/html/2609.28919",
  "execSummary": "**The headline:** AI now writes code faster than companies can check it. The firms getting real returns limit what AI is allowed to ship. They are not the ones letting it run freely.\n\n### The Picture\n\nMost companies have given engineers AI coding tools, and many now let agents (software that acts on its own, without prompting) write and submit whole changes. According to McKinsey, investment in this kind of agent-driven development grew more than twelvefold from 2025 to 2026. Yet only a quarter of companies see meaningful acceleration, and productivity fell at 30% of them after adoption. A small group is pulling ahead by keeping AI on narrow jobs that can be checked: DoorDash's agents produced usable fixes for 45 of 50 sampled stale feature switches at about $4.79 each, and Adyen's automated upgrades are merged 70% of the time. The rest are learning that writing code was never the constraint. Checking it is, and companies that have not redesigned review are piling up risk faster than output.\n\n### This Fortnight\n\n- **Two surveys published this month show AI code winning approval at review, then failing in production.** Qodo found 89% of organizations have had an AI-related production incident, and only 3.7% of engineering leaders call their processes adequate. New Relic found 94% of US tech leaders rate AI code highly at review, yet 78% see more incidents once it ships. If your AI dashboards track output rather than incidents and rework, they are measuring the wrong thing.\n\n- **Duolingo now lets a risk-scoring bot approve its lowest-risk 10% of code changes without a human reviewer.** Median time to merge fell from 18 hours to 12 across about 300 engineers, and changes to cloud infrastructure and repositories under SOX or ISO audit are excluded. Separately, the startup Corridor showed that switching the underlying AI model alone moved wrong approvals from 42.9% to 5%. Automated approval can work, but only inside tight boundaries and on a model you have tested yourself.\n\n- **Researchers at UpGuard found 16,326 Supabase databases with tables anyone could read, largely because AI tools created them without the default protections.** AI-assisted development now creates more than 60% of new databases on the platform, and the safety defaults covered only tables made by hand in the dashboard. Any platform designed around human setup may leave gaps when agents do the setup, so it is worth auditing what your tools have built.\n\n- **A zero-click flaw dubbed Plugin4Shell let attackers run code through all four major AI coding assistants, from OpenAI, Anthropic, Google and GitHub.** Patching has been uneven across vendors. The tools your engineers run with broad access to code and credentials are now an attack surface in their own right, and they belong in your security reviews.\n\n- **Microsoft relaunched its Copilot app on September 25 with a Code builder that lets non-technical staff create apps and automations by describing them.** It is billed by usage rather than per seat, and the apps run on a governed hosting layer with an administrator kill switch. Fewer than 7% of Microsoft's 450 million-plus commercial Microsoft 365 seats carry Copilot licenses, so this is Microsoft's push to reach everyone else in your organization.\n\n### Coming Up\n\n- **Microsoft's Code builder reaches early-access customers over the coming weeks, with gaps analysts at Moor Insights & Strategy say remain before general availability (out of beta, generally available).** No one yet owns abandoned apps, there is no way to move apps elsewhere, and usage-based bills are hard to predict. Decide who owns citizen-built apps and set spending caps before staff start building, not after.\n\n- **The EU has deferred the AI Act's human-oversight rules for high-risk AI to December 2027, and coding tools are not usually high-risk systems in any case.** The case for a human at the merge gate rests on practice, not law: practitioner guidance already treats an AI approval as something that cannot stand in for a human sign-off on authentication, payments or data migrations. If your teams are piloting auto-approval, write down where a named person must still own the merge decision.\n\n- **Cost is replacing code quality as the thing that ends AI coding programs.** Earlier this year Uber used up its AI coding budget and Microsoft withdrew Anthropic's coding tool from most engineers, and Gartner forecasts that 40% of agentic AI projects will be cancelled by 2027. Research on a simulated 10,000-seat company found smarter model routing recovers only 14 to 21% of spend. Treat AI coding costs as a budget line with hard limits, not an open meter.\n\n### What's Hard About This\n\n- **Every productivity gain in writing code turns into a queue at review.** Linear's data across 47,900 workspaces shows teams tripling pull requests (proposed code changes) from 21 to 65 a week while total developer time rose. Buying faster generation without adding review capacity simply moves the bottleneck.\n\n- **Human reviewers get less strict the more AI code they see.** A study of 11,429 reviews found approval rates for AI code rising from 30.5% to 36.6% as reviewers grew used to it. Keeping a human-in-the-loop (a person reviews each AI output before it ships) is only a safeguard if the reviewer stays skeptical, and volume works against that.\n\n- **AI succeeds on bounded, repeatable work and fails at scale.** On whole-codebase migrations, only 5.4% of AI agent runs pass every check. Replacing large legacy systems end to end remains out of reach, and business cases built on that promise should be discounted accordingly.",
  "headline": "AI now writes code faster than companies can check it. The firms getting real returns limit what AI is allowed to ship. They are not the ones letting it run freely.",
  "execSummarySections": [
    {
      "id": "the-picture",
      "title": "The Picture",
      "body": "Most companies have given engineers AI coding tools, and many now let agents (software that acts on its own, without prompting) write and submit whole changes. According to McKinsey, investment in this kind of agent-driven development grew more than twelvefold from 2025 to 2026. Yet only a quarter of companies see meaningful acceleration, and productivity fell at 30% of them after adoption. A small group is pulling ahead by keeping AI on narrow jobs that can be checked: DoorDash's agents produced usable fixes for 45 of 50 sampled stale feature switches at about $4.79 each, and Adyen's automated upgrades are merged 70% of the time. The rest are learning that writing code was never the constraint. Checking it is, and companies that have not redesigned review are piling up risk faster than output."
    },
    {
      "id": "this-fortnight",
      "title": "This Fortnight",
      "body": "- **Two surveys published this month show AI code winning approval at review, then failing in production.** Qodo found 89% of organizations have had an AI-related production incident, and only 3.7% of engineering leaders call their processes adequate. New Relic found 94% of US tech leaders rate AI code highly at review, yet 78% see more incidents once it ships. If your AI dashboards track output rather than incidents and rework, they are measuring the wrong thing.\n\n- **Duolingo now lets a risk-scoring bot approve its lowest-risk 10% of code changes without a human reviewer.** Median time to merge fell from 18 hours to 12 across about 300 engineers, and changes to cloud infrastructure and repositories under SOX or ISO audit are excluded. Separately, the startup Corridor showed that switching the underlying AI model alone moved wrong approvals from 42.9% to 5%. Automated approval can work, but only inside tight boundaries and on a model you have tested yourself.\n\n- **Researchers at UpGuard found 16,326 Supabase databases with tables anyone could read, largely because AI tools created them without the default protections.** AI-assisted development now creates more than 60% of new databases on the platform, and the safety defaults covered only tables made by hand in the dashboard. Any platform designed around human setup may leave gaps when agents do the setup, so it is worth auditing what your tools have built.\n\n- **A zero-click flaw dubbed Plugin4Shell let attackers run code through all four major AI coding assistants, from OpenAI, Anthropic, Google and GitHub.** Patching has been uneven across vendors. The tools your engineers run with broad access to code and credentials are now an attack surface in their own right, and they belong in your security reviews.\n\n- **Microsoft relaunched its Copilot app on September 25 with a Code builder that lets non-technical staff create apps and automations by describing them.** It is billed by usage rather than per seat, and the apps run on a governed hosting layer with an administrator kill switch. Fewer than 7% of Microsoft's 450 million-plus commercial Microsoft 365 seats carry Copilot licenses, so this is Microsoft's push to reach everyone else in your organization."
    },
    {
      "id": "coming-up",
      "title": "Coming Up",
      "body": "- **Microsoft's Code builder reaches early-access customers over the coming weeks, with gaps analysts at Moor Insights & Strategy say remain before general availability (out of beta, generally available).** No one yet owns abandoned apps, there is no way to move apps elsewhere, and usage-based bills are hard to predict. Decide who owns citizen-built apps and set spending caps before staff start building, not after.\n\n- **The EU has deferred the AI Act's human-oversight rules for high-risk AI to December 2027, and coding tools are not usually high-risk systems in any case.** The case for a human at the merge gate rests on practice, not law: practitioner guidance already treats an AI approval as something that cannot stand in for a human sign-off on authentication, payments or data migrations. If your teams are piloting auto-approval, write down where a named person must still own the merge decision.\n\n- **Cost is replacing code quality as the thing that ends AI coding programs.** Earlier this year Uber used up its AI coding budget and Microsoft withdrew Anthropic's coding tool from most engineers, and Gartner forecasts that 40% of agentic AI projects will be cancelled by 2027. Research on a simulated 10,000-seat company found smarter model routing recovers only 14 to 21% of spend. Treat AI coding costs as a budget line with hard limits, not an open meter."
    },
    {
      "id": "whats-hard-about-this",
      "title": "What's Hard About This",
      "body": "- **Every productivity gain in writing code turns into a queue at review.** Linear's data across 47,900 workspaces shows teams tripling pull requests (proposed code changes) from 21 to 65 a week while total developer time rose. Buying faster generation without adding review capacity simply moves the bottleneck.\n\n- **Human reviewers get less strict the more AI code they see.** A study of 11,429 reviews found approval rates for AI code rising from 30.5% to 36.6% as reviewers grew used to it. Keeping a human-in-the-loop (a person reviews each AI output before it ships) is only a safeguard if the reviewer stays skeptical, and volume works against that.\n\n- **AI succeeds on bounded, repeatable work and fails at scale.** On whole-codebase migrations, only 5.4% of AI agent runs pass every check. Replacing large legacy systems end to end remains out of reach, and business cases built on that promise should be discounted accordingly."
    }
  ],
  "url": "https://www.thestateofplay.ai/domain/software-development",
  "license": "CC BY 4.0",
  "licenseUrl": "https://creativecommons.org/licenses/by/4.0/",
  "generatedAt": "2026-10-01"
}