The State of Play

A living index of AI adoption across industries — where established practice meets the bleeding edge
UPDATED DAILY
← All domains

⌨️ Software Engineering

AI across the development lifecycle — writing, reviewing, testing, and shipping code. Code completion is established and IDE-native; agentic coding and AI-driven CI/CD are advancing fast but half the domain remains bleeding-edge. The widest maturity spread of any domain: a few practices are table stakes while many are still experimental.

24 practices: 2 established, 4 good practice, 11 leading edge, 7 bleeding edge

The Headline

AI now writes code faster than companies can check it. The firms getting real returns limit what AI is allowed to ship. They are not the ones letting it run freely.

24
practices tracked
18
at the frontier
1
moving this fortnight
4,075
evidence items

The Picture

Most companies have given engineers AI coding tools, and many now let agents (software that acts on its own, without prompting) write and submit whole changes. According to McKinsey, investment in this kind of agent-driven development grew more than twelvefold from 2025 to 2026. Yet only a quarter of companies see meaningful acceleration, and productivity fell at 30% of them after adoption. A small group is pulling ahead by keeping AI on narrow jobs that can be checked: DoorDash's agents produced usable fixes for 45 of 50 sampled stale feature switches at about $4.79 each, and Adyen's automated upgrades are merged 70% of the time. The rest are learning that writing code was never the constraint. Checking it is, and companies that have not redesigned review are piling up risk faster than output.

This Fortnight

  • Two surveys published this month show AI code winning approval at review, then failing in production. Qodo found 89% of organizations have had an AI-related production incident, and only 3.7% of engineering leaders call their processes adequate. New Relic found 94% of US tech leaders rate AI code highly at review, yet 78% see more incidents once it ships. If your AI dashboards track output rather than incidents and rework, they are measuring the wrong thing.

  • Duolingo now lets a risk-scoring bot approve its lowest-risk 10% of code changes without a human reviewer. Median time to merge fell from 18 hours to 12 across about 300 engineers, and changes to cloud infrastructure and repositories under SOX or ISO audit are excluded. Separately, the startup Corridor showed that switching the underlying AI model alone moved wrong approvals from 42.9% to 5%. Automated approval can work, but only inside tight boundaries and on a model you have tested yourself.

  • Researchers at UpGuard found 16,326 Supabase databases with tables anyone could read, largely because AI tools created them without the default protections. AI-assisted development now creates more than 60% of new databases on the platform, and the safety defaults covered only tables made by hand in the dashboard. Any platform designed around human setup may leave gaps when agents do the setup, so it is worth auditing what your tools have built.

  • A zero-click flaw dubbed Plugin4Shell let attackers run code through all four major AI coding assistants, from OpenAI, Anthropic, Google and GitHub. Patching has been uneven across vendors. The tools your engineers run with broad access to code and credentials are now an attack surface in their own right, and they belong in your security reviews.

  • Microsoft relaunched its Copilot app on September 25 with a Code builder that lets non-technical staff create apps and automations by describing them. It is billed by usage rather than per seat, and the apps run on a governed hosting layer with an administrator kill switch. Fewer than 7% of Microsoft's 450 million-plus commercial Microsoft 365 seats carry Copilot licenses, so this is Microsoft's push to reach everyone else in your organization.

Coming Up

  • Microsoft's Code builder reaches early-access customers over the coming weeks, with gaps analysts at Moor Insights & Strategy say remain before general availability (out of beta, generally available). No one yet owns abandoned apps, there is no way to move apps elsewhere, and usage-based bills are hard to predict. Decide who owns citizen-built apps and set spending caps before staff start building, not after.

  • The EU has deferred the AI Act's human-oversight rules for high-risk AI to December 2027, and coding tools are not usually high-risk systems in any case. The case for a human at the merge gate rests on practice, not law: practitioner guidance already treats an AI approval as something that cannot stand in for a human sign-off on authentication, payments or data migrations. If your teams are piloting auto-approval, write down where a named person must still own the merge decision.

  • Cost is replacing code quality as the thing that ends AI coding programs. Earlier this year Uber used up its AI coding budget and Microsoft withdrew Anthropic's coding tool from most engineers, and Gartner forecasts that 40% of agentic AI projects will be cancelled by 2027. Research on a simulated 10,000-seat company found smarter model routing recovers only 14 to 21% of spend. Treat AI coding costs as a budget line with hard limits, not an open meter.

What's Hard About This

  • Every productivity gain in writing code turns into a queue at review. Linear's data across 47,900 workspaces shows teams tripling pull requests (proposed code changes) from 21 to 65 a week while total developer time rose. Buying faster generation without adding review capacity simply moves the bottleneck.

  • Human reviewers get less strict the more AI code they see. A study of 11,429 reviews found approval rates for AI code rising from 30.5% to 36.6% as reviewers grew used to it. Keeping a human-in-the-loop (a person reviews each AI output before it ships) is only a safeguard if the reviewer stays skeptical, and volume works against that.

  • AI succeeds on bounded, repeatable work and fails at scale. On whole-codebase migrations, only 5.4% of AI agent runs pass every check. Replacing large legacy systems end to end remains out of reach, and business cases built on that promise should be discounted accordingly.

2012View this domain's timeline →Today

Practices in this Domain (24)

PRACTICETIERTREND
Adversarial test generationLEADING EDGE— Steady
Agentic coding for exploration & prototypingLEADING EDGE— Steady
Agentic coding for production integrationBLEEDING EDGE— Steady
Agentic coding with full autonomyBLEEDING EDGE— Steady
AI-assisted code review with auto-approveBLEEDING EDGE— Steady
AI-assisted code review with suggestionsLEADING EDGE— Steady
AI-assisted test generationBLEEDING EDGE— Steady
API & schema generation from natural languageBLEEDING EDGE— Steady
Architecture documentation & specification writingGOOD PRACTICE— Steady
Chat-based code assistance & debuggingLEADING EDGE— Steady
CI/CD & infrastructure-as-code generationLEADING EDGE— Steady
Code & API documentation generationLEADING EDGE— Steady
Code refactoring & technical debt managementBLEEDING EDGE— Steady
Code search & codebase Q&ALEADING EDGE— Steady
Copilot-style inline code autocompleteESTABLISHED— Steady
Dependency management & cross-repository impact analysisLEADING EDGE— Steady
Deployment risk assessment & rollout managementESTABLISHED— Steady
Legacy code analysis & migrationLEADING EDGE— Steady
Multi-agent development pipelinesBLEEDING EDGE↑ Accelerating
Natural language to code for non-developersLEADING EDGE— Steady
Performance optimisation & query tuningGOOD PRACTICE— Steady
Security-focused code reviewLEADING EDGE— Steady
Test coverage analysis & gap identificationGOOD PRACTICE— Steady
Visual regression testing & self-healing test maintenanceGOOD PRACTICE— Steady
Read the full technical briefing (1,545 words) →

Where AI Stands in Software Engineering

Software engineering has the widest maturity spread of any domain AI touches, and the gap is getting wider. At one end, inline completion has become plumbing. GitHub's rebuilt Copilot suggestion model served more than 61 billion successful requests in 90 days, but interest has moved on. Copilot's share of Stack Overflow's survey fell from 67% to 51% as Cursor and Claude Code posted the fastest first-year debuts on record. Autocomplete-only tools keep 31% of users after 90 days, against 78% for agentic tools. At the other end, agents now write production systems. GitHub says agents wrote most of the 800,000-plus lines of Rust in its rewritten Copilot runtime, landed across 128 pull requests by mainly one developer in a few months. Linear's telemetry across 47,900 workspaces shows agents authoring about half of all issues and team pull-request throughput tripling from 21 to 65 a week. It also shows total developer time going up, not down.

That last detail is the arc of the domain. Producing code is no longer the constraint; checking it is. McKinsey data reported by CIO Dive shows investment in agentic software development growing more than twelvefold from 2025 to 2026. Yet only a quarter of companies see meaningful acceleration, and productivity fell in 30% after adoption. Qodo's survey of 800 developers and leaders found 89% of organisations had suffered an AI-related production incident, and only 3.7% of leaders think their processes are adequate. A study of 6,774 merged agent pull requests found they attract follow-up fixes at 1.62 times the odds of human ones, and 61.4% had no recorded human review. GitClear finds refactoring has shrunk from 21% of changed lines in 2022 to 3.8%. This explains why the practices on firmest ground are the ones that constrain rather than generate. Feature flags and progressive delivery, architecture specifications that agents execute against, and query tuning validated against EXPLAIN plans are mature because they put limits on what gets shipped.

Where momentum is real, the work is bounded and deterministic. Adyen's OpenRewrite rollout produced 4,000 automated merge requests with a 70% merge rate. DoorDash's four-agent pipeline produced usable pull requests for 45 of 50 sampled stale feature flags at about $4.79 each. Duolingo lets a risk bot approve only the lowest-risk tenth of its changes. Where the work is open-ended, it stalls. Omdia finds only 10% of application-development leaders using fully autonomous AI. IDC finds only 29% of production agents interact with one another, and just 7% of enterprises have advanced multi-agent orchestration. On whole-repository migrations, only 5.4% of agent runs pass every evaluation stage. One further feature sets this domain apart: it generates its own telemetry at enormous scale. The gap between how productive AI feels and how productive it measures is therefore better documented here than anywhere else.

What's New, 2026-09-15 to 2026-09-29

This fortnight brought no change in standing, only more evidence behind the verification problem. The McKinsey, Qodo and New Relic surveys arrived together. New Relic found 94% of US technology leaders rate AI code higher at review, yet 78% report more incidents once it ships. The merge decision is where the practical detail sharpened. Duolingo's pull-request risk bot now auto-approves 10% of changes across roughly 300 engineers, cutting median merge time from 18 hours to 12. Finance, audited and infrastructure repositories are explicitly excluded. Corridor, which made human review optional in July, showed that the choice of model alone can move false approvals from 42.9% (GPT-5.5) to 5.0% (GPT-5.6 Terra). On disputed changes, Grok-4.6 wrongly approved 81.7%. Harness found 76% of respondents believe they could disable a misbehaving agent within 15 minutes, but only 33% have a kill switch and 19% an automatic release gate. BCG reports that nearly two-thirds of AI front-runner firms have swapped upfront sign-off for staged rollouts. Even autocomplete produced an instructive result. GitHub's unified suggestion model showed no gain in acceptance, and the biggest improvement came from a client-side fix. A separate study found developers choosing between five AI suggestions picked the most and least secure at the same rate.

Security and the non-developer edge produced the sharpest news. UpGuard found 16,326 Supabase databases with publicly readable tables. AI-assisted development now creates more than 60% of new Supabase databases, and row-level security was on by default only for tables made in the dashboard, not those created by agents in SQL. Plugin4Shell, a zero-click remote-code-execution flaw, hit Codex, Claude Code, Gemini CLI and Copilot. Endor Labs found Opus 5.5 produced secure code on only 33.5% of tasks. On adversarial testing, NIST's CAISI scored Z.ai's GLM-5.3 at 40.4% on SEC-Bench Pro against 90.2% for the leading US model. Scale AI reported automated red-teaming broke a client's system in 3% of attempts, against 68% of sessions for human testers. Microsoft relaunched its Copilot app on 25 September with a Code builder for non-technical staff, billed by consumption and published to a governed runtime. Fewer than 7% of more than 450 million commercial Microsoft 365 seats carry Copilot licences. On legacy systems, Deloitte India and Cast formed an Asia-Pacific modernisation alliance on 23 September, while Ensono's survey showed most enterprises extending legacy estates in place rather than exiting them.

Key Tensions

  • Throughput outruns verification capacity. Linear's data shows weekly pull requests tripling while total developer time rises, because review and coordination absorb the gain. CircleCI's CTO reports main-branch success at a five-year low of 70.8%, and Qodo finds engineering leaders name reviewing AI code as their main delivery constraint. Every productivity gain upstream becomes a queue downstream.

  • Confidence at review, incidents in production. New Relic's respondents rate AI code highly at review and still see more incidents after it ships. A study of 11,429 reviews found approval rates for AI code rising from 30.5% to 36.6% as reviewers got used to it. Harness shows the same pattern in operations: most believe they can stop a bad agent quickly, and a third have the switch to do it.

  • Bounded determinism beats open-ended autonomy. Agents succeed when they call real engines or work to narrow scopes: JetBrains Rider's refactoring skill cut agent task time by 83%, and Adyen's recipe-driven merge requests land 70% of the time. Scale breaks them. RepoMod-Bench records accuracy falling from 91.3% to 15.3% as codebases grow, and only 5.4% of agent runs on whole-repository migrations pass every stage.

  • The toolchain has become the attack surface. Plugin4Shell and the earlier GitSpawn disclosures show coding agents themselves carrying zero-click and pre-trust execution flaws. Platform defaults built for humans leave tables exposed when agents create them, as the Supabase findings show. Agents that install packages directly also bypass Dependabot's new release cooldown, the ecosystem's main defence against freshly published malicious versions.

  • Token economics now decide deployments. In OpenChamber's analysis of 66,320 Reddit complaints, token cost rose from 9.1% to 13.7% of posts while buggy code fell from 13.1% to 9.6%. OpenChamber competes with several of the tools it counted. Earlier this year Uber exhausted its AI coding budget and Microsoft withdrew Claude Code from most engineers. Research on a simulated 10,000-seat enterprise finds cache-safe model routing recovers only 14–21% of spend, so cost control is a governance problem rather than a tuning one.

Top 10 Evidence Items

  1. Migrating the GitHub Copilot runtime to Rust, using Copilot (case-study) — The headline case for agents writing production-scale code, with review depth left unquantified — exactly the confidence gap the briefing flags. https://github.blog/ai-and-ml/generative-ai/migrating-the-github-copilot-runtime-to-rust-using-copilot/
  2. Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests (research-paper) — Hard evidence that agent PRs draw more follow-up fixes and skip review far more often, grounding the 'checking it is the constraint' claim. https://arxiv.org/html/2609.26847
  3. AI Code Generation Scaled. Verification Didn't. (industry-report) — Direct quantification of the verification gap: near-universal incidents against almost no leaders trusting their own processes. https://futurumgroup.com/insights/ai-code-generation-scaled-verification-didnt/
  4. Teaching Engineers, Trusting AI: How Education Enabled Autonomous Code Review (conference-talk) — Shows bounded, exclusion-scoped autonomy in merge decisions actually working at production scale, contrasting with open-ended failures elsewhere. https://www.infoq.com/presentations/duolingo-ai-literacy-code-review/
  5. How Changing Models Cut False PR Approvals by 80% (case-study) — Demonstrates that model choice alone swings false-approval rates by an order of magnitude, sharpening the merge-decision tension. https://www.corridor.dev/blog/how-changing-models-cut-false-pr-approvals
  6. 16,326 Supabase Databases Had Public Data: How to Fix Missing RLS (news-coverage) — Shows the toolchain's attack surface concretely: agent defaults bypassing security protections humans would have set. https://windowsforum.com/news/16-326-supabase-databases-had-public-data-how-to-fix-missing-rls.446355/
  7. A zero-click RCE flaw in AI coding agents could have exposed enterprise systems (news-coverage) — A zero-click RCE across nearly every major coding agent turns the abstract 'toolchain as attack surface' point into a dated, concrete event. https://www.infoworld.com/article/4223907/a-zero-click-rce-flaw-in-ai-coding-agents-could-have-exposed-enterprise-systems.html
  8. DoorDash Uses Multi-Agent LLM System to Remove 60,000 Stale Feature Flags (news-coverage) — The clearest working example of bounded, deterministic agent scope succeeding where open-ended autonomy stalls. https://www.webpronews.com/doordash-uses-multi-agent-llm-system-to-remove-60000-stale-feature-flags/
  9. Harness report finds enterprise AI agent confidence outpaces actual governance controls (news-coverage) — Captures the confidence-versus-control gap directly: most believe they could stop a bad agent, few actually can. https://digitalisationworld.com/news/23784-harness-report-finds-enterprise-ai-agent-confidence-outpaces-actual-governance-controls
  10. Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise (research-paper) — Shows that even with routing, most of the cost problem persists, reframing token economics as a governance rather than tuning problem. https://arxiv.org/html/2609.28919