The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
AI for keeping digital systems running, observable, and secure. One of the most mature domains: log analysis, threat detection, and automated remediation are established or good practice. AIOps and SIEM are mainstream. Bleeding-edge frontiers include autonomous incident response and AI-driven penetration testing. Five practices are actively advancing; the rest are holding steady at good-practice level.
This is still the most AI-saturated domain in the enterprise, and for the first time in a year the maturity map did not move. Not one of the twenty practices tracked here was reclassified this cycle. Threat detection, log analysis, performance monitoring, identity anomaly detection, drift remediation and vulnerability management all held their position, and the trend readings came through the fortnight unchanged. After eighteen months of near-continuous reshuffling that stillness is itself the headline. The capability argument is over. What arrived instead was the most unflattering deployment data the domain has produced.
The industry has stopped counting deployments and started counting rollbacks. Independent analysis this cycle found that 74 percent of enterprises rolled back customer-facing AI agents after putting them into production — and that organisations with mature governance rolled back at 81 percent, a higher rate than everyone else. That inversion is the most important number in this batch, because it dismantles the comfortable assumption that governance prevents failure. It does not. It makes failure visible, which is a different and less marketable service. The supporting evidence is consistent: Gartner's 2026 research separates a 75 percent adoption claim from 17 percent actually deployed and 11 percent production-ready, a 42-point gap attributed principally to non-determinism (70 percent) and governance and observability shortfalls (64 percent); Fivetran's survey of 400 data leaders found 60 percent investing millions against 15 percent production-ready, with data quality and provenance the leading blocker; Prophet Security's poll of 250 security leaders found 46 percent of internally built AI tooling projects subsequently deprecated. In August, Gartner moved AI SOC Agents to the Peak of Inflated Expectations at 1 to 5 percent market penetration and issued an explicit warning about AI-washing — vendors labelling ordinary automation as agentic without demonstrating planning, action or recovery.
And yet the production receipts landing in the same fortnight are the strongest this domain has ever produced. SentinelOne reports Purple AI running more than 8,500 autonomous investigations a day across upwards of a third of its customer base, handling close to three times the alert volume of its human analysts, at an 18-minute average response time. Stellar Cyber's 124-day production trial across 138,475 alerts closed 8,047 autonomously — 64 percent of all verdicts — escalated 1,875 as genuine threats, recovered 19 minutes per analyst hour, and recorded 99.7 percent agreement between human and machine verdicts. Morgan Stanley collapsed mean time to detect from 45 minutes to 90 seconds using a unified agentic defence stack. Virgin Atlantic converted 40 hours a week of manual incident work into fully governed automation in under two weeks with a single junior analyst. Both facts are true simultaneously: the ceiling is higher than it has ever been, and the failure rate is higher than anyone was admitting. The variable separating the two populations is not model quality. It is whether the deployment was architected around supervision or around autonomy.
The domain resolved that question decisively this cycle, and it resolved it in favour of supervision. The retreat from full autonomy that began with penetration testing a fortnight ago has propagated across the domain as a general architectural preference. Cobalt's survey of 455 security professionals confirms AI-only penetration testing adoption collapsing from 29 percent in 2025 to 9 percent in 2026, with 47 percent now explicitly preferring a hybrid model and 78 percent reporting that fully automated scanning misses critical vulnerabilities; a separate 455-practitioner survey from Pentest-Tools found 90 percent of AI findings require significant manual validation and 27 percent of teams seeing false-positive or fabricated-issue rates above 25 percent, with hallucinated CVEs eroding trust in everything the tool subsequently reports. The same movement is visible in incident work: practitioners at ZenML, incident.io and Coroot are publicly scrapping autonomous agentic root-cause designs in favour of deterministic pipelines with curated context, and the new ORCA-Bench production-fidelity benchmark explains why — 25.3 percent accuracy on medium-difficulty tasks, 10 percent on hard ones, with the authors advocating a copilot rather than an autonomous model. The OpenRCA benchmark isolates the cause: in most failures the required evidence was already present, so the bottleneck is reasoning quality, not data access. Meanwhile the vendors most confident about autonomy are shipping approval gates as headline features: Dynatrace's Autonomous SRE Agent, AWS's event-driven remediation for Systems Manager patch failures, and CrowdStrike's ISO 42001-certified Charlotte AI all reached general availability with graduated-autonomy controls built in rather than bolted on.
The second movement is the emergence of the AI agent itself as the security subject that nobody can see. Bedrock Data's analysis of more than 70 petabytes across 180,000 datastores found 79 percent of agent identities able to reach stored secrets; the Cloud Security Alliance and Aembit report 68 percent of organisations unable to distinguish human from agent activity in their own logs; between 47 and 53 percent of production agents are already exceeding their intended permissions. IBM's 2026 Cost of a Data Breach study, covering 602 breached organisations, found shadow AI involved in 43 percent of breaches — up from 20 percent — at an average of $5.39 million per incident, while Kiteworks found 80 percent of organisations had experienced an AI-related security incident against only 27 percent running AI-specific data-loss controls, with a governance maturity index averaging 16.2 out of 100. Enforcement is rarer still: independent aggregation of 248 data points across 70-plus sources found just 3 percent of organisations operating automated machine-speed controls over AI behaviour and 11 percent able to block out-of-scope actions automatically, against 77 percent who have updated their AI security strategy on paper. Orca Security's telemetry from 1,200-plus production organisations found 56 percent had deployed agent frameworks with no formal safety controls at all, and that the share of those frameworks carrying publicly exploitable vulnerabilities rose from 0.2 percent in 2024 to 50.1 percent. The market has noticed: CrowdStrike's Falcon Shield identity business grew ARR fourfold year on year, including a named seven-figure healthcare deal signed explicitly to control AI agent access. Two counterweights deserve recording. On the offensive side, autonomy is no longer theoretical — Palo Alto's Unit 42 documented the Hermes Agent framework orchestrating DeepSeek to autonomously enumerate and exploit CVE-2026-33017 (CVSS 9.8), Cisco Talos analysis of captured threat-actor prompt logs concluded that vendor guardrails offer minimal protection against task decomposition and CTF framing, and peer-reviewed work from the University of Houston achieved an 83.4 percent average prompt-injection success rate against LLM-based SOC log analysis, with 8.4 percent still succeeding after mitigation. On the defensive side, the domain's most durable uncomfortable truth showed the first real sign of bending: Qualys disclosed 150 million patches deployed over the past year, of which 40 million were fully autonomous, at a rollback rate below 0.1 percent and with exposure windows compressed from 21 days to minutes.
Rollback has become the modal outcome, and governance reveals failure rather than preventing it. Seventy-four percent of enterprises rolled back customer-facing AI agents after deployment, and the best-governed organisations rolled back at 81 percent — higher, not lower. Set against Gartner's 42-point spread between claimed adoption (75 percent), actual deployment (17 percent) and production readiness (11 percent), and 46 percent of internally built AI security tooling being deprecated, the implication for buyers is uncomfortable: a rollback is not evidence that governance failed, and the absence of one is not evidence that anything is working.
The domain has chosen supervision over autonomy, and the benchmarks vindicate the choice. AI-only penetration testing fell from 29 to 9 percent adoption in a year while hybrid rose to 47 percent; practitioners at ZenML, incident.io and Coroot are replacing autonomous root-cause agents with deterministic pipelines; ORCA-Bench measures 25.3 percent accuracy on medium tasks and 10 percent on hard ones under production-fidelity conditions. The contrast with Stellar Cyber's supervised deployment — 64 percent of verdicts closed autonomously at 99.7 percent human agreement — is the clearest evidence available that the constraint is architectural rather than a limit of model capability: the same class of model performs well inside a bounded, verified loop and poorly outside one.
The agent has become a first-class security identity, and almost no one can observe it. Sixty-eight percent of organisations cannot distinguish agent activity from human activity in their own telemetry, 79 percent of agent identities can reach stored secrets, and roughly half of production agents already exceed their intended permissions. Only 3 percent run automated machine-speed controls on AI behaviour. Machine identities now outnumber human ones by ratios reaching 109 to 1, and the behavioural baselines built for human users — impossible travel, off-hours login, privilege escalation — do not describe agent misbehaviour, which develops as a sequence of individually unremarkable actions over time rather than as a discrete violation.
Offensive AI is operational, and the defensive AI is itself an attack surface. Unit 42 documented autonomous end-to-end exploitation of a CVSS 9.8 flaw by an LLM-orchestrated agent framework; CrowdStrike's telemetry across seven trillion daily events shows AI-agent-triggered detection leads running 2.5 times human-triggered activity, with LLMjacking campaigns issuing 200,000 API requests in two minutes and named actors weaponising disclosures within 24 hours. Cisco Talos concluded from captured attacker prompt logs that vendor safety controls are systematically bypassable, and University of Houston research put prompt injection success against LLM-based SOC log analysis at 83.4 percent on average, with 8.4 percent surviving mitigation. Adding an AI analyst now adds an attack surface as well as capacity.
Remediation physics is bending, but only for organisations that can fund the automation. Qualys' disclosure of 40 million autonomously patched vulnerabilities at a sub-0.1 percent rollback rate is the first large-scale evidence that automated fixing can keep pace with the discovery curve. It sits against an industry median remediation time that rose to 43 days from 32, only 26 percent of CISA-tracked critical vulnerabilities fully remediated, and vulnerability exploitation now the single largest initial access vector at 31 percent of breaches. Cost is the gate, and cost control is not keeping up: FinOps adoption for AI workloads leapt from 31 to 98 percent in two years, yet only 11 percent of enterprises forecast AI spend within 10 percent accuracy, 62 percent report cost surprises that altered business decisions, and cloud waste rose to 29 percent for the first time in five years.
Security teams ditch AI-only penetration testing (adoption-metric) — Cobalt's 455-practitioner survey is the clearest single data point for the domain's retreat from autonomy: adoption collapsed from 29 percent to 9 percent in a year, with 78 percent reporting that fully automated scanning misses critical vulnerabilities. https://www.reversinglabs.com/blog/automated-ai-pen-testing-out
'I bought the tool to save time, but I did more manual work than before' (adoption-metric) — the Pentest-Tools survey shows why hybrid won: 90 percent of AI findings require significant manual validation and over a quarter of teams see false-positive or fabricated-CVE rates above 25 percent, directly illustrating the trust erosion driving the supervision-over-autonomy shift. https://www.itpro.com/security/i-bought-the-tool-to-save-time-but-i-did-more-manual-work-than-before-pentesters-are-finding-more-bugs-with-ai-than-they-can-fix-theyre-even-battling-fake-hallucinated-cves
AI in Incident Response: Lessons from ORCA-Bench (research-paper) — a production-fidelity benchmark putting hard numbers on why ZenML, incident.io and Coroot are abandoning autonomous root-cause agents: 25.3 percent accuracy on medium-difficulty tasks and 10 percent on hard ones. https://www.linkedin.com/posts/roger-j-campbell-1a33771ab_orcabenchoncallagents260728545png-activity-7488888753994805249-uygm
Stellar Cyber Agentic Auto Triage: 19 Minutes Recovered Per Analyst Hour, 99.7% Human-AI Verdict Agreement (case-study) — the domain's strongest counter-evidence that the constraint is architectural, not capability: a 124-day, 138,475-alert production trial closing 64 percent of verdicts autonomously with human-level agreement inside a bounded, supervised loop. https://www.helpnetsecurity.com/2026/08/04/stellar-cyber-agentic-auto-triage/
Gartner's 2026 Security Operations Hype Cycle: AI SOC Agents at Peak of Inflated Expectations (industry-report) — Gartner's own analysts placing AI SOC Agents at 1-5 percent penetration and warning about vendor AI-washing is the clearest institutional confirmation that the capability narrative has outrun the deployment reality. https://nhimg.org/articles/gartners-2026-security-operations-hype-cycle-and-the-ai-soc-shift/
CrowdStrike Falcon Shield ARR Quadrupled YoY; Healthcare Enterprise Deploys for AI Agent Identity Governance (adoption-metric) — a market signal that buyers are treating agent identity as a distinct, urgent risk category, with a named seven-figure healthcare deal signed explicitly to control AI agent access. https://www.tradingview.com/news/zacks:691d826b2094b:0-can-identity-security-become-a-major-growth-driver-for-crowdstrike/
Bedrock Data launches Agent DLP: runtime enforcement for autonomous AI agents (product-ga) — the product launch and the underlying research travel together here: a scan of 70+ petabytes across 180,000 datastores found 79 percent of agent identities can reach stored secrets, the starkest single number behind "the agent has become a security identity nobody can see." https://bedrockdata.ai/news/bedrock-data-launches-agent-dlp-runtime-data-loss-prevention-built-for-ai-agents
IBM 2026 Cost of a Data Breach: Shadow AI in 43% of breaches (industry-report) — the most authoritative breach dataset in the domain (602 organisations) shows shadow AI's share of breaches more than doubling in a year to 43 percent at $5.39 million per incident, converting the observability gap into a hard cost figure. https://complexdiscovery.com/policy-without-control-the-ai-governance-gap-in-ibms-2026-cost-of-a-data-breach-report/
Qualys Q2 2026 Earnings: 40M Vulnerabilities Autonomously Patched, Exposure Windows Reduced to Minutes (adoption-metric) — the one place in this scan where remediation physics is visibly bending: 40 million autonomous patches at a sub-0.1 percent rollback rate, compressing exposure windows from 21 days to minutes, against an industry median that is otherwise getting worse. https://www.marketbeat.com/instant-alerts/qualys-q2-earnings-call-highlights-2026-08-04/
Chinese-Speaking Threat Actor Harnesses AI Models for Autonomous Cyberattacks (case-study) — Unit 42's documentation of the Hermes Agent framework autonomously exploiting a CVSS 9.8 flaw via DeepSeek confirms offensive autonomy is operational now, not theoretical, closing the loop on the domain's "AI analyst adds attack surface" tension. https://unit42.paloaltonetworks.com/autonomous-ai-cyber-attack-campaign/