Perly Consulting │ Beck Eco

The State of Play

A living index of AI adoption across industries — where established practice meets the bleeding edge
UPDATED DAILY

The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.

The Daily Dispatch

A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.

Pick a role above to explore practices

BLEEDING EDGE

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🔬 RESEARCH & KNOWLEDGE
⚖️ LEGAL, COMPLIANCE & RISK
🎧 CUSTOMER OPERATIONS
🏛️ AI GOVERNANCE & SAFETY
📊 DATA & ANALYTICS
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💼 SALES & REVENUE
🎬 CREATIVE & GENERATIVE MEDIA
👁️ COMPUTER VISION & SENSING
💹 FINANCE & ACCOUNTING
🔄 OPERATIONS & PROCESS AUTOMATION
🚗 AUTONOMOUS SYSTEMS & VEHICLES
🦾 PHYSICAL AI & ROBOTICS
🎓 EDUCATION & LEARNING
PERSONAL EFFECTIVENESS

LEADING EDGE

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🔬 RESEARCH & KNOWLEDGE
⚖️ LEGAL, COMPLIANCE & RISK
🎧 CUSTOMER OPERATIONS
🏛️ AI GOVERNANCE & SAFETY
📊 DATA & ANALYTICS
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💼 SALES & REVENUE
🎬 CREATIVE & GENERATIVE MEDIA
👁️ COMPUTER VISION & SENSING
💹 FINANCE & ACCOUNTING
🔄 OPERATIONS & PROCESS AUTOMATION
👥 PEOPLE & TALENT
🚗 AUTONOMOUS SYSTEMS & VEHICLES
🦾 PHYSICAL AI & ROBOTICS
🎓 EDUCATION & LEARNING
PERSONAL EFFECTIVENESS

GOOD PRACTICE

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🔬 RESEARCH & KNOWLEDGE
⚖️ LEGAL, COMPLIANCE & RISK
🎧 CUSTOMER OPERATIONS
🏛️ AI GOVERNANCE & SAFETY
📊 DATA & ANALYTICS
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💼 SALES & REVENUE
🎬 CREATIVE & GENERATIVE MEDIA
👁️ COMPUTER VISION & SENSING
💹 FINANCE & ACCOUNTING
🔄 OPERATIONS & PROCESS AUTOMATION
👥 PEOPLE & TALENT
🚗 AUTONOMOUS SYSTEMS & VEHICLES
🦾 PHYSICAL AI & ROBOTICS
🎓 EDUCATION & LEARNING
PERSONAL EFFECTIVENESS

ESTABLISHED

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💹 FINANCE & ACCOUNTING
👥 PEOPLE & TALENT

🎯 Product & Design

AI applied from user research through to shipped product experience. Wide maturity spread: A/B testing and analytics are established, prototyping and design systems are good practice, but nearly half the domain is bleeding-edge — generative UI, autonomous UX research, and AI-native product frameworks are experimental. Most practices are stalled, with more energy in tooling announcements than production adoption.

13 practices: 2 established, 4 good practice, 4 leading edge, 3 bleeding edge

Where AI Stands in Product & Design

The tooling contest in this domain finished some time ago. Every practice we track now has production-grade agentic AI available to it — autonomous analytics agents that investigate anomalies unprompted, specification agents that draft PRDs and score backlogs, design agents that generate and police component libraries, research platforms that synthesise hundreds of interviews and dispatch the findings straight into a ticketing system. What changed over the past fortnight is that these tools stopped being islands. The Model Context Protocol has become the domain's connective tissue: Figma's MCP server reached general availability on a stateless protocol, Anthropic shipped Claude Design with a closed loop back into Claude Code, Webflow's MCP 2.0 added governance controls with named enterprise deployments at Arkose Labs and Amazon Ads, UserTesting shipped an MCP server letting researchers create studies without leaving a chat window, Mixpanel's AI suite went GA across 29,000 organisations, and Dovetail's Channels 2.0 now classifies feedback from thirty-plus sources and dispatches evidence-backed ideas into Claude Code, Cursor, Linear or Jira with account context attached. The research-to-spec-to-design-to-code-to-audit pipeline is being wired together end to end.

That is the good news, and it is genuinely significant. The uncomfortable news arrived in the same window: the handoffs now being automated are precisely where quality has always been enforced. Practitioner testing published in late July found that 100% of AI-generated front-end code sampled failed accessibility standards, with the root cause identified as models faithfully reproducing inaccessible training data rather than sloppy implementation. AI design-to-code tools generate 322% more privilege-escalation paths than human-authored equivalents. Not one of 200 surveyed site-reliability leaders reports high confidence in AI-generated code once deployed. Two-thirds of product managers still substantially rewrite AI-generated PRDs. Design-system agents plateau at roughly 95% compliance against a well-governed system, and a documented case study showed fourteen individually clean AI-generated buttons producing eleven undocumented padding values, with four times the year-two maintenance cost. Automating a handoff does not remove the work the handoff was doing.

The domain's maturity spread remains unusually wide, and it is worth being precise about it. Two practices are genuinely established infrastructure — A/B testing, where the question is execution quality rather than adoption, and behavioural session replay, now a baseline feature rather than a differentiator. A middle band of mainstream-but-throttled capabilities covers accessibility auditing, competitive analysis, journey mapping and research synthesis: the tools work, the organisations mostly do not use them well. The frontier — feature prioritisation, product analytics interpretation, and wireframe-to-code conversion — remains genuinely unsettled: powerful, demonstrably risky, and far from routine. Cutting across all of it is a demand-side ceiling that is close to unique to this domain, because so much of its output is consumer-facing. YouGov polling of roughly 10,000 people finds 51% uncertain or sceptical about AI-generated content and only 15% trusting a brand more for using it; DoubleVerify's 22,000-respondent study finds 42% saying low-quality AI advertising damages brand trust. No amount of tooling maturity lifts that ceiling.

What's New, 2026-07-18 to 2026-08-01

Design-to-code reached a platform-consolidation inflection, with four significant shipments inside two weeks. Figma's workflow lab (22 July) lets designers deploy code changes directly to GitHub without raising an engineering ticket, closing the handoff loop at production scale; its MCP server hit GA on 28 July with a stateless protocol, removing the session-affinity barrier to cloud deployment. Anthropic's Claude Design (19 July) is the first vendor-integrated design-to-production pipeline, pairing conversational wireframing with closed-loop handoff to Claude Code. Builder 2.0 (23 July) introduced multiplayer editing in which designers, PMs and QA edit live code on production branches, collapsing design review into concurrent iteration. Webflow MCP 2.0 (21 July) added brand governance, with over 30% enterprise adoption reported. Elsewhere, Productboard's Spark completed GA on 31 July across feedback analysis, spec drafting, competitive research and post-launch evaluation; Adobe shipped Coworker Chat and a Data Insights Agent into Customer Journey Analytics with Corteva Agriscience and ResMed confirmed in production; Meta deployed an LLM-driven agentic recommender for connected-TV discovery, with Alibaba Cloud, Huawei Noah and Baidu validating multi-agent recommendation architectures in A/B-tested production; and agentic accessibility remediation crossed from announcement to GA as Siteimprove shipped bulk PDF remediation, Evinced launched Autopilot, Resolve and Harness for CI/CD and batch fixing, and Digital.ai embedded Deque-powered WCAG scanning directly into mainstream QA reports.

Against that, the fortnight produced the most pointed counter-evidence this domain has seen in months. Sinch surveyed 2,527 senior decision-makers across ten countries and found 74% have already rolled back or shut down a customer-facing AI agent after deployment — while 62% still have agents live and 98% are increasing AI investment. The finding that matters is counterintuitive: organisations with mature governance frameworks roll back more, at 81%, because rigorous testing surfaces failures that weaker organisations simply never detect. Root cause analysis put 80% of those failures on data fragmentation and API integration complexity, not model accuracy. That aligns with an independent analysis of more than 10,000 enterprise AI failures showing hallucination now accounts for under 10% of incidents, down from being the dominant 2023–24 concern, with execution and escalation breakdowns at 31.1% taking its place. Meanwhile the Ninth Circuit revived a California wiretapping claim in Mikulsky v. Bloomingdale's over FullStory's keystroke and URL capture, with the California CIPA docket now at 3,968 cases — 81% of the national total — and a sector audit finding session replay running on 59 major US healthcare sites despite the patient-data capture risk. Netflix detailed a production anytime-valid inference platform providing real-time quality control across 1,000-plus concurrent experiments, while experimentation leaders at DoorDash, Robinhood and Intuit reframed maturity as a shift from building more tests to ensuring decision systems actually change outcomes. No practice changed tier or trend this cycle. That stability is itself the signal: the tooling contest has resolved, and the bottleneck has settled firmly onto plumbing, governance and human review capacity.

Key Tensions

  • Automating a handoff concentrates risk where the checkpoint used to be. The MCP wiring shipped this fortnight removes the human pauses between research, spec, design, code and audit — and every piece of new quality evidence lands precisely in those gaps. All AI-generated code sampled in late-July practitioner testing failed accessibility standards; AI design-to-code output carries 322% more privilege-escalation paths; Faros's analysis of 22,000 developers shows AI adoption raising throughput 33.7% while bugs per developer rose 54% and incidents per pull request rose 242.7%. Designer confidence (91% report better quality, 89% report faster work) and operator reality have diverged sharply, with senior engineers now losing up to a third of their week to triaging AI-generated failures.

  • The dominant failure mode has moved from hallucination to plumbing. This is the single most consequential reframing in the current evidence. Analysis of over 10,000 enterprise AI deployments puts hallucination below 10% of failures and execution/escalation breakdowns at 31.1%; Sinch attributes 80% of agent rollbacks to data fragmentation and integration complexity rather than model accuracy. The industry spent three years hardening against confidently invented outputs and largely succeeded. The money and attention now need to move to data quality, defined scope, observability and identity governance — where only 18% of leaders report confidence even though 40% already run agents in production.

  • Enforcement, not generation, has become the deliverable. Across four separate practices the same conclusion crystallised this cycle. In design systems, practitioner consensus now positions CI-based mechanical checks (token resolution) and judgment checks (rubric-based grading) as the actual product, with Alibaba's Canvas shipping a recipe layer encoding composition intent rather than component inventory. In UX copy, vendor guidance converged on five-step guardrail patterns, four-layer governance architectures and weekly audit rubrics, after one mid-market brand discovered 43% of its product descriptions had drifted off-brand through untracked prompt changes. In requirements, Altium's agentic offering continuously quality-scans imported documents behind human-approval gates. And the hard part is human: practitioner analysis finds most brand-voice programmes fail not for lack of tooling but because writers will not adopt the enforcement layer into daily workflow.

  • ROI evidence is bifurcating on survivorship, not on truth. SoundHound's production data shows 96% of live agentic systems met or exceeded ROI targets, with 74% positive inside year one. A synthesis of BCG, KPMG, MIT and McKinsey 2026 surveys finds only 5–8% of organisations reporting measurable ROI against 98% adoption and average budgets of $186M. Both are correct. Gartner and CISA data explain the gap: 88% of agent pilots never reach production at all, and Sinch's 74% rollback rate culls much of the remainder. Systems that survive to steady-state production do pay — Clayco reports a 93% productivity gain and $12M projected return; Adobe's Firefly ARR approached $300M with AI-first ARR past $500M, tripling year on year. The question for any given organisation is not whether AI in product work can pay, but whether its own deployment will survive long enough to.

  • The measurement layer has become a legal target. Behavioural analytics is this domain's most established practice and is now among its most legally exposed. The Ninth Circuit's revival of the Mikulsky CIPA claim gives the California wiretapping theory appellate life against session replay vendors; the state's tracker stands at 3,968 cases, 81% of the national docket, spanning retail, technology and professional services, and healthcare sites are running replay tools over pages that capture protected health information. In accessibility the parallel is starker: 983 of 3,948 US accessibility lawsuits filed in 2025 — 24.9% — targeted sites that already had a compliance overlay installed, and the Department of Justice extended ADA Title II deadlines to 2027–28 explicitly because generative AI "does not yet reliably automate the remediation of inaccessible content at scale." Buying a tool is not a defence; documented source-code remediation is.

Top 10 Evidence Items

  1. Every AI-Generated Component Is Correct. Together They're Chaos. (case-study) — The domain's cleanest illustration of "enforcement, not generation, is the deliverable": fourteen individually correct AI-generated buttons produced eleven undocumented padding values and quadrupled year-two maintenance cost, showing that automation without a governing checkpoint accelerates debt rather than removing it. https://www.codexical.com/posts/2026-07-19-ai-generated-ui-design-debt
  2. Why AI-generated code keeps failing accessibility checks (opinion) — Practitioner root-cause analysis behind the fortnight's starkest statistic — 100% of sampled AI-generated front-end code failing accessibility standards — because models faithfully reproduce inaccessible patterns baked into their training data, not because implementation is sloppy. https://sturit.medium.com/why-ai-generated-code-keeps-failing-accessibility-checks-72942319af81
  3. 983 ADA lawsuits in 2025 hit sites that already had an accessibility widget (adoption-metric) — Directly substantiates the "buying a tool is not a defence" tension: nearly a quarter of 2025 US accessibility litigation targeted sites that had already installed a compliance overlay, showing tooling adoption and actual remediation are two different things. https://screenmy.site/blog/accessibility-widget-lawsuits-2025
  4. Mikulsky v. Bloomingdale's: The 9th Circuit CIPA Reversal (news-coverage) — The appellate decision that gives California's wiretapping theory new legal teeth against session-replay vendors, turning the domain's most mature, "baseline feature" practice into one of its most exposed. https://consentpixel.com/blogs/mikulsky-v-bloomingdales/
  5. Healthcare tracking report 2026: Are healthcare companies one audit away from a compliance crisis? (industry-report) — Independent audit of 59 major US healthcare sites finding session replay tools capturing pages with protected health information, the sector-specific edge case behind the domain's "measurement layer as legal target" tension. https://piwik.pro/healthcare-website-tracking-report-2026/
  6. Enterprise AI Failure Modes Have Shifted. Hallucinations Are No Longer the Problem. (industry-report) — The single most consequential reframing in this cycle's evidence: analysis of 10,000-plus enterprise AI failures puts hallucination under 10% of incidents while execution and escalation breakdowns hit 31.1%, redirecting where governance investment needs to go. https://forkast.news/enterprise-ai-failure-modes-have-shifted-hallucinations-are-no-longer-the-problem/
  7. Enterprise AI ROI Measurement in 2026: Only 5-8% of Companies See Real Returns on $186M Budgets (industry-report) — The survivorship-bias half of the ROI tension: a synthesis of BCG, KPMG, MIT and McKinsey 2026 data showing measurable returns concentrated in a small minority of deployments despite near-universal adoption, the counterweight to vendor success-story data. https://valueaddvc.com/blog/enterprise-ai-roi-in-2026-what-companies-are-actually-measuring-and-finding
  8. Workflow lab: Deploying designs directly with Figma Make (product-ga) — Figma's own account of the shipment that best captures the fortnight's platform-consolidation inflection, letting designers push code changes straight to GitHub and closing the research-to-spec-to-design-to-code loop the summary describes. https://www.figma.com/blog/workflow-lab-deploying-designs-directly-with-figma-make/
  9. Channels 2.0: from customer feedback to code in one click (product-ga) — Dovetail's official announcement of the MCP-wired pipeline that classifies feedback from thirty-plus sources and dispatches evidence-backed ideas straight into Claude Code, Cursor, Linear or Jira — the connective-tissue story at the centre of "Where AI Stands." https://dovetail.com/blog/suns-out-channels/
  10. How AI drove Shopify back to clean code (news-coverage) — Named production case of a major platform re-architecting its design system specifically to be machine-legible, evidence that "the design system's real job is catching the AI" is now operational reality rather than commentary. https://www.theregister.com/devops/2026/07/25/how-ai-drove-shopify-back-to-clean-code/5277901