{
  "id": "product-design",
  "label": "Product & Design",
  "description": "AI applied from user research through to shipped product experience. Wide maturity spread: A/B testing and analytics are established, prototyping and design systems are good practice, but nearly half the domain is bleeding-edge — generative UI, autonomous UX research, and AI-native product frameworks are experimental. Most practices are stalled, with more energy in tooling announcements than production adoption.",
  "icon": "🎯",
  "filters": [
    "building",
    "creating"
  ],
  "hasSummary": true,
  "hasExecSummary": true,
  "practiceCount": 13,
  "evidenceCount": 2372,
  "practices": [
    {
      "slug": "ab-test-design-and-analysis",
      "name": "A/B test design & analysis",
      "tier": "established",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that helps design experiments, determines sample sizes, analyses results, and identifies statistically significant outcomes. Includes automated experiment design and Bayesian analysis; distinct from marketing attribution which analyses campaign effectiveness rather than product experiments.",
      "evidenceCount": 202
    },
    {
      "slug": "accessibility-auditing-and-remediation",
      "name": "Accessibility auditing & remediation",
      "tier": "good-practice",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that audits digital products for accessibility compliance (WCAG) and suggests or implements remediation. Includes automated WCAG testing and remediation code generation; distinct from accessibility support in personal effectiveness which assists individual users rather than auditing products.",
      "evidenceCount": 194
    },
    {
      "slug": "behavioural-analytics-session-replay-and-interaction-patterns",
      "name": "Behavioural analytics — session replay & interaction patterns",
      "tier": "established",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that analyses session replays, heatmaps, and interaction patterns to identify UX issues and optimisation opportunities. Includes rage-click detection and interaction flow analysis; distinct from user research synthesis which analyses qualitative rather than behavioural data.",
      "evidenceCount": 195
    },
    {
      "slug": "competitive-product-analysis",
      "name": "Competitive product analysis",
      "tier": "good-practice",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that monitors and analyses competitor products, features, pricing, and positioning to inform product strategy. Includes automated feature comparison and competitor tracking; distinct from competitive positioning in marketing which analyses messaging rather than product capabilities.",
      "evidenceCount": 154
    },
    {
      "slug": "design-system-generation-and-enforcement",
      "name": "Design system generation & enforcement",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that generates design system components and enforces consistency across product interfaces. Includes component variant generation and design lint checking; distinct from brand-voice workflows which enforce written style rather than visual design.",
      "evidenceCount": 160
    },
    {
      "slug": "feature-prioritisation-and-roadmap-support",
      "name": "Feature prioritisation & roadmap support",
      "tier": "bleeding-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that helps prioritise features by synthesising customer signals, business impact, and engineering effort estimates. Includes RICE/ICE scoring assistance and roadmap scenario modelling; distinct from backlog management which organises work rather than prioritising outcomes.",
      "evidenceCount": 145
    },
    {
      "slug": "personalisation-engine-design-and-tuning",
      "name": "Personalisation engine design & tuning",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that designs and optimises personalisation rules and recommendation algorithms within products. Includes recommendation system tuning and personalisation strategy testing; distinct from personalised content delivery in marketing which uses personalisation rather than building it.",
      "evidenceCount": 227
    },
    {
      "slug": "product-analytics-interpretation-and-insight",
      "name": "Product analytics interpretation & insight",
      "tier": "bleeding-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that analyses product usage data and surfaces actionable insights about feature adoption, retention drivers, and user behaviour. Includes automated insight generation and metric explanation; distinct from automated EDA which analyses any data rather than specifically product metrics.",
      "evidenceCount": 172
    },
    {
      "slug": "requirements-prd-and-user-story-generation",
      "name": "Requirements, PRD & user story generation",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that generates product requirements documents, user stories, and acceptance criteria from research, feedback, and stakeholder input. Includes PRD drafting and story decomposition; distinct from feature prioritisation which ranks rather than defines requirements.",
      "evidenceCount": 178
    },
    {
      "slug": "user-journey-mapping-from-behavioural-data",
      "name": "User journey mapping from behavioural data",
      "tier": "good-practice",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that constructs user journey maps from actual behavioural data rather than assumptions, revealing real navigation patterns. Includes path analysis and journey clustering; distinct from customer journey analysis in customer ops which focuses on post-sale support rather than product usage.",
      "evidenceCount": 194
    },
    {
      "slug": "user-research-and-feedback-synthesis",
      "name": "User research & feedback synthesis",
      "tier": "good-practice",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that synthesises user research transcripts, survey data, and customer feedback into themes and feature signals. Includes automated affinity mapping and sentiment-driven feature prioritisation; distinct from product analytics which analyses behavioural data rather than qualitative feedback.",
      "evidenceCount": 183
    },
    {
      "slug": "ux-copy-generation-and-voice-enforcement",
      "name": "UX copy generation & voice enforcement",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that generates UX microcopy and enforces brand voice and tone guidelines across product interfaces. Includes context-aware microcopy creation and tone consistency checking; distinct from brand-voice workflows in marketing which target external content rather than product UI.",
      "evidenceCount": 183
    },
    {
      "slug": "wireframe-generation-and-design-to-code-conversion",
      "name": "Wireframe generation & design-to-code conversion",
      "tier": "bleeding-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that generates wireframes and prototypes from descriptions and converts designs into production code. Includes text-to-wireframe tools and Figma-to-code conversion; distinct from design system generation which creates reusable component libraries rather than individual screens.",
      "evidenceCount": 185
    }
  ],
  "summary": "## Where AI Stands in Product & Design\n\nProduct and design is where AI has done most to automate the making of artefacts and least to change how decisions get made. Wireframes, PRDs, user stories, microcopy, research summaries, battlecards and component libraries can now be generated in minutes. The vendors selling them have settled on the same plumbing: MCP servers that let Claude, ChatGPT and Cursor reach directly into Figma, Amplitude, Mixpanel, UserTesting, Klue, Microsoft Clarity and a growing list of brand-governance tools. Among practitioners, use of these tools is close to universal: 91% of designers use AI weekly and Maze finds 84% of researchers using it for synthesis. Organisations have not yet learned to trust the output and act on it. Tempo's survey of 300 planning leaders found 91% piloting AI but only 26% using it to decide what gets built. A 75-organisation benchmark found 99% investing and 1% claiming mature capability. Q3 enterprise surveys put 74% in production but only 5% quantifying the impact. The UK government's 1,000-licence Copilot trial recorded 72% user satisfaction and no measured productivity gain.\n\nThe domain's two long-established practices show where the younger ones are heading. A/B testing has had sophisticated platforms for a decade. Datadog paid $220M for Eppo, OpenAI paid $1.1B for Statsig, and CUPED, sequential testing and Bayesian engines are now commodities. Even so, false-positive rates above 26% persist at organisations such as Microsoft and Netflix, and Kohavi's replication work found most celebrated \"winning patterns\" fail to reproduce at 2.4 million users per test. Session replay is standard in Datadog, New Relic, Amplitude and PostHog, yet has the worst ratio of installs to sustained use in its category. Here, mature tooling has not produced good execution. AI is now repeating that pattern, faster, in prioritisation, analytics interpretation, requirements and design-to-code. In each of these, only a handful of well-governed teams run AI reliably in production.\n\nWhere there is real momentum, it comes from grounding, not generation. The deployments with clean numbers all constrain the model with a machine-readable source of truth. Coinbase's Figma Code Connect integration cut agent token use by 22.5% and task time by 22.3%, and stopped the agent inventing icons. OpenText halved ticket turnover from 4.2 to 2.1 days using JSON component contracts and lint gates. Amplitude's own data team raised agent answer accuracy to 80–90% with a semantic layer. Where that scaffolding is missing, the output erodes the qualities the domain exists to protect. WebAIM's audit of the top million sites finds 95.9% of home pages failing WCAG, the first worsening in six years, and attributes it partly to AI-assisted development. An audit of 276 production sites found 94% using design tokens but only 3.2% applying them consistently. This domain's bottleneck is organisational discipline, not model capability, and more of that discipline is needed each quarter.\n\n## What's New, 2026-09-12 to 2026-09-26\n\nThe past fortnight mostly confirmed existing patterns rather than moving them, but it sharpened two points. First, enforcement tooling became concrete. shadcn-ui/lint, an open-source linter built for coding agents with 2.8k GitHub stars, reports violations falling to zero after one correction round across more than 150 task runs on several frontier models. An MIT-licensed npm MCP server now ships 29 WCAG, performance and responsive-layout rules. OpenText's Figma-to-code pipeline and a Tallinn consultancy's token pipeline, with CI gates on every push, showed verification catching wrong values before they reach components. A 62-participant study showed why these gates matter: raw AI-generated screens gave 63% task accuracy against 100% for human-designed ones. The failures included clickable divs instead of buttons, missing focus trapping, and touch targets of 24–32px against a 44px guideline. Requirements generation produced its clearest positive cases so far. A six-developer Thoughtworks team went from 15 to 27 user stories per iteration over eight months, and IBM reports requirements gathering 30–40% faster after grounding its tool in platform, regulatory and coding standards. In regulated sectors, Jama Software set out the provenance records auditors will expect under ISO 26262 and DO-178C, which turns audit trails into an explicit barrier to adoption.\n\nSecond, the return-on-investment figures got worse. In brand-voice surveys, 48% now call AI adoption a \"massive disappointment\", up from 34%, and 29% report significant returns. Of 20 voice-enforcement tools surveyed, only two score voice fidelity numerically, and only one publishes how. This comes even as Writer shipped Agent Memory and Enterprise Brain, and brand rules began travelling over MCP servers from Pulumi, Frontify, Canva, Monotype, Adobe and Markup AI. Maze found 17% of research teams with integrated process ownership, against 84% using AI for synthesis, and Forrester's Q3 Wave named data governance, not AI capability, as the constraint on feedback programmes. Deque surveyed 200 engineering leaders: 64% named accessibility the top driver of post-production rework, even though their teams prompt agents to write accessible code. On the positive side, Amplitude's award entrants showed governed analytics agents paying off: Infosys, Algolia, and LIFULL, which cut metric investigations from 90 to 15 minutes. Vendor risk surfaced in wireframing, where Uizard has reportedly been frozen since Miro acquired it and Balsamiq ends Desktop sales on 31 December. None of this materially changed how mature any practice is.\n\n## Key Tensions\n\n- **Artefacts are cheap, judgement is not.** Every generative practice in the domain breaks at the same point: the output looks finished before it is right. Mountain Goat's worked example had AI writing acceptance criteria for an upload form when the real intent was extracting data from a photo. DORA data links a 25% rise in AI adoption with a 7.2% fall in delivery stability. Review capacity, not generation capacity, now sets the pace.\n\n- **Grounding matters more than model choice.** Reliable deployments stand out for the context they give the model, not for which model they use. Benchmarks cited by Atlan put analytics-agent accuracy at 21% without a semantic layer and 95% with one, and MCP-governed data access reaches about 90% on Claude, GPT and Gemini alike. The investment that pays off is unglamorous: tokens, metric definitions, spec files and lint gates.\n\n- **The plumbing ships faster than the proof.** MCP has become the domain's common connector. Klue, Mixpanel, UserTesting, Figma, Clarity and brand-governance vendors all expose one, and wireframe buyers now treat an MCP server as an evaluation criterion. Measurement has not kept up. Brand MCP servers distribute rules without judgement, 59% of enterprises report rising wasted AI spend, and only 5% quantify the impact of AI they already run in production.\n\n- **Generated output creates debt that older practices must pay down.** Accessibility auditing and design-system enforcement are becoming the clean-up crew for AI-generated interfaces. WebAIM records average errors per page up 10.1% to 56.1, and audits of Lovable-built storefronts found eight critical accessibility failures that did not change however the prompts were phrased. Code duplication in committed code has risen to 15.7%, against 8.3% in 2021.\n\n- **User behaviour is moving out of sight of the instruments.** Journey mapping, session replay and attribution assume the user is on your page. Only 8.8% of AI-influenced visits are tracked correctly, one estimate says more than 20% of purchase paths now begin in LLM conversations, and agents acting over MCP leave no trace in session replay. The engines that now shape those journeys have biases of their own: across 12 models, LLM shopping assistants recommended established vendors 31.6 percentage points more often than new entrants.\n\n## Top 10 Evidence Items\n\n1. **OpenText: agentic two-stage Figma-to-code pipeline with JSON component contracts and lint gates reduces ticket turnover 50% (4.2→2.1 days)** (case-study) — A clean grounding case: JSON component contracts and lint gates halved ticket turnover, the domain's best evidence that constraint beats generation. https://www.uxbybrett.com/case-study-opentext\n2. **Where AI-Generated Design Breaks UX Laws: user study of 62 participants finds 63% task accuracy on raw AI screens vs 100% human-designed; documents keyboard/touch-target/accessibility failures** (research-paper) — An empirical failure: ungrounded AI screens hit 63% task accuracy against 100% for human-designed ones, and the failures were accessibility and touch-target basics. https://us.headtopics.com/news/where-ai-generated-design-breaks-ux-laws-87702651\n3. **shadcn-ui/lint: agent-first Tailwind design system linter with near-universal enforcement across 150+ task runs** (significant-repo) — Shows enforcement tooling becoming concrete, with agent violations falling to zero after one correction round. https://github.com/shadcn-ui/lint\n4. **As Enterprise AI Enters the Value-Maxxing Era: What Development Teams Can Learn from Digital Accessibility** (opinion) — Deque's survey shows that telling agents to write accessible code still leaves compliance debt, the briefing's generated-output-creates-debt tension. https://www.unite.ai/ai-token-costs-accessibility-roi/\n5. **Enterprise AI Has an Execution Problem: What the Latest Research Says** (news-coverage) — Market data for the plumbing-outruns-proof tension: 74% run AI in production, yet only 5% quantify the impact. https://gritdaily.com/enterprise-ai-has-an-execution-problem/\n6. **Atlan: Data-readiness impact on analytics agent accuracy — peer-reviewed metrics from 21% to 95% via semantic layer, 90% to 98%+ via governed definitions** (opinion) — Quantifies the grounding thesis: analytics-agent accuracy of 21% without a semantic layer against 95% with one. https://atlan.com/know/ai-agent/data-for-ai/what-makes-data-ai-ready/\n7. **Session replay adoption crisis: highest tool-abandonment ratio in category** (opinion) — Shows mature tooling failing to produce good execution: session replay has the worst install-to-sustained-use ratio in its category. https://www.dynoweb.app/blog/best-shopify-session-replay-apps\n8. **A/A Testing and Pre-Result Platform Validation** (tutorial) — Documents endemic false-positive problems in A/B platforms at Microsoft, evidence that maturity of tooling does not mean maturity of practice. https://dev.to/4thwithme/ab-testing-how-to-test-your-test-enn\n9. **Thoughtworks: Quantifying AI adoption—from initial challenges to doubling speed** (case-study) — Requirements generation's clearest positive case, and it came from a small team with lightweight, grounded workflows. https://www.thoughtworks.com/en-us/insights/blog/machine-learning-and-ai/quantifying-ai-adoption-from-initial-challenges-to-doubling-speed\n10. **Marketing Funnel Fragmentation: How AI Broke Attribution** (opinion) — Shows user behaviour moving out of sight of the instruments, with only 8.8% of AI-influenced visits tracked correctly. https://aisearch.similarweb.com/blog/marketing-funnel-fragmentation/",
  "execSummary": "**The headline:** AI now drafts most product and design work in minutes, but few companies trust it to decide anything. The ones pulling ahead gave it clean rules and definitions first.\n\n### The Picture\n\nMost product and design teams already use AI every week, and 91% of designers do. Very few have turned that use into decisions. Of 300 planning leaders surveyed, 91% are piloting AI but only 26% use it to decide what gets built, and while three quarters of enterprises run AI in production, only 5% can put a number on the payoff. The small group pulling ahead did unglamorous groundwork first: agreed metric definitions, machine-readable design rules and automated checks, so the AI works from one source of truth. If your teams are producing faster without that groundwork, you are in the pack and quietly building up rework.\n\n### This Fortnight\n\n- **Tools that automatically check AI-generated interface code against company design rules are now freely available.** An open-source checker built for coding agents (software that acts on its own without being prompted) is gaining traction, and OpenText used similar automated checks to halve the time taken to close design tickets, from 4.2 to 2.1 days. Enforcing your own design standards on AI output is now a cheap engineering decision, not a research project.\n\n- **A 62-person study found that users completed tasks correctly only 63% of the time on raw AI-generated screens, against 100% on human-designed ones.** The failures were basic: clickable areas that were not real buttons, and touch targets too small for fingers. Unchecked AI screens are a usability and accessibility liability, not a shortcut.\n\n- **AI-drafted requirements produced their clearest wins yet, but only where the tool was grounded in company standards.** Over eight months, a six-developer Thoughtworks team went from 15 to 27 user stories per iteration. IBM reported similar gains after feeding its tool platform, regulatory and coding standards. The gain comes from the context you supply, not from the tool you buy.\n\n- **Returns from AI brand-voice tools got worse: 48% of adopters now call it a massive disappointment, up from 34% a year earlier.** This is happening even as Adobe, Canva, Frontify and others start piping brand rules straight into AI tools. Almost none of the voice-checking tools can show how they score whether output is on-brand, so ask that question before renewing.\n\n- **Named companies showed AI analytics assistants paying off in production where the underlying data was governed.** Japanese property portal LIFULL cut metric investigations from 90 minutes to 15, and Infosys and Algolia reported production wins of their own. What they share is governed data, which is where to invest before you expand analytics AI.\n\n### Coming Up\n\n- **Balsamiq stops selling its desktop product on December 31, and Uizard has reportedly been frozen since Miro acquired it.** Wireframing tools are consolidating, and buyers now treat an MCP server (a connector standard for AI tools) as a baseline requirement. List the design tools your teams depend on and plan any exits before renewal dates, not after.\n\n- **US state and local government accessibility deadlines arrive from April 2027, and European enforcement is already under way.** A French court upheld a €500-a-day penalty against Carrefour even though its site passed most automated checks, because what counts is the experience of real disabled users. If AI is building your customer-facing screens, require accessibility testing in the build process now.\n\n- **In safety-regulated industries, auditors will expect a record of how every AI-drafted requirement was produced.** Jama Software has set out the audit trail that automotive and aerospace standards will demand. Without it, AI output counts as undocumented work. If you operate in a regulated sector, get legal and engineering to agree logging rules before teams scale up AI drafting.\n\n### What's Hard About This\n\n- **AI makes drafts look finished before they are right, so the pace is now set by how much your people can review.** DORA research, a long-running study of software delivery, links a 25% rise in AI adoption to a 7.2% fall in delivery stability. Speeding up generation without adding reviewers mostly produces faster rework.\n\n- **Which AI you buy matters less than the definitions you give it.** Benchmarks put analytics-assistant accuracy at 21% without a shared dictionary of business metrics and 95% with one. The money that pays off goes on agreed definitions, design rules and spec files, not on the next tool.\n\n- **More and more customer journeys now start inside AI chatbots, where your analytics cannot see them.** One estimate finds only 8.8% of AI-influenced website visits are tracked correctly, and AI tools acting on a customer's behalf leave no trace in session recordings. Treat conversion and attribution reports as incomplete, and ask your analytics team how they plan to measure this.",
  "headline": "AI now drafts most product and design work in minutes, but few companies trust it to decide anything. The ones pulling ahead gave it clean rules and definitions first.",
  "execSummarySections": [
    {
      "id": "the-picture",
      "title": "The Picture",
      "body": "Most product and design teams already use AI every week, and 91% of designers do. Very few have turned that use into decisions. Of 300 planning leaders surveyed, 91% are piloting AI but only 26% use it to decide what gets built, and while three quarters of enterprises run AI in production, only 5% can put a number on the payoff. The small group pulling ahead did unglamorous groundwork first: agreed metric definitions, machine-readable design rules and automated checks, so the AI works from one source of truth. If your teams are producing faster without that groundwork, you are in the pack and quietly building up rework."
    },
    {
      "id": "this-fortnight",
      "title": "This Fortnight",
      "body": "- **Tools that automatically check AI-generated interface code against company design rules are now freely available.** An open-source checker built for coding agents (software that acts on its own without being prompted) is gaining traction, and OpenText used similar automated checks to halve the time taken to close design tickets, from 4.2 to 2.1 days. Enforcing your own design standards on AI output is now a cheap engineering decision, not a research project.\n\n- **A 62-person study found that users completed tasks correctly only 63% of the time on raw AI-generated screens, against 100% on human-designed ones.** The failures were basic: clickable areas that were not real buttons, and touch targets too small for fingers. Unchecked AI screens are a usability and accessibility liability, not a shortcut.\n\n- **AI-drafted requirements produced their clearest wins yet, but only where the tool was grounded in company standards.** Over eight months, a six-developer Thoughtworks team went from 15 to 27 user stories per iteration. IBM reported similar gains after feeding its tool platform, regulatory and coding standards. The gain comes from the context you supply, not from the tool you buy.\n\n- **Returns from AI brand-voice tools got worse: 48% of adopters now call it a massive disappointment, up from 34% a year earlier.** This is happening even as Adobe, Canva, Frontify and others start piping brand rules straight into AI tools. Almost none of the voice-checking tools can show how they score whether output is on-brand, so ask that question before renewing.\n\n- **Named companies showed AI analytics assistants paying off in production where the underlying data was governed.** Japanese property portal LIFULL cut metric investigations from 90 minutes to 15, and Infosys and Algolia reported production wins of their own. What they share is governed data, which is where to invest before you expand analytics AI."
    },
    {
      "id": "coming-up",
      "title": "Coming Up",
      "body": "- **Balsamiq stops selling its desktop product on December 31, and Uizard has reportedly been frozen since Miro acquired it.** Wireframing tools are consolidating, and buyers now treat an MCP server (a connector standard for AI tools) as a baseline requirement. List the design tools your teams depend on and plan any exits before renewal dates, not after.\n\n- **US state and local government accessibility deadlines arrive from April 2027, and European enforcement is already under way.** A French court upheld a €500-a-day penalty against Carrefour even though its site passed most automated checks, because what counts is the experience of real disabled users. If AI is building your customer-facing screens, require accessibility testing in the build process now.\n\n- **In safety-regulated industries, auditors will expect a record of how every AI-drafted requirement was produced.** Jama Software has set out the audit trail that automotive and aerospace standards will demand. Without it, AI output counts as undocumented work. If you operate in a regulated sector, get legal and engineering to agree logging rules before teams scale up AI drafting."
    },
    {
      "id": "whats-hard-about-this",
      "title": "What's Hard About This",
      "body": "- **AI makes drafts look finished before they are right, so the pace is now set by how much your people can review.** DORA research, a long-running study of software delivery, links a 25% rise in AI adoption to a 7.2% fall in delivery stability. Speeding up generation without adding reviewers mostly produces faster rework.\n\n- **Which AI you buy matters less than the definitions you give it.** Benchmarks put analytics-assistant accuracy at 21% without a shared dictionary of business metrics and 95% with one. The money that pays off goes on agreed definitions, design rules and spec files, not on the next tool.\n\n- **More and more customer journeys now start inside AI chatbots, where your analytics cannot see them.** One estimate finds only 8.8% of AI-influenced website visits are tracked correctly, and AI tools acting on a customer's behalf leave no trace in session recordings. Treat conversion and attribution reports as incomplete, and ask your analytics team how they plan to measure this."
    }
  ],
  "url": "https://www.thestateofplay.ai/domain/product-design",
  "license": "CC BY 4.0",
  "licenseUrl": "https://creativecommons.org/licenses/by/4.0/",
  "generatedAt": "2026-10-01"
}