The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.
A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.
Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail
AI that synthesises user research transcripts, survey data, and customer feedback into themes and feature signals. Includes automated affinity mapping and sentiment-driven feature prioritisation; distinct from product analytics which analyses behavioural data rather than qualitative feedback.
AI-powered synthesis of qualitative user feedback -- interviews, surveys, open-ended responses -- is a proven capability with mature tooling and documented enterprise ROI. The question is no longer whether it works but why it has stalled. Over half of UX researchers now use AI for synthesis, and vendor-commissioned studies report ROI figures from 236% to 665%. Yet adoption remains confined to large enterprises with established research operations, and the category shows clear signs of a maturity plateau. The binding constraints are organisational, not technical: integration complexity favours vendor-led deployments over internal builds, practitioners hide AI tool use from colleagues even while reporting productivity gains, and hallucination risks demand human oversight that erodes the speed advantage. A critical capability boundary has emerged: AI excels at descriptive synthesis tasks (extracting, coding, clustering feedback) but fails at interpretive synthesis requiring judgment about meaning and implication—a 192-study systematic review confirms GenAI effective for coding (62.5% of studies) but only 10.9% achieve pattern-based theme generation. A deeper constraint affects global research programs: LLMs systematically bias toward Western moral frameworks when interpreting human values, and AI moderators produce shallower data and miss cultural cues when interviewing non-Western participants—exposing what one practitioner analysis termed 'epistemic colonialism automated at scale.' Consensus has settled on AI as an efficiency multiplier for mechanical theme extraction and summarisation -- compressing weeks of affinity mapping into hours -- rather than a replacement for interpretive research judgment. The result is a good-practice capability whose rollout challenge is less about tooling maturity and more about embedding AI-assisted synthesis into research workflows with validated accuracy for global populations, governance preventing hallucinated synthesis from driving product decisions, and clear acknowledgment that humans own meaning and judgment.
Three platforms dominate the vendor landscape -- UserTesting, Dovetail, and Thematic -- each with GA AI features and named enterprise customers including Amazon, Canva, Meta, and Mayo Clinic. UserTesting's latest Forrester TEI study documents 665% ROI with measurable business outcomes: 60% conversion improvement and 140% lift in customer spend; independent G2 market leader verification (July 2026) shows sustained adoption with 96% 4–5-star ratings from 283 verified customers and 89% recommending the platform. Dovetail has pushed furthest on workflow integration, with 3.0 (Fall 2025) and May 2026 releases shipping AI Agents, Dashboards, Figma integration, and Chat--a multi-source synthesis interface with transparent 'Show Thinking' panels that reveal which sources were scanned and how many interviews read, directly addressing transparency concerns raised in practitioner research. July 2026's Sun's Out launch closes critical integration gaps: 10 new feedback integrations (Qualtrics, Salesforce, Pendo, PostHog, ServiceNow, HubSpot, SurveyMonkey) auto-pull customer feedback from support tickets, surveys, and product events without manual export, while 8 MCP connectors (Canva, Linear, Hex, Salesforce, Slack, Notion, Gmail, Snowflake) push synthesized insights back into teams' existing tools via Chat and Agents—enabling continuous feedback synthesis and action without tool-switching. Channels 2.0 automates the full feedback-to-action pipeline: classifies qualitative feedback from 30+ sources, connects individual quotes to customer business context (account plan, ARR), synthesizes into evidence-backed ideas, and dispatches to Claude Code, Cursor, Linear, or Jira with full customer context attached, demonstrating synthesis-to-action operational maturity. June 2026 releases accelerate momentum: Dovetail's Deep Research mode enables reasoning across multi-source customer data (research sessions, support tickets, sales calls) for complex strategic synthesis; quantified outcomes show product managers reducing workload from 100 to 10 hours per week and teams saving 38+ hours weekly; deployment scale has accelerated sharply, with PM interview cadence doubling from 4 to 9 per quarter at median penetration and top-quartile teams running 21+ interviews quarterly, driven by maturation of AI moderation and async participation workflows. July 2026 product updates signal ecosystem-wide AI-tool integration: UserTesting's MCP Server (July 29) enables researchers to create studies, recruit participants, and analyze insights without switching from Claude, ChatGPT, Figma Make, or other AI clients; the enhanced Figma plugin now supports multiple tasks and prototypes in a single study, reducing design-research iteration cycles. Thematic documents 92% time reduction in feedback analysis with $4.8M incremental revenue generation. Outset extended synthesis beyond text to multi-modal analysis (facial cues, physical interaction) via April 2026 visual intelligence suite. Enterprise adoption has accelerated: 61% of enterprises with >1,000 employees deployed AI text analytics (up from 38% in 2023), with synthesis accuracy reaching 87-92% on structured feedback tasks and ROI of 3.2x within two years for mature VOC programs. Insights teams adoption jumped to 72% (up from 31% in 2024), with synthesis cost compressed from $50-500K to $2-8K per study. Peer-reviewed benchmarking (May 2026, arXiv) validates that LLM-based synthesis achieves high speed (28x improvement over manual coding in 20 minutes) while confirming quality trade-offs: exploratory tasks benefit from AI acceleration but precision-critical work requires human validation to shift quality burden back to researchers. Market analysis projects the text analytics segment reaching $18 billion by 2028. However, emerging research reveals critical limitations for global deployment: PNAS study (July 2026) confirms LLMs systematically prioritize Western moral frameworks when interpreting values, underestimating non-Western participant concerns. Empirical testing shows AI-moderated interviews with Afro-descendant and Latine participants produce shallower data and miss cultural cues versus human moderation, while practitioner analysis documents how Western-trained AI models misclassify non-Western communication patterns as 'tangential' or 'low relevance'—exposing a capability ceiling for inclusive global research that training alone cannot resolve. Critical practitioner assessment (July 2026) identifies specific implementation failure modes requiring mandatory human oversight: context compression during AI summarization, gap-filling with assumptions rather than reported uncertainty, premature contradiction resolution replacing messy tensions where insight lives, and confidence inflation in outputs--demonstrating that careless AI deployment accelerates synthesis misuse and requires methodological discipline. Meanwhile, adoption barriers persist on the ground: 93% of collected customer feedback never gets analyzed, only 17% of organizations use LLMs for feedback analytics despite tool availability, and organizational readiness gaps (51% of researchers lack evaluation processes, 13% have formal integration) remain the binding constraint on expansion beyond large enterprises.
Mid-2026 adoption data confirms rapid expansion at breadth but reveals persistent organizational barriers. Perspective AI's survey of 300 product teams (June 2026) documents synthesis as mainstream: 88% use AI for analysis and feedback, with 80% incorporating it somewhere in workflow; research democratization has tripled from 8% to 22% of organizations where research is essential to all strategic levels. Cycle-time compression is dramatic: teams that previously took 3 weeks now complete synthesis in 3 days (91% reduction). However, organizational readiness remains fragmented: 87% of 400+ researchers across academia and enterprise use AI weekly, yet only 13% have formal integration and 51% lack evaluation processes despite 52% verifying outputs—indicating individual adoption decoupled from institutional governance. The competitive landscape has shifted sharply: traditional platforms' $40k enterprise contracts compete against AI-native alternatives at $30-80/month flat, with concurrent interview scaling from 4-6 human moderators to hundreds simultaneously, representing 1,000x cost compression. Yet critical risks persist in deployment: 47% of enterprise AI users have made major business decisions based on hallucinated synthesis content, with documented examples of entire features built on fabricated user preference findings. Burke, Inc.'s synthetic data analysis (June 2026) documents that LLM-based synthetic panels produce false conclusions in 60% of tested business scenarios—a structural limitation independent of model selection. Academic research confirms that synthetic respondents (AI-generated personas) provide only 1.4 percentage-point improvement over unpersonalized baselines across 1,784 real human studies, with documented distortions limiting their use to rehearsal and ideation rather than evidence gathering. Practitioner quality concerns sharpen the picture: 58% of product professionals now use AI (up from 44% in 2024), but AI-generated themes frequently miss deeper context and underlying anxiety drivers that human researchers identify; 21% of practitioners cite speed-quality tension as their biggest challenge. New governance frameworks (April-May 2026) emphasize source verification, construct validity checks, and human-in-the-loop review as prerequisites for responsible deployment, positioning synthesis outputs as 'prediction not verification' rather than fact.
Reliability and governance represent the binding constraints on category expansion beyond enterprise segment. Industry assessment (Greenbook, June 2026) finds 95% of users report flaws in AI synthesis, with synthetic data losing momentum across stakeholder segments and 35% of research firms reporting staff displacement driven by task automation rather than job elimination. Hallucination mitigation shows technical promise: multi-model verification architecture reduces hallucination rates by 61% (from 8.3% to 3.2% across enterprise deployments), but this adds complexity and cost unsuitable for mid-market adoption. Practitioner research (June 2026) documents concrete failure modes: AI synthesis flagged 11 usability problems in one project but 10 were false positives or hallucinations—requiring manual quality gates that erode the speed advantage. Critical research (April 2026) documents that AI-based research methodologies fail at adoption scale: systems designed from users' stated preferences achieve only 57.7% accuracy, underperforming naive baselines, and deployment variance is extreme (bottom-quartile teams reach 12-18% daily active users vs. top-quartile 82-88% within 90 days). The differentiator is understanding the problem before building—a research design issue, not a technology one. Vendor-led implementations succeed at roughly twice the rate of internal builds, and a 42-day average project cycle suggests the bottleneck is process and research methodology, not processing power. Industry consensus has shifted from AI-as-replacement toward responsible AI augmentation: vendors explicitly position AI as effective for accelerating interpretation and synthesis automation while humans own meaning, impact, and decisions. This maturation signals the category has settled into a sustainable but bounded equilibrium: proven value for large enterprises with mature research operations and research discipline, persistent structural barriers preventing expansion to mid-market segments rooted in adoption methodology and organizational readiness rather than tooling capability, and synthesis accuracy constraints that training improvements alone cannot resolve.
— UserTesting GA: MCP Server integrations into Claude, ChatGPT, Figma; enhanced Figma plugin supporting multi-task studies—enabling research synthesis embedded in AI and design workflows where decisions happen.
— Critical assessment of AI misuse patterns: context compression, gap-filling, premature contradiction resolution, confidence inflation—documents real failure modes when AI is deployed carelessly without human verification and reflexivity requirements.
— Independent G2 market leadership: UserTesting (283 reviews, 96% 4–5 stars, 89% recommend); User Interviews (1000+ reviews, 4.6/5, 95% 4–5 stars, 87% recommend)—signals sustained adoption and satisfaction at scale.
— Peer-reviewed framework in International Journal of Social Research Methodology proposing methodologically congruent GenAI use: AI as assistant not analyst, researcher-led synthesis with human-owned interpretation, addressing methodological integrity concerns.
— Ecosystem convergence signal: category-wide move to AI-moderated interviews at scale, MCP integrations enabling Claude/ChatGPT/Cursor access to research data, and agentic execution enabling autonomous study management.
— Dovetail Sun's Out launch: 10 new feedback integrations (Qualtrics, Salesforce, Pendo, etc.) + 8 MCP connectors enabling synthesis without tool-switching, closing feedback collection and intelligence distribution gaps.
— Dovetail Channels 2.0 GA: automated feedback classification from 30+ sources, revenue-ranked prioritization, and one-click dispatch to Claude Code/Linear/Jira with customer context attached—demonstrates synthesis-to-action operational maturity.
— Practitioner task-by-task breakdown: transcription/translation/coding work reliably; sampling, rapport, contradiction-interpretation fail—documents capability boundaries and safeguard requirements (never accept code without traceable source sentence).