{
  "id": "research-analysis",
  "label": "Research & Knowledge",
  "description": "AI for finding, synthesising, verifying, and preserving organisational knowledge. Mostly leading-edge: literature review, competitive intelligence, and knowledge management tools are maturing quickly with five practices actively advancing. The main constraint is hallucination risk — fact-checking and source verification still require human oversight in high-stakes contexts.",
  "icon": "🔬",
  "filters": [
    "discovering"
  ],
  "hasSummary": true,
  "hasExecSummary": true,
  "practiceCount": 14,
  "evidenceCount": 2485,
  "practices": [
    {
      "slug": "academic-patent-and-technical-literature-analysis",
      "name": "Academic, patent & technical literature analysis",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that reviews academic papers, patents, and technical literature to identify prior art, synthesise findings, and map research landscapes. Includes citation network analysis and prior art search; distinct from general research retrieval which handles broader information needs.",
      "evidenceCount": 220
    },
    {
      "slug": "continuous-research-monitoring-and-alerting",
      "name": "Continuous research monitoring & alerting",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that continuously monitors sources for new information on defined topics and alerts users to significant developments. Includes automated literature watch and competitive signal monitoring; distinct from deep research which conducts one-off investigations rather than ongoing surveillance.",
      "evidenceCount": 115
    },
    {
      "slug": "document-summarisation-and-synthesis",
      "name": "Document summarisation & synthesis",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that summarises individual documents and synthesises information across multiple sources into coherent outputs. Includes executive summary generation and cross-document theme extraction; distinct from deep research which autonomously gathers sources rather than summarising provided ones.",
      "evidenceCount": 210
    },
    {
      "slug": "domain-specific-rag-and-cross-corpus-question-answering",
      "name": "Domain-specific RAG & cross-corpus question answering",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that performs retrieval-augmented generation over proprietary domain-specific corpora and answers questions across multiple knowledge bases. Includes specialised embedding and retrieval for technical domains; distinct from enterprise search which targets general internal documentation.",
      "evidenceCount": 194
    },
    {
      "slug": "due-diligence-research-automation",
      "name": "Due diligence research automation",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that automates research components of due diligence for M&A, investment, and partnership decisions. Includes automated company profiling and risk flag identification; distinct from financial auditing which examines internal records rather than external research.",
      "evidenceCount": 179
    },
    {
      "slug": "email-thread-summarisation-and-key-point-extraction",
      "name": "Email thread summarisation & key point extraction",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that summarises long email threads, extracts key decisions and open questions, and highlights action items. Includes thread digest generation and decision extraction; distinct from email triage which prioritises rather than summarises.",
      "evidenceCount": 185
    },
    {
      "slug": "enterprise-search-and-rag",
      "name": "Enterprise search & RAG",
      "tier": "good-practice",
      "trend": "steady",
      "blockerType": null,
      "description": "AI-powered search and retrieval-augmented generation across internal documentation and enterprise systems. Includes cross-system search federation and context-aware answer generation; distinct from domain-specific RAG which targets specialised corpora rather than general enterprise knowledge.",
      "evidenceCount": 174
    },
    {
      "slug": "knowledge-management-capture-taxonomy-and-curation",
      "name": "Knowledge management — capture, taxonomy & curation",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that captures institutional knowledge, generates taxonomies and ontologies, and maintains organisational knowledge structures. Includes automated knowledge graph construction and expert knowledge extraction; distinct from enterprise search which retrieves rather than organises knowledge.",
      "evidenceCount": 165
    },
    {
      "slug": "market-and-competitive-intelligence-gathering",
      "name": "Market & competitive intelligence gathering",
      "tier": "good-practice",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that systematically gathers and synthesises competitive intelligence from public sources, filings, and market data. Includes automated competitor monitoring and market signal aggregation; distinct from competitive product analysis in product & design which compares features rather than market positioning.",
      "evidenceCount": 187
    },
    {
      "slug": "meeting-intelligence-transcription-summaries-and-actions",
      "name": "Meeting intelligence — transcription, summaries & actions",
      "tier": "good-practice",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that transcribes meetings, generates summaries, extracts action items, and tracks follow-up completion. Includes speaker attribution and automated task assignment; distinct from email summarisation which processes written rather than spoken communication.",
      "evidenceCount": 189
    },
    {
      "slug": "multi-step-autonomous-deep-research",
      "name": "Multi-step autonomous deep research",
      "tier": "bleeding-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI agents that conduct multi-step research autonomously — formulating queries, reading sources, following leads, and synthesising findings. Includes tools like Gemini Deep Research and Perplexity Pro; distinct from single-query retrieval which answers from a single search round.",
      "evidenceCount": 164
    },
    {
      "slug": "single-query-research-retrieval-and-summary",
      "name": "Single-query research retrieval & summary",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that retrieves relevant information from a single search query and synthesises a coherent answer with source attribution. Includes search-augmented generation and cited responses; distinct from deep research which conducts multi-step autonomous investigation.",
      "evidenceCount": 173
    },
    {
      "slug": "trend-identification-and-horizon-scanning",
      "name": "Trend identification & horizon scanning",
      "tier": "leading-edge",
      "trend": "accelerating",
      "blockerType": null,
      "description": "AI that identifies emerging trends and weak signals across large volumes of publications, filings, and discussions. Includes early signal detection and trend trajectory modelling; distinct from social listening which monitors social platforms rather than scanning broad information sources.",
      "evidenceCount": 148
    },
    {
      "slug": "verification-fact-checking-citations-and-source-quality",
      "name": "Verification — fact-checking, citations & source quality",
      "tier": "leading-edge",
      "trend": "steady",
      "blockerType": null,
      "description": "AI that verifies citations, validates factual claims, assesses source quality, and ranks source reliability. Includes automated reference checking and misinformation detection; distinct from research retrieval which finds information rather than verifying it.",
      "evidenceCount": 182
    }
  ],
  "summary": "## Where AI Stands in Research & Knowledge\n\nResearch and knowledge work is the domain where AI capability has most clearly outrun AI trustworthiness, and the gap is now measurable in hours rather than anecdotes. Almost every practice here is technically mature and commercially available: Clarivate reports 85% of IP professionals using AI tools, up from 57% in 2023; Thomson Reuters' CoCounsel Legal serves over a million professionals across 107 countries; Gemini Notebook counts 600,000-plus organisations; AlphaSense sits at roughly $700M ARR with 85% of the S&P 100 as customers. What has not arrived is the productivity dividend those numbers imply. A study of 637 scientists found 89% spend more than a tenth of their AI-derived time savings verifying, debugging or fact-checking the output, and 46% spend more than a quarter of it. Questel's survey of IP professionals found 88% spend half their time reviewing AI output. Fish & Richardson cut prior-art search from 14 hours to 3.2, then logged a 34% error rate on the resulting claim charts. This is the domain's defining arithmetic: the saving is real, the verification cost is real, and in high-stakes work the second frequently cancels the first.\n\nThe structural reason is that a research output is a claim about the world that somebody else will act on, and nothing in the stack verifies it automatically. The evidence on that failure has hardened from alarming to epidemiological. A synthesis of four independent audits found 146,932 hallucinated citations inserted into the 2025 literature alone, with fabrication rates rising twelvefold in three years — from one in 2,828 papers in 2023 to one in 277 by early 2026 — and 98.4% of flagged papers receiving no publisher action. The University of Auckland found AI-hallucinated references in 2% of more than 3,000 peer-reviewed papers. Across NeurIPS, ICLR, ICML and ACL, between 18.7% and 26.2% of *accepted* papers contain at least one hallucinated reference, and the rating gap between papers with and without them is statistically indistinguishable from noise: peer review, the discipline's designated verification layer, is not catching this. Nor are the tools. CASRAI's evaluation of five commercial citation-verification products found all of them unreliable for unsupervised deployment, each trading false positives against missed hallucinations. The University of Michigan, across 670,000 trials, found language models weight source *popularity* over source *reliability* by a factor of two.\n\nWhere momentum genuinely persists, a regulator is usually supplying it. Continuous pharmacovigilance monitoring is the domain's one advancing practice, and not because the technology is better — it is because EU Implementing Regulation 2025/1466 obliges every marketing authorisation holder to integrate EudraVigilance signal detection into a continuous management system, AMLA's guidelines extend equivalent duties to financial crime monitoring across 27 member states from July 2027, and the FDA now pipes near-real-time trial safety signals to sponsors including Amgen and AstraZeneca. Veeva's Falcon Safety and Syneos Health's Hyderabad hub are production deployments built against a mandate. Contrast the unmandated case: Cloud Security Alliance research found 94.1% of AI-flagged security alerts were false positives across 16.9 million records, and Gartner expects only 15% of SOC AI pilots to show measurable results despite 70% adoption. Absent an external obligation, alert-tuning discipline collapses and the signal drowns. The same bifurcation runs through every other practice here: narrow, supervised retrieval and triage pays reliably — IPRally, Patsnap and Cypris all beat general frontier models on prior-art recall, and domain-adapted embeddings lift retrieval precision 20–40% over off-the-shelf models — while autonomous synthesis remains bounded by hallucination burden and a validation overhead that approximately equals generation time.\n\n## What's New, 2026-09-09 to 2026-09-23\n\nTwo practices moved from advancing to stalled this cycle: academic and patent literature analysis, and verification itself. Both moved for the same reason — the enforcement layer arrived and immediately exposed how much unverified output is in circulation. EMNLP's new citation policy desk-rejected 258 submissions, sanctioned 1,166 authors and banned 35 repeat offenders from EMNLP 2027. The USPTO issued its first AI-predicated attorney discipline order in *In re Brian E. Mitchell*, publicly reprimanding hallucinated citations to the intrinsic patent record. LexisNexis shipped Protégé to general availability claiming 70–90% time reduction at Schott Pharma, while an independent vendor assessment put legal-reasoning error rates at 69–88% despite 90–94% search F1 — the clearest statement yet that retrieval and reasoning are separate capabilities with separate failure profiles. Hallucinated citations in SIGCSE computing-education proceedings rose from three to seventeen year on year. Meanwhile Meta replaced professional fact-checking with community notes across sixteen Spanish-speaking countries, removing verification capacity precisely where AI-generated misinformation is growing.\n\nThe measurement story sharpened everywhere else. A production audit of Perplexity found 34.7% of 1,826 citations across 310 factual questions failed verification, with a quarter of cited pages never captured by the Internet Archive — and Perplexity's chatbot traffic share reportedly collapsed to 1.3%. Salesforce disclosed that a RAG system scoring above 90% on benchmarks fell to 46% on real enterprise documents, recovering to 84.4% only after intelligent parsing and reranking; IDC now puts the share of AI proofs-of-concept that never reach wide deployment at 88%. Kolena's benchmark of 45 models across 264 runs found field-level extraction accuracy unstable week to week on identical schemas, and separate analysis showed 97% per-field accuracy translating to roughly 38% of summaries being fully correct — a metric-inflation problem that explains why procurement decisions keep failing to predict deployment outcomes. In knowledge management the infrastructure improved markedly — AWS shipped its Context Ontology Accelerator to general availability under Apache 2.0, Google rebranded Dataplex to a Gemini-powered Knowledge Catalog, Databricks released Genie Ontology — against failure-mode analysis finding fewer than 15% of enterprise knowledge-graph pilots pass pilot stage, with 67% of abandonments blamed on expertise gaps and programme costs running into eight figures. On deep research, Google's Deep Research agents reached general availability through the Gemini API and OpenResearcher's distilled Nemotron-3-Nano matched GPT-4.1 on BrowseComp-Plus, while an independent twenty-brief benchmark found no tool clearing 50% requirement coverage and IBM measured a 24.4-point consistency gap across repeated runs of the same agent. Meeting intelligence gained a tier-one platform entrant in Webex and a bot-free challenger in Notion, even as a federal court in California held that an AI notetaker can function as a third-party eavesdropper rather than a tool of the consenting host — a ruling that dismantles the single-host-consent defence enterprises had relied on, with BIPA statutory damages of $1,000 to $5,000 per violation behind it.\n\n## Key Tensions\n\n- **The verification tax cancels the productivity gain.** Across surveyed scientists, 89% spend more than 10% of their AI time savings on verification and 46% spend more than a quarter; Questel found 88% of IP professionals spend half their time reviewing AI output. Where verification is skipped the cost reappears downstream — Fish & Richardson's 14-to-3.2-hour prior-art gain came with a 34% claim-chart error rate. Until validation is cheaper than generation, headline time savings will keep failing to reach the P&L.\n\n- **Benchmarks systematically overstate production behaviour.** Salesforce's RAG system scored above 90% on benchmarks and 46% on real enterprise documents; offline RAGAS scores of 0.92 on gold datasets correspond to roughly 0.78 on live traffic. Kolena found field-level accuracy unstable week to week across 45 models on identical schemas, and 97% per-field accuracy yields only about 38% fully correct summaries. Procurement decisions made on vendor benchmarks are being made on the wrong number.\n\n- **Courts and conferences are now doing the governing.** EMNLP desk-rejected 258 submissions and banned 35 authors from its 2027 edition; the USPTO issued its first AI-predicated attorney discipline order; Florida's AOSC26-12 requires certification of citation accuracy in every filing; a Munich court held Google liable for AI Overview statements, rejecting the platform-neutrality defence. Verification has shifted from a quality preference to a liability exposure, which changes who signs off and how slowly.\n\n- **Mandate, not enthusiasm, is the only reliable adoption engine.** The domain's one genuinely advancing practice is continuous pharmacovigilance monitoring, compelled by EU Implementing Regulation 2025/1466 and extended to financial crime by AMLA from July 2027. Where no mandate exists, discipline decays: 94.1% of AI-flagged security alerts were false positives across 16.9 million records, and Gartner expects just 15% of SOC AI pilots to show measurable results. Voluntary governance has not scaled; obligation has.\n\n- **The ceiling sits below the model, in retrieval and curation.** Anthropic's own remediation data shows keyword search plus reranking cuts RAG failures by 67%, with reranking alone the single largest contributor — ahead of swapping embedding models. Domain-adapted embeddings lift retrieval precision 20–40% over generic ones, and a Michigan study across 670,000 trials found models weight source popularity over reliability by two to one. Organisations upgrading models while leaving corpora, chunking and permissions untouched are optimising the wrong layer.\n\n## Top 10 Evidence Items\n\n1. **Enterprise RAG Pipeline Failure and Recovery: Benchmark-to-Production Accuracy Gap at Salesforce** (case-study) — The Salesforce 90%-to-46% collapse is the single clearest demonstration that benchmarks overstate production behaviour, cited three times across the briefing. https://engineering.salesforce.com/enterprise-ai-accuracy-building-a-more-trustworthy-rag-application/\n2. **How to use AI for patent review and analysis in 2026?** (tutorial) — Fish & Richardson's 14-to-3.2-hour gain paired with a 34% claim-chart error rate is the defining arithmetic of the verification tax made concrete. https://patentreviewpro.com/knowledge/how_to_use_ai_for_patent_review_and_analysis_in_2026.php\n3. **Audit of 111M references finds 146,932 fake 2025 citations | AI Weekly** (news-coverage) — The 146,932-hallucinated-citation audit is the epidemiological evidence that fabrication has moved from anecdote to measurable scale. https://aiweekly.co/alerts/audit-of-111m-references-finds-146932-fake-2025-citations\n4. **How Accurate Are AI-Based Patent Analysis Tools? | iLumos** (opinion) — Shows retrieval and reasoning as separate capabilities with separate failure profiles - 90-94% search F1 against 69-88% legal-reasoning error. https://ilumos.ai/blog/ai-patent-analysis-tools-accuracy\n5. **Cloud Security Alliance: AI-Generated SOC Alerts Flood to 94.1% Noise Rate** (research-paper) — 94.1% false-positive rate is the sharpest evidence that unmandated continuous monitoring collapses under its own noise. https://labs.cloudsecurityalliance.org/research/csa-research-note-ai-soc-alert-noise-20260914-csa-styled/\n6. **Polpharma/Syntro PSUR Automation Pilot: 100% Completeness, 98% Accuracy** (case-study) — The named pharmacovigilance pilot hitting 98% accuracy shows what mandate-driven adoption actually achieves against the SOC counterexample. https://startsmartcee.org/case-studies/how-do-you-scale-ai-inside-a-pharma-giant-testing-ai-in-a-pharmaceutical-environment-the-syntro-and-polpharma-experience/\n7. **Enterprise knowledge graph failure analysis (15% pilot success, 67% expertise gap)** (opinion) — Sub-15% pilot success with 67% blamed on expertise gaps undercuts the wave of knowledge-graph product launches in the same cycle. https://atlan.com/know/ai-agent/knowledge-graph/enterprise-knowledge-graph-pitfalls/\n8. **Otter.ai August 13, 2026 Ruling: Third-Party Eavesdropper Precedent Upholds Class Claims** (opinion) — The Otter.ai eavesdropper ruling is the clearest instance of courts now doing the governing rather than vendors or standards bodies. https://basilai.app/articles/2026-09-16-after-otter-ruling-ai-notetaker-procurement-checklist-enterprise-buyers.html\n9. **Computing education papers show hallucinated citations rising sharply (SIGCSE 2026)** (research-paper) — SIGCSE's citation count rising from three to seventeen year-on-year is the concrete instance behind the claim that verification itself stalled this cycle. https://arxiv.org/abs/2609.16574\n10. **AI Deep Research: Codex vs Claude vs Grok vs Exa—Independent DR-20 Benchmark** (opinion) — No deep-research tool clearing 50% requirement coverage punctures the GA announcements for Gemini Deep Research and OpenResearcher in the same period. https://aimultiple.com/ai-deep-research",
  "execSummary": "**The headline:** AI research tools now save real hours. Checking what they produce eats most of those hours back — and courts, regulators and journals have started punishing the organizations that skip the checking.\n\n### The Picture\n\nAlmost every tool in this field is already bought and installed. Eighty-five percent of intellectual property professionals now use AI tools, up from 57% three years ago; Thomson Reuters' legal assistant serves over a million professionals across 107 countries; Google's research notebook counts more than 600,000 organizations. What almost nobody has is the productivity dividend those numbers imply, because verifying the output consumes a large share of the time saved. The organizations genuinely pulling ahead are, with few exceptions, the ones a regulator obliged to build governance first — continuous drug-safety monitoring in pharma is the one practice here that is clearly advancing. Everyone else is in the pack, and the window to build verification discipline voluntarily is closing as courts, regulators and publishers start imposing it.\n\n### This Fortnight\n\n- **A major AI research conference rejected 258 papers over fabricated citations and banned 35 authors, while the US patent office disciplined its first attorney for the same offense.** Both landed in the same fortnight, and both target the individual signer rather than the software. Anything your organization files, publishes or submits now carries a personal accountability trail, so someone senior has to own the citation check before it goes out.\n\n- **A federal court in California held that an AI meeting notetaker can count as a third-party eavesdropper rather than a tool of the host who consented to recording.** That undercuts the single-host-consent policy most companies rely on, and Illinois biometric law attaches damages of $1,000 to $5,000 per violation, multiplied across every participant in every recorded meeting. If notetakers have spread through your teams without a consent policy, that is now a quantifiable liability rather than an IT question.\n\n- **Salesforce disclosed that a system feeding its AI internal documents to work from scored above 90% in testing and 46% on real company files.** It recovered to 84% only after rebuilding how documents were parsed and results re-ranked — not by changing the AI model. IDC now puts the share of AI pilots that never reach wide deployment at 88%, and this gap is the main reason: vendor scores are measured on clean data your filing system does not resemble.\n\n- **LexisNexis moved its legal AI assistant out of beta claiming 70 to 90% time savings, while a separate assessment of legal AI tools put reasoning error rates at 69 to 88%.** The same assessment found document-search accuracy of 90 to 94%. Finding the right authority and reasoning correctly about it are separate capabilities with separate failure rates, so buy these tools for retrieval and keep qualified people on the judgment.\n\n- **A study of 637 scientists found 89% spend more than a tenth of their AI time savings checking the output, and 46% spend more than a quarter of it.** A parallel survey of intellectual property professionals put it higher: 88% spend half their working time reviewing what the AI produced. The savings in your business case are real, but a large and measurable share of them never reaches the bottom line.\n\n### Coming Up\n\n- **EU anti-money-laundering rules will require continuous, AI-assisted transaction monitoring across all 27 member states from July 2027.** Pharmaceutical firms already live under an equivalent obligation for drug-safety signals, and it is the one corner of this field where adoption reliably works. Where no mandate exists, discipline decays — one large study found 94% of AI-flagged security alerts were false positives. If a regulator is about to compel you, the build starts in this budget cycle.\n\n- **Certification duties are spreading: Florida now requires court filers to certify every cited authority exists, and a Munich court held Google liable for statements in its AI-generated summaries.** Judges and regulators are converging on one principle — whoever publishes an AI-generated claim owns it, and pointing at the model is not a defense. Put a named sign-off step on anything that leaves the building, before a court invents one for you.\n\n- **The open web that AI research tools can read is shrinking, with Cloudflare's default blocking potentially cutting up to 30% of top sites out of AI-generated answers.** Publishers with paid licensing deals already earn a 48% citation premium in ChatGPT answers, and roughly one in six sources these tools retrieve is itself AI-generated. Assume your teams are researching from a narrower and more commercially shaped slice of the web than six months ago, and check anything consequential against a primary source.\n\n### What's Hard About This\n\n- **In high-stakes work, the checking costs roughly what the generating saved.** One law firm cut prior-art searching from 14 hours to 3.2 using AI, then logged a 34% error rate on the resulting analysis. Until validating an output is cheaper than producing it, headline time savings will keep failing to show up in the accounts.\n\n- **Vendor benchmarks systematically overstate how these tools behave on your own data.** A benchmark of 45 models found accuracy swinging week to week on identical document templates, and 97% accuracy per data field yields only about 38% of summaries fully correct end to end. Insist on a paid trial against your real documents before signing anything.\n\n- **The limiting factor sits below the AI model, in your documents, indexes and permissions.** Anthropic's own remediation data shows that adding keyword search and re-ranking results cut retrieval failures by 67%, a bigger gain than swapping the underlying model. Fewer than 15% of corporate knowledge-graph projects get past pilot, with two thirds of abandonments blamed on missing in-house expertise — budget for librarians and data engineers, not just licenses.",
  "headline": "AI research tools now save real hours. Checking what they produce eats most of those hours back — and courts, regulators and journals have started punishing the organizations that skip the checking.",
  "execSummarySections": [
    {
      "id": "the-picture",
      "title": "The Picture",
      "body": "Almost every tool in this field is already bought and installed. Eighty-five percent of intellectual property professionals now use AI tools, up from 57% three years ago; Thomson Reuters' legal assistant serves over a million professionals across 107 countries; Google's research notebook counts more than 600,000 organizations. What almost nobody has is the productivity dividend those numbers imply, because verifying the output consumes a large share of the time saved. The organizations genuinely pulling ahead are, with few exceptions, the ones a regulator obliged to build governance first — continuous drug-safety monitoring in pharma is the one practice here that is clearly advancing. Everyone else is in the pack, and the window to build verification discipline voluntarily is closing as courts, regulators and publishers start imposing it."
    },
    {
      "id": "this-fortnight",
      "title": "This Fortnight",
      "body": "- **A major AI research conference rejected 258 papers over fabricated citations and banned 35 authors, while the US patent office disciplined its first attorney for the same offense.** Both landed in the same fortnight, and both target the individual signer rather than the software. Anything your organization files, publishes or submits now carries a personal accountability trail, so someone senior has to own the citation check before it goes out.\n\n- **A federal court in California held that an AI meeting notetaker can count as a third-party eavesdropper rather than a tool of the host who consented to recording.** That undercuts the single-host-consent policy most companies rely on, and Illinois biometric law attaches damages of $1,000 to $5,000 per violation, multiplied across every participant in every recorded meeting. If notetakers have spread through your teams without a consent policy, that is now a quantifiable liability rather than an IT question.\n\n- **Salesforce disclosed that a system feeding its AI internal documents to work from scored above 90% in testing and 46% on real company files.** It recovered to 84% only after rebuilding how documents were parsed and results re-ranked — not by changing the AI model. IDC now puts the share of AI pilots that never reach wide deployment at 88%, and this gap is the main reason: vendor scores are measured on clean data your filing system does not resemble.\n\n- **LexisNexis moved its legal AI assistant out of beta claiming 70 to 90% time savings, while a separate assessment of legal AI tools put reasoning error rates at 69 to 88%.** The same assessment found document-search accuracy of 90 to 94%. Finding the right authority and reasoning correctly about it are separate capabilities with separate failure rates, so buy these tools for retrieval and keep qualified people on the judgment.\n\n- **A study of 637 scientists found 89% spend more than a tenth of their AI time savings checking the output, and 46% spend more than a quarter of it.** A parallel survey of intellectual property professionals put it higher: 88% spend half their working time reviewing what the AI produced. The savings in your business case are real, but a large and measurable share of them never reaches the bottom line."
    },
    {
      "id": "coming-up",
      "title": "Coming Up",
      "body": "- **EU anti-money-laundering rules will require continuous, AI-assisted transaction monitoring across all 27 member states from July 2027.** Pharmaceutical firms already live under an equivalent obligation for drug-safety signals, and it is the one corner of this field where adoption reliably works. Where no mandate exists, discipline decays — one large study found 94% of AI-flagged security alerts were false positives. If a regulator is about to compel you, the build starts in this budget cycle.\n\n- **Certification duties are spreading: Florida now requires court filers to certify every cited authority exists, and a Munich court held Google liable for statements in its AI-generated summaries.** Judges and regulators are converging on one principle — whoever publishes an AI-generated claim owns it, and pointing at the model is not a defense. Put a named sign-off step on anything that leaves the building, before a court invents one for you.\n\n- **The open web that AI research tools can read is shrinking, with Cloudflare's default blocking potentially cutting up to 30% of top sites out of AI-generated answers.** Publishers with paid licensing deals already earn a 48% citation premium in ChatGPT answers, and roughly one in six sources these tools retrieve is itself AI-generated. Assume your teams are researching from a narrower and more commercially shaped slice of the web than six months ago, and check anything consequential against a primary source."
    },
    {
      "id": "whats-hard-about-this",
      "title": "What's Hard About This",
      "body": "- **In high-stakes work, the checking costs roughly what the generating saved.** One law firm cut prior-art searching from 14 hours to 3.2 using AI, then logged a 34% error rate on the resulting analysis. Until validating an output is cheaper than producing it, headline time savings will keep failing to show up in the accounts.\n\n- **Vendor benchmarks systematically overstate how these tools behave on your own data.** A benchmark of 45 models found accuracy swinging week to week on identical document templates, and 97% accuracy per data field yields only about 38% of summaries fully correct end to end. Insist on a paid trial against your real documents before signing anything.\n\n- **The limiting factor sits below the AI model, in your documents, indexes and permissions.** Anthropic's own remediation data shows that adding keyword search and re-ranking results cut retrieval failures by 67%, a bigger gain than swapping the underlying model. Fewer than 15% of corporate knowledge-graph projects get past pilot, with two thirds of abandonments blamed on missing in-house expertise — budget for librarians and data engineers, not just licenses."
    }
  ],
  "url": "https://www.thestateofplay.ai/domain/research-analysis",
  "license": "CC BY 4.0",
  "licenseUrl": "https://creativecommons.org/licenses/by/4.0/",
  "generatedAt": "2026-10-01"
}