Perly Consulting │ Beck Eco

The State of Play

A living index of AI adoption across industries — where established practice meets the bleeding edge
UPDATED DAILY

The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.

The Daily Dispatch

A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.

Pick a role above to explore practices

BLEEDING EDGE

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🔬 RESEARCH & KNOWLEDGE
⚖️ LEGAL, COMPLIANCE & RISK
🎧 CUSTOMER OPERATIONS
🏛️ AI GOVERNANCE & SAFETY
📊 DATA & ANALYTICS
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💼 SALES & REVENUE
🎬 CREATIVE & GENERATIVE MEDIA
👁️ COMPUTER VISION & SENSING
💹 FINANCE & ACCOUNTING
🔄 OPERATIONS & PROCESS AUTOMATION
🚗 AUTONOMOUS SYSTEMS & VEHICLES
🦾 PHYSICAL AI & ROBOTICS
🎓 EDUCATION & LEARNING
PERSONAL EFFECTIVENESS

LEADING EDGE

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🔬 RESEARCH & KNOWLEDGE
⚖️ LEGAL, COMPLIANCE & RISK
🎧 CUSTOMER OPERATIONS
🏛️ AI GOVERNANCE & SAFETY
📊 DATA & ANALYTICS
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💼 SALES & REVENUE
🎬 CREATIVE & GENERATIVE MEDIA
👁️ COMPUTER VISION & SENSING
💹 FINANCE & ACCOUNTING
🔄 OPERATIONS & PROCESS AUTOMATION
👥 PEOPLE & TALENT
🚗 AUTONOMOUS SYSTEMS & VEHICLES
🦾 PHYSICAL AI & ROBOTICS
🎓 EDUCATION & LEARNING
PERSONAL EFFECTIVENESS

GOOD PRACTICE

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🔬 RESEARCH & KNOWLEDGE
⚖️ LEGAL, COMPLIANCE & RISK
🎧 CUSTOMER OPERATIONS
🏛️ AI GOVERNANCE & SAFETY
📊 DATA & ANALYTICS
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💼 SALES & REVENUE
🎬 CREATIVE & GENERATIVE MEDIA
👁️ COMPUTER VISION & SENSING
💹 FINANCE & ACCOUNTING
🔄 OPERATIONS & PROCESS AUTOMATION
👥 PEOPLE & TALENT
🚗 AUTONOMOUS SYSTEMS & VEHICLES
🦾 PHYSICAL AI & ROBOTICS
🎓 EDUCATION & LEARNING
PERSONAL EFFECTIVENESS

ESTABLISHED

⌨️ SOFTWARE ENGINEERING
✍️ CONTENT & MARKETING
🛡️ IT OPERATIONS & SECURITY
🎯 PRODUCT & DESIGN
💹 FINANCE & ACCOUNTING
👥 PEOPLE & TALENT

🔬 Research & Knowledge

AI for finding, synthesising, verifying, and preserving organisational knowledge. Mostly leading-edge: literature review, competitive intelligence, and knowledge management tools are maturing quickly with five practices actively advancing. The main constraint is hallucination risk — fact-checking and source verification still require human oversight in high-stakes contexts.

14 practices: 3 good practice, 10 leading edge, 1 bleeding edge

Where AI Stands in Research & Knowledge

This is the domain where AI's raw capability is least in dispute and its trustworthiness is most in dispute. Every link in the research chain — find, retrieve, read, summarise, synthesise, verify, remember — now has commodity tooling shipped by a hyperscaler and deployed at population scale. Google's Gemini Notebook counts 13 million individual users and 600,000 organisations. Ninety-four percent of B2B buyers report using AI for vendor research, and 69% changed their vendor choice after an AI interaction. Enterprise RAG is the dominant architecture at 73% of 847 surveyed Fortune 500 and FTSE 350 companies. Yet of the fourteen practices tracked here, only three have settled into genuine good practice: enterprise search and retrieval, meeting transcription, and competitive intelligence gathering. Everything upstream of "fetch a document and quote it accurately" remains a supervised activity, and the reason is arithmetic rather than pessimism. Research is a chain, and chains multiply. At 95% per-step accuracy, a twenty-step research task succeeds 36% of the time. Enterprise RAG postmortems put 73% of failures at the retrieval layer, not generation. A 124-study systematic review of AI literature synthesis found 80.6–96.5% sensitivity on structured screening tasks collapsing to 4.6% precision on interpretative synthesis.

The dividing line that organises the whole domain is between grounded and open-ended work. Where the source material is bounded, curated, and in front of the model, the results are real and large. Morgan Stanley synthesises across 100,000 proprietary research documents for 16,000 advisers. JPMorgan cut M&A document review from multi-hour cycles to under thirty seconds. A $1.2 billion private equity acquisition processed 14,000 data-room documents in 72 hours and surfaced three undisclosed change-of-control consents that produced a $40 million price reduction. Gong passed $500 million ARR with named outcomes at BruntWork (8,000 agents, double the revenue per agent) and Yotpo (account research from thirty minutes to thirty seconds). AlphaSense is estimated at $700 million ARR and a $7.5 billion valuation ahead of a possible listing. Where the model must go and find things for itself, the picture inverts. Columbia's Tow Center measured Perplexity at a 37% error rate across 1,600 queries — and that was the best result of the eight tools tested. No large language model, under any prompting regime tested across four models, exceeded a 50% citation-existence rate. Specialised legal tools do not escape it either: Lexis+ at 65% accuracy with 17% hallucination, Westlaw at 42% and 33%.

What has changed in 2026 is that the money has begun to notice. KPMG's survey of 2,145 leaders found 49% had narrowed, delayed, or paused agentic AI rollouts on cost-value mismatch, with only 7% reporting established ROI against 76% claiming meaningful business value. Gartner forecasts that more than 40% of agentic projects will be cancelled by year-end. A Forbes survey found 74% enterprise deployment against 50% unable to measure return at all. Meanwhile a Y Combinator-backed startup founded specifically to automate private-equity due diligence pivoted wholesale into construction compliance within six months of launch — a quiet but pointed signal about where the durable economics sit. The domain is entering a phase in which the thing being purchased is no longer capability but verification: hallucination-detection software is growing at 25.8% CAGR toward $981 million by 2031, and corporate legal AI adoption jumped from 44% to 87% year on year largely on the back of compliance controls rather than productivity features.

What's New, 2026-07-29 to 2026-08-12

No practice changed tier or trend this fortnight, which is itself the story: this is a domain consolidating, not moving. What did move was the evidence base, and it moved in three coordinated directions. First, the citation-integrity numbers converged from three independent sources into something that now reads as a structural condition rather than a scandal. A Lancet audit of 2.5 million PubMed Central papers put fabricated references at 1 in 277 papers in early 2026, up twelvefold from 1 in 2,828 in 2023, with 98% of flagged papers receiving no publisher action. CiteMe's first-party audit of 47,098 references found a 27.2% fabrication rate with 84.4% of bibliographies containing at least one fake citation. ACL 2026 desk-rejected more than a hundred already-accepted papers during camera-ready review over hallucinated references. And the detection layer failed its own audit: CASRAI tested five leading citation-checking tools and found all of them unreliable for unsupervised deployment, each trading false positives against missed fabrications. A University of Michigan study across 670,000 trials and thirteen models found LLMs weight source popularity over source reliability by a factor of two and perform near chance on source discernment — a foundational problem for any architecture that asks a model to police itself.

Second, and more usefully, the fortnight produced the first architectural intervention with clean numbers behind it. Google Research's Science One framework retrieves citations at generation time rather than generating text and attaching references afterwards; tested on 337 citations it produced zero phantom references against a 21% baseline. EviGraph, which maintains an explicit evidence graph across research stages and repairs inconsistencies, improved claim-support rate by 40%. A peer-reviewed benchmark on a 300-document TikTok litigation corpus put source-grounded Gemini Notebook at 13% hallucination against 40% for ChatGPT and Gemini on identical material. NewsGuard shipped a production news service generating exclusively from 12,000 pre-vetted outlets with mandatory citations, against an independent finding that unconstrained models spread false claims 35% of the time. Legal technology has already drawn the conclusion and abandoned AI-generated citations entirely in favour of retrieval-only systems with separate mandatory verification. Third, the retrieval orthodoxy inverted. A controlled scaling study across 28 corpus tiers from 1.7 million to 601 million tokens found plain BM25 lexical search overtaking agentic-first retrieval at around 10 million tokens and holding a roughly twenty-point accuracy margin at full scale — a finding echoed by a typed knowledge graph over 690 skills underperforming a plain hybrid lexical-plus-dense ranker by 11.2 points, and by practitioner benchmarking showing semantic reranking can actively regress accuracy on structured tabular data.

Key Tensions

  • Fabricated citations are compounding faster than the tools built to catch them. The Lancet's twelvefold rise to 1 in 277 papers, CiteMe's 27.2% fabrication rate across 47,098 references, and Michigan's finding that models rank sources by popularity at twice the rate they rank by reliability all point the same way. CASRAI's verdict that all five leading detectors are unsuitable for unsupervised use closes the obvious escape route: you cannot solve a model-generated verification problem with another model. Institutional enforcement is arriving faster than technical remedy — ACL's hundred-plus desk rejections, Florida's statewide filing-certification rule, and a global tally of documented court hallucination cases now well past 1,700 across more than forty countries.

  • Grounding at generation time is the only intervention with unambiguous evidence behind it. Every credible improvement this cycle came from constraining what the model may cite rather than from making the model larger. Science One's zero-from-337 result, Gemini Notebook's 13%-versus-40% margin on an identical corpus, NewsGuard's pre-vetted 12,000-source pool, and EviGraph's 40% claim-support gain share one design principle: the citation is retrieved as the text is written, not attached afterwards. Eighteen months of frontier capability gains produced no measurable reliability improvement for production research agents; architecture did.

  • More sophisticated retrieval is losing to simpler retrieval at enterprise scale. The BM25 scaling result is the sharpest version, but it is not isolated: a typed knowledge graph lost by 11.2 points to a plain hybrid ranker with 98.6% of its edges connecting items already surfaced; generic cross-encoder rerankers trained on web data degrade precision on scientific corpora; semantic reranking regressed tabular accuracy from 100% to 97%. Where agentic retrieval did improve, the lever was procedural discipline rather than capability — a "read-gate" invariant forcing agents to actually read retrieved evidence before answering added 14.9 to 19.9 accuracy points across 12,000 multi-hop trajectories. Organisations mid-way through a vector-database migration should read these results before the next architecture review.

  • The boundary between document work and judgment work is now drawn with numbers, not intuition. In due diligence, AI reliably compresses document-heavy work — contract review, data-room summarisation, expert-call preparation — at 31% genuine enterprise integration, while deal sourcing shows a 64% effectiveness gap and portfolio monitoring a 75% failure rate. The failure mode is precision, not capability: across Kira Systems, Luminance, and Harvey AI, 20–35% of flagged items are stale records — dissolved subsidiaries, expired UCC filings, terminated consents — adding over $36,000 in unbudgeted attorney time per deal because platforms optimise for recall and lack temporal logic. Wakefield found 88% of CFOs using agentic AI, 86% reporting hallucinations, and 14% trusting it.

  • Availability has decoupled from use, and use has decoupled from value. Seventy-two percent of knowledge workers have access to meeting AI through bundled platforms; 41% use it monthly. Only 40% of Copilot licensees maintain active use. Fifty-eight percent of meeting-intelligence deployments stall at "recorded but unused" by month nine without active manager coaching. Even where output is produced, one analysis found AI meeting summaries omitting 97% of significant decisions and actions — an omission failure invisible to hallucination metrics — while a 124-study review identified condition-dropping, hedge-flattening, and limitation-omission as structural effects of compression rather than isolated errors. The organisational work of making summaries load-bearing is consistently larger than the technical work of producing them.

Top 10 Evidence Items

  1. 1 in 277 PubMed Papers Now Cite Fake References (adoption-metric) — The single number that anchors the domain's central tension: a Lancet audit of 2.5 million papers puts fabricated citations twelvefold above 2023 levels, with 98% of flagged papers receiving no publisher action. https://casrai.org/news/fabricated-citations-1-in-277-pubmed-lancet-audit

  2. Hallucinated Citation Checkers: What They Miss (research-paper) — CASRAI's test of five leading detection tools found all of them unreliable for unsupervised use, closing off the obvious "use another model to police the model" fix and forcing the domain toward architectural rather than detective solutions. https://casrai.org/news/hallucinated-citation-detection-tools-what-they-can-and-cannot-do-2026

  3. Closing the Loop: How Scientists Are Teaching AIs to Keep Their Promises (case-study) — Case study of Google Research's Science One framework, which retrieves citations at generation time and produced zero phantom references against a 21% baseline — the fortnight's clearest evidence that grounding-by-architecture beats grounding-by-afterthought. https://akmaier.substack.com/p/closing-the-loop-how-scientists-are

  4. NotebookLM Review 2026: Honest Verdict After the Rename (opinion) — Independent review grounded in the peer-reviewed Hagar et al. benchmark showing source-grounded Gemini Notebook at 13% hallucination against 40% for ChatGPT and Gemini on an identical 300-document corpus — the same architectural point as Science One, replicated on a different product. https://gemini-notebook-hub.online/guides/notebooklm-review

  5. BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms (research-paper) — Controlled 28-tier study spanning 1.7 million to 601 million tokens finds plain lexical search overtaking agentic-first retrieval at around 10 million tokens and holding a 20-point margin at full scale, inverting a year of enterprise-RAG orthodoxy. https://arxiv.org/abs/2607.26497

  6. A typed knowledge graph over 690 skills performed 11.2 points worse than a plain hybrid ranker (research-paper) — Corroborates the BM25 finding from a different angle: 98.6% of the graph's edges connected items already surfaced by simpler retrieval, undercutting the case for knowledge-graph complexity as a default architecture. https://mindpattern.ai/s/2026-08-08-a-typed-knowledge-graph-over-690-skills-performed-11-2-points-worse-than-a-plain-hybr

  7. Dili's Pivot Away From PE Due Diligence Automation: Market Viability Signal (significant-repo) — An uncomfortable market signal: a YC-backed startup founded specifically to automate PE/VC due diligence via LLM extraction abandoned the space for construction compliance within six months, quietly confirming where the durable economics of automated judgment actually sit. https://www.teardown.ai/companies/dili

  8. The Legal AI "Dead Record" Problem (opinion) — Documents the precision failure underlying due diligence's document-versus-judgment boundary: 20–35% of items flagged by Kira Systems, Luminance, and Harvey AI are stale records — dissolved subsidiaries, expired UCC filings — adding over $36,000 in unbudgeted attorney time per deal. https://www.thelegalstack.org/posts/the-legal-ai-dead-record-problem-why-ai-due-diligence-tools

  9. Your AI Meeting Summary Is Probably Missing Something (opinion) — Names the omission failure mode directly: AI meeting summaries dropping 97% of significant decisions and actions, an error invisible to any hallucination metric and a sharp illustration of why availability has decoupled from value across meeting intelligence. https://hackernoon.com/your-ai-meeting-summary-is-probably-missing-something

  10. AlphaSense $700M ARR Analysis: IPO Trajectory and Market Dominance (adoption-metric) — The market's verdict on bounded, grounded research work: AlphaSense at an estimated $700 million ARR and $7.5 billion valuation ahead of a possible listing, the clearest capital-markets confirmation that curated-corpus retrieval is where the money is actually landing. https://sacra.com/research/700m-yr-sacra-for-public-markets/