Perly Consulting │ Beck Eco

The State of Play

A living index of AI adoption across industries — where established practice meets the bleeding edge
UPDATED DAILY

The AI landscape doesn't move in one direction — it lurches. Some techniques leap from experiment to table stakes in a single quarter; others stall against regulatory walls, technical ceilings, or organisational inertia that no amount of hype can dislodge. Knowing which is which is the hard part. The State of Play cuts through the noise with a rigorously maintained index of AI techniques across every major business domain — classified by maturity, evidenced by real-world adoption, and updated daily so you always know where you stand relative to the field. Stop guessing. Start knowing.

The Daily Dispatch

A daily newsletter distilling the past two weeks of movement in a domain or two — delivered to your inbox while the index updates in the background.

AI Maturity by Domain

Each dot marks the weighted maturity of practices within a domain — hover for a brief summary, click for more detail

DOMAIN
BLEEDING EDGEESTABLISHED

Spreadsheet & data task automation

BLEEDING EDGE

TRAJECTORY

Stalled

AI that automates spreadsheet tasks including formula creation, data analysis, pivot tables, and formatting from natural language. Includes Excel/Sheets AI assistants and formula generation; distinct from natural language to SQL which queries databases rather than manipulating spreadsheets.

OVERVIEW

AI-driven spreadsheet automation entered late 2026 facing a critical bifurcation: elite scaled deployment in finance and bounded workflows (52% accounting adoption, 250% ROI, case studies showing 75% cycle time reduction and +9.4% revenue per rep) versus a widespread adoption/usage paradox that reveals the true mainstream constraint is organizational, not technical. Microsoft's Q3 FY26 earnings report 20 million paid Copilot seats growing 250% YoY, with half of Fortune 500 adopting Copilot Cowork within 6 months—yet 43.7% of organizations report their AI implementations are "implemented but not used," and 95% of enterprise AI pilots fail to deliver measurable P&L impact. All major vendors (Microsoft, Google, OpenAI, Anthropic) have independently converged on spreadsheet agents as core platform strategy, and Copilot Cowork GA demonstrates tested 30-40% cost efficiency gains on multi-step workflows. The capability ceiling is well-established: peer-reviewed benchmarks confirm frontier models "frequently fall short of professional finance standards" on multi-step logic, and independent analysis documents specific competitive displacement (GitHub Copilot lost to Cursor and Claude Code). Practice remains at bleeding-edge: elite deployment in bounded finance and operations workflows with clear ROI; mainstream stalled by organizational adoption barriers (training, manager modeling, governance) not feature maturity, compounded by ROI measurement failure (only 5-8% of enterprises achieve at-scale returns despite $186M average budgets).

CURRENT LANDSCAPE

August 2026 evidence reinforces the dual reality: platform convergence at scale paired with organizational adoption barriers and competitive displacement risks. Microsoft Q3 FY26 earnings report 20 million paid Microsoft 365 Copilot seats (250% YoY growth, fastest quarterly adds since launch) with monthly engagement now at Outlook levels; Copilot Cowork GA (June 2026) tested 30-40% cost efficiency gains and now comprises 49% of all Copilot tasks (up from 29%), with half of Fortune 500 adopted within 6 months. All major vendors (Google, Microsoft, OpenAI, Anthropic) have GA spreadsheet automation with multilingual support and agent coordination: Google's July 2026 Workspace features include Gemini formula error diagnosis (70.48% accuracy) expanded to 28 languages; Microsoft's work IQ context integration supports 30+ file types up to 50MB; Anthropic's Claude for Excel add-in (GA July 2026) offers multi-sheet reading with formula explanation and tracked changes. Yet specific capability benchmarks remain firm: SpreadsheetBench v2 (321 real workflows) tops at 34.89%; Vals AI Finance Agent capped at 23% on financial modeling; independent research reveals 70-80% accuracy on extraction but only 25-30% on decision-making tasks. Real production failures persist: FP&A teams report disabling Copilot within six weeks of June 2026 release due to unreliable multi-sheet and sign-convention handling.

The adoption/usage paradox is the dominant signal for mainstream stagnation. Deloitte's 3,235-leader survey (24 countries) shows 74% intend AI agent deployment by 2027, but only 21% have mature governance (53-point gap); critically, 43.7% of organizations implementing AI report it is "implemented but not used" due to organizational barriers (training gaps, manager modeling, permission issues)—a signal that features and adoption metrics diverge sharply from realized value. Industry-wide ROI measurement failure dominates: 95% of enterprise AI pilots show zero measurable P&L impact; only 5-8% achieve at-scale returns despite $186M average budgets; the root cause is integration gaps (tools not connected to underlying workflows) and budget misalignment toward sales/marketing instead of back-office automation. Finance sector remains strongest signal: accounting 52% adoption with 250% ROI, AR automation 40% payment acceleration, three-statement modeling 30→4 minutes, case studies showing Cloud Supply Chain 75% cycle time reduction and sales operations +9.4% revenue per rep via agentic automation. Specialized bounded-scope adoption signals emerge: Tracelight AI (error detection, not autonomous generation) deployed with 7 of top 10 management consulting firms and PE funds >$600B AUM. Mainstream adoption stalls: 60% achieve minimal ROI on investment; CSA May 2026 survey documents 82% discovered ungoverned shadow agents, 65% experienced security incidents, 61% reported data exposure. Strategic risks emerging: GitHub Copilot lost market lead to Cursor and Claude Code despite Microsoft's $13B OpenAI investment and distribution dominance—indicating competitive displacement in aligned markets. The practice bifurcates sharply: finance operations with bounded scope and verification gates show sustained ROI at scale; error-detection and review workflows (not autonomous generation) achieve enterprise scaling; mainstream knowledge work adoption blocked by organizational change barriers, not capability maturity. Practice remains at bleeding-edge: elite deployment in finance and operations with proven ROI; mainstream advancement blocked by adoption/usage gaps, ROI measurement failures, and competitive displacement in aligned segments.

TIER HISTORY

ResearchJun-2023 → Jun-2023
Bleeding EdgeJun-2023 → present

EVIDENCE (140)

— Microsoft internal case study: Cloud Supply Chain pilot deployed 70+ specialized agents reducing cycle time 75%; sales organization achieved +9.4% revenue per rep and -20% deal close cycle via agentic automation.

— Copilot Cowork GA with 30M+ paid seats; multi-step workflows now 49% of all tasks (up from 29%); tested 30-40% cost efficiency gain vs single-model alternatives; half of Fortune 500 adopted within 6 months of preview.

— Strategic analysis: Microsoft's <4.5% conversion (20M of 450M seats) despite $13B OpenAI investment; GitHub Copilot lost market lead to Cursor and Claude Code; spreadsheet automation faces competitive displacement despite distribution dominance.

Google Workspace Updates: Google SheetsProduct Launches

— Google GA announcements: Gemini-powered formula error diagnosis with one-click debugging and fix suggestions (70.48% accuracy); Fill with Gemini expanded to 28 languages; multi-series scatter and combo charts.

— Microsoft fiscal Q3 2026 earnings: M365 Copilot paid seats surpassed 20 million with 250% YoY growth; monthly engagement now at Outlook levels; GitHub Copilot Enterprise adoption nearly 140,000 orgs (tripled YoY).

— Independent analyst (GCN): 5 million seats added in Q2 (fastest quarterly growth since launch); companies with 50k+ seats quadrupled YoY; Accenture's 743k-seat deployment is largest publicly announced Copilot rollout.

— Enterprise adoption analysis: 43.7% of AI implementations are 'implemented but not used'; barriers are organizational (training gaps, manager modeling, permissions issues) not technical, revealing adoption/usage gap despite high licensing.

— Multi-source synthesis (MIT, BCG, KPMG, McKinsey, 2,100+ execs): 95% of AI pilots fail with zero measurable P&L impact; only 5-8% achieve at-scale ROI; failure rooted in workflow integration gaps and budget misalignment to back-office automation.

HISTORY

  • 2023-H1: M365 Copilot enters early access with Excel integration announced; initial user feedback reveals limitations in complex formula generation.
  • 2023-H2: Copilot reaches GA in October with formula suggestions and Python integration. Peer-reviewed research documents accuracy breakdown on complex problems. Practical deployments of ChatGPT+Sheets for data analysis appear. Feature availability gaps and access friction limit real-world rollout.
  • 2024-Q1: Google Sheets ecosystem solidifies as dominant platform with multiple AI integration approaches (native, third-party add-ons like SheetGen, app scripts). Third-party tool ecosystem expands (GptExcel, others). Copilot remains unavailable in desktop Excel despite GA claims; licensing/access issues persist. Research benchmarks confirm formula generation accuracy challenges via NL2Formula dataset.
  • 2024-Q2: Adoption metrics emerge: automated reporting platforms show 60%+ organizational adoption and 80% time savings. Real-world case studies of scale deployments (1,700+ response surveys with ChatGPT analysis). Institutional adoption programs launched (M365 Copilot training bootcamps). Desktop Excel access gaps persist despite training programs.
  • 2024-Q3: No significant new evidence of capability expansion or adoption barrier shifts captured during this window.
  • 2024-Q4: Google launches Gemini AI integration in Google Sheets and new =AI() function via Workspace Labs. Microsoft ships Copilot Lite in Microsoft 365 Family plans, expanding distribution. User feedback and peer-reviewed research document persistent reliability gaps: Copilot in Excel fails on common tasks (find-replace, pivot table analysis) despite GA claims. Formal trustworthiness framework published identifying hallucination risks in formula generation.
  • 2025-Q1: Google ships Gemini GA for chart generation and Python-driven insights in Sheets (January). Microsoft expands Copilot in Excel with Python-driven advanced analytics for forecasting and risk analysis (March). Meanwhile, SPREADSHEETBENCH benchmark reveals 75%+ of models score below 24% accuracy on real-world Excel forum queries. Alteryx analyst survey confirms adoption paradox: 70% say AI improves productivity, yet 76% still rely on spreadsheets and 45% spend 6+ hours weekly on data prep. Emerging third-party solutions (GRID) begin bridging AI-to-spreadsheet logic gaps. Capability expansion continues, but accuracy and logic interpretation barriers remain blocking broader tier advancement.
  • 2025-Q2: Ecosystem expansion: Sourcetable launches AI-powered spreadsheet platform with $4.3M seed funding (April), signaling continued venture interest. Production deployment issues emerge: NHS experiences Copilot outage blocking Excel file analysis in April (resolved November). Access friction persists: Copilot unavailable for Family/Personal tiers, causing user frustration documented across support forums. Vendor acknowledgment: Coherent publishes analysis arguing AI cannot reliably replace Excel for complex models, must remain complementary. Market adoption stalled: analyst productivity claims hold at 70%, but data prep time-sinks (6+ hours weekly, 45% of analysts) remain unaddressed. Core barriers unchanged: formula accuracy, access complexity, organizational rollout friction, trustworthiness concerns.
  • 2025-Q3: Vendor use-case contraction becomes visible: Microsoft ships =COPILOT formula but explicitly warns it unsuitable for accuracy-dependent work (financial, legal, calculations). Google ships incremental Gemini features for Sheets (table auto-formatting, =AI() text function). Paradigm AI and other third-party entrants launch with specialized agent ecosystems (5,000+ agents) as alternative to general-purpose tools. Landmark production deployment: UK government trial of M365 Copilot (1,000 licenses) finds no productivity gains and actual slowdown on complex analysis. Analyst adoption metrics frozen: 70% report productivity gains, 76% still rely on spreadsheets, 45% still spend 6+ hours on data prep. Market bifurcating into specialized text-focused and text-excluded use cases. Analyst skepticism deepens as vendor warnings narrow the applicable use case.
  • 2025-Q4: Google expands Gemini for multi-table analysis (October); Microsoft deprecates Copilot application skills in Excel (removal Feb 2026). Third-party ecosystem matures with 10+ distinct platforms (Julius, Equals, Arcwise, Rows, SheetGod, GPTExcel, SheetAI, Coefficient, Quadratic, others). Critical case studies emerge: companies with high-volume spreadsheet work (Lula Commerce, REVOLVE) abandon spreadsheets for dedicated BI platforms rather than adopt AI spreadsheet tools. Quadratic critiques Excel AI architecture as unsuitable for production data work. Feature expansion continues to mask use-case contraction and user exodus to non-spreadsheet platforms.
  • 2026-Jan: Adoption bifurcation crystallizes: financial services teams deploy spreadsheet automation at scale (accounting sector 52% adoption, 250% ROI, 20-30% capacity gains) while mainstream adoption stalls (56% of CEOs report no AI ROI). Microsoft scales back Copilot integration due to user trust concerns and privacy friction, signaling vendor pivot from aggressive embedding to tactical deployment. Third-party ecosystem remains active but no major breakthroughs. Core barrier: structured domains see strong unit ROI while long-tail knowledge workers face trust, cost, and governance obstacles preventing mainstream scaling.
  • 2026-Feb: Vendor feature expansion continues: Microsoft releases Agent Mode availability in EU and local file querying with Copilot Chat; Google announces admin usage reports for Gemini and forecasting in Connected Sheets via BigQuery ML. However, structural adoption barriers intensify: Microsoft 365 Copilot security bug (CW1226324) bypasses Data Loss Prevention policies, exposing confidential data; technical analysis shows 83% of AI-generated formulas fail under conditional formatting; organizational surveys reveal only 3% of enterprises highly transformed with AI while 72% remain early-stage. Accounting case studies document strong sector adoption (60% of firms, 25-35% time savings, 90-day ROI), yet broader market stalled by compliance concerns, formula fragility, and skills readiness gaps (61% use AI daily but only one-third prepared to adapt).
  • 2026-Mar: Adoption bifurcation sharpens: Microsoft Copilot seats reach 15M with 160% YoY growth and 3x increase in large deployments (35k+ seats); 60% Fortune 500 now deployed with 20-40% measured productivity on Excel data analysis. Google launches Fill with Gemini for Sheets with real-time web data integration and pattern summarization. Yet critical barriers persist: Excel data analysis achieves only 20% adoption despite wide distribution; complex workbooks score worse than random guessing (82% best-case vs >50% human baseline); governance gaps cause 60% of AI projects to fail; and measured ROI remains elusive despite high engagement. Finance sector remains strongest signal (52% adoption, 250% ROI). Core tension unresolved: feature acceleration in vendors masks fundamental accuracy, governance, and trust gaps preventing mainstream tier advancement.
  • 2026-Apr: Trust fracture and market bifurcation deepen. Microsoft's own terms of service contradict marketing claims: Copilot labeled "entertainment only" in ToS while marketing emphasizes productivity gains, revealing 3.3% real market penetration. Microsoft restricts free Copilot Chat access in Excel effective April 15, requiring $30/month per user licensing, signaling major pricing barrier and adoption friction. Meanwhile, leading third-party tool GPT for Work reaches 7M+ installations (ranked #1 in Kinross 2026 report, 4.9★ rating) and Google Workspace reports 128 customer deployments with named-org metrics (Geotab 89% adoption, Docusign 80% positive, Pinnacol 96% time savings). Microsoft ships Work IQ context-aware Copilot, Claude Opus 4.6 model support, and Copilot Notebooks for Excel generation. Yet Gartner forecast warns >40% of agentic AI projects will be discontinued by 2027 due to unclear ROI, uncontrolled costs, and insufficient data. Practice remains at bleeding-edge: strong specialized adoption in accounting, but mainstream tier advancement blocked by trust gaps between vendor marketing and contractual disclaimers, persistent formula accuracy failures under conditional formatting (83%), and licensing/pricing barriers preventing organizational scale.
  • 2026-May: Governance crisis crystallizes as primary adoption barrier, superseding feature maturity concerns. OpenAI ships ChatGPT for Excel/Sheets sidebar GA (May 8) with enterprise deployment options, removing app-store friction. CSA survey (418 security professionals) reveals 82% of organizations discovered ungoverned shadow automation agents created without IT/security/governance knowledge; 65% experienced security incidents; 61% reported data exposure from AI agents. Ramp Labs Sheets AI vulnerability (detected May 7, patched March 16) demonstrated prompt injection exploits allowing agents to exfiltrate financial data via formula injection. New peer-reviewed benchmark (WorkstreamBench, arXiv:2605.22664) evaluated LLM agents on end-to-end professional spreadsheet tasks in finance: Claude family leads overall but "frequently fall short of professional finance standards," with performance degrading sharply on tasks beyond simple chained calculations—a critical negative signal for production finance deployment. Enterprise adoption framework from EPC Group (200+ deployments, Fortune 500) documents 90-120 days to value versus 6-12 month industry baseline, with structured rollouts achieving 60-75% DAU versus 15-25% without structure. Finance automation ROI benchmarks (Kwestra, citing Hackett/Gartner data) show 45% cost reduction in AP/AR/close processes with 41% reaching breakeven within 12 months. Independent testing (Neuriflux, May 2026) documents Copilot in Excel achieving 7/10 formula generation success on first attempt with 12-second latency on 8K-row datasets. Finance sector shows bifurcated signal: AR automation metrics strong (40% payment acceleration, 90% error reduction, 91% mid-market success); but Deloitte survey of 1,300+ finance leaders shows 63% deployed financial automation yet only 21% report clear ROI—measurement gap masks benefit realization. Named case studies demonstrate payback: Mayo Clinic RPA deployment yielded 84,000 annual staff hours (18:1 ROI); City of Los Angeles licensing automation reduced processing from 45 to 6 days. Core tension: deep governance gaps (shadow agents, data exposure, lifecycle control, permission drift) now block mainstream adoption more than feature maturity or pricing, and even leading models fall short of professional finance standards on complex tasks.
  • 2026-Jun: Platform convergence and governance crisis arrived simultaneously. Microsoft Copilot Agent Mode (GA April 22) delivered +67% Excel engagement, +50% retention, and 65% satisfaction, while Google's semantic layer (Cloud Next) enabled end-to-end natural language spreadsheet construction and Microsoft Copilot Cowork (GA June 16) extended autonomous task execution to emails, meetings, and files—all major vendors converging on spreadsheets as a core AI substrate. Yet CVE-2026-42824 (SearchLeak, CVSS 10/10) exposed one-click data exfiltration from Copilot Enterprise Search via prompt injection, and EY's 150,000-user deployment confirmed that structured governance is the critical differentiator (94% monthly usage) vs. ungoverned rollouts (15–25% DAU). ROI data remained bifurcated: 60% of enterprises achieve minimal AI value despite investment; only 21% of finance leaders who deployed automation report clear ROI; while finance automation (accounting 52% adoption, 250% returns) continues to be the practice's strongest signal. Independent analysis confirmed spreadsheets are the AI substrate of choice by default—occupied by millions of users before AI design choices were made—rather than by design merit.
  • 2026-Jul: Capability expansion continued across all major platforms: Google GA'd a one-click formula error diagnosis and fix in Sheets (expanded to 28 languages), and Microsoft scheduled Excel text-column analysis for survey categorization for July 2026 GA—extending automation into knowledge-work classification tasks. GPT-5.4 Thinking achieved 87.3% on three-statement financial models (vs 43.7% for GPT-5), the strongest single-model result to date for finance workflows. However, the capability ceiling remains firm: SpreadsheetBench v2 (real business workflows) tops at 34.89%, an independent analyst benchmark documented real-world accuracy degradation from 97% (month 1) to 83% (month 8), and Deloitte's 3,235-leader survey found only 21% with mature governance against 74% intending agent deployment by 2027—a 53-point gap that remains the binding constraint on mainstream adoption. Later in the month, Claude's Microsoft 365 add-ins (Excel, Word, PowerPoint) reached GA across all paid plans alongside Claude Cowork's GA for agentic multi-step spreadsheet workflows, becoming a third major vendor entry into the space. New finance-specific benchmarks (Vals AI, Micro1) sharpened the capability ceiling further: models score 70-80% on extraction/calculation but only 23-30% on financial modeling and judgment tasks, and a practitioner review reported most FP&A teams quietly disabled Excel Copilot within six weeks of its June release due to unreliable multi-sheet and sign-convention handling.
  • 2026-Aug: Microsoft confirmed 20M+ paid Copilot seats (250% YoY, Q3 FY26 earnings) with a single-quarter add of 5M seats and Accenture's 743k-seat rollout now the largest publicly announced Copilot deployment, while an internal Cloud Supply Chain pilot of 70+ specialized agents cut cycle time 75% and sales operations saw +9.4% revenue per rep from agentic automation. Google shipped GA Gemini formula-error diagnosis (70.48% accuracy) in Sheets expanded to 28 languages, but a multi-source synthesis (MIT, BCG, KPMG, McKinsey) found 95% of AI pilots show zero measurable P&L impact and only 5-8% achieve at-scale ROI, and strategic analysis noted GitHub Copilot has lost competitive ground to Cursor and Claude Code despite Microsoft's distribution dominance and $13B OpenAI investment.

TOOLS