💹 Finance & Accounting
AI for financial operations, reporting, planning, and risk management. Over half the practices are good practice: fraud detection, expense management, invoice processing, and financial forecasting have mainstream adoption. Regulatory compliance and audit automation are advancing. The domain is tightly clustered around good-practice with minimal bleeding-edge — finance favours proven, auditable tools over experimental ones.
The Headline
Finance now uses AI to catch fraud and capture invoices. Anything an auditor has to sign off still needs a named person, and this fortnight's evidence shows that is unlikely to change soon.
The Picture
Most finance organizations now use AI somewhere. KPMG puts the figure at 75%, up from 30% in 2024, but most of that use is shallow. Fraud scoring, invoice capture, expense coding and bank matching are now routine, standard technology: 90% of euro-area banks use AI to detect fraud, and nobody serious is still debating whether invoice capture works. A small group is pulling ahead with one design: fixed, rule-based software does the arithmetic, AI drafts, and a named person signs. Performance Food Group automates 97% of its journal entries this way, and Cisco machine-drafts most of its annual-report management commentary before lawyers review it. Everyone else is stuck in extended pilots. Deloitte and Trintech find that 63% of companies have deployed AI in the month-end close, yet only 21% see measurable value. What separates the leaders is clean data and controls. The software is not the difference.
This Fortnight
A survey of 687 finance leaders found that only 28% are comfortable letting AI make routine finance decisions. The same PEX study found that interest in AI-drafted financial reports far outruns use, and respondents named accuracy as their top barrier. Plan your deployments around AI as a drafter whose work gets reviewed. Don't plan on it deciding.
The UK's audit regulator found that none of the six largest UK audit firms formally measures how its AI tools affect audit quality. Gartner separately found that nearly all audit leaders use some AI but only a small minority run formal, routine use cases. Ask your audit firm how it validates the AI it runs over your ledger before you rely on its reduced sample sizes.
A peer-reviewed study found AI tools misread figures from SEC filings up to 29% of the time. About half the errors were scaling mistakes, such as confusing thousands with millions. Error rates fell below 1% when the same figures came from structured, tagged filing data. Putting data in a clean structure buys more accuracy than switching to a better model.
A Forrester study commissioned by Airwallex found that over a third of finance leaders using AI are pausing further expansion. The main reasons were inexperience and difficulty proving returns beyond short-term gains, and payback expectations for finance AI are stretching from one year toward three or four. If your board expects fast returns, reset that expectation now.
Collections is where AI agents, software that acts without being prompted, are producing the clearest wins. LA Federal Credit Union cut 30-to-60-day delinquencies from about $12 million to $3.8 million in 21 days. Collections suits agents because an outcome, a payment or no payment, arrives within weeks, so this is a sensible place to test them.
Coming Up
US audit standard amendments take effect on December 15, 2026, bringing AI that touches journal entries and general-ledger postings inside the documented control environment. Auditors will expect you to have mapped every AI touchpoint, named a human validator for each and set a rule for when models are revalidated. Build that inventory before year-end fieldwork begins.
Tax practitioners enter the 2027 filing season under binding IRS duties to verify AI output. The Tax Court has already sanctioned a filing built on AI-invented case citations, and one benchmark found general chatbots accurate only 12% of the time on hard, multi-part tax scenarios. Ask your in-house team and outside advisers for their written AI verification procedures.
Seven more US states will require licensed-physician review before AI-issued health-claim denials from January 2027. They join 13 states with 2026 rules requiring human review before AI-only claim denials, while federal fair-lending enforcement is shifting to state regimes. If you deny claims or credit with AI, budget for audit trails that work state by state.
What's Hard About This
Accounting needs identical answers every time, and today's AI does not work that way. On a close-task benchmark, the best model produced a consistent correct answer across eight runs only 2.6% of the time, and a close runs to roughly 40 steps in sequence. Keep calculations in rule-based systems and limit AI to proposing and drafting.
Most AI failures in finance start in the data that goes into the AI. Only 11% of procurement organizations have a unified source-to-pay data model, and Fieldguide traced most mismatches between its AI and human auditors to missing firm context rather than AI error. Money spent on data foundations will pay back before money spent on new tools.
Accountability for AI output is personal and set in law, so review work eats into the savings. Sarbanes-Oxley, IRS practitioner rules and fair-lending law all put a human's name on the result, and Sage measured finance teams spending about 13 hours a week checking AI output. Count that review time in every business case.
Practices in this Domain (16)
| PRACTICE | TIER | TREND |
|---|---|---|
| Audit anomaly detection & trail analysis | LEADING EDGE | — Steady |
| Budget variance analysis & narrative explanation | GOOD PRACTICE | — Steady |
| Cash flow prediction | GOOD PRACTICE | — Steady |
| Collections, payments & purchase order automation | GOOD PRACTICE | — Steady |
| Contract pricing analysis | GOOD PRACTICE | — Steady |
| Credit risk assessment & scoring | LEADING EDGE | ↑ Accelerating |
| Expense management & reconciliation | GOOD PRACTICE | — Steady |
| Financial close & revenue recognition | LEADING EDGE | — Steady |
| Financial forecasting & scenario modelling | GOOD PRACTICE | — Steady |
| Financial reporting & management narrative generation | BLEEDING EDGE | — Steady |
| Fraud detection & behavioural analytics | ESTABLISHED | — Steady |
| Insurance underwriting & claims processing | GOOD PRACTICE | — Steady |
| Intercompany transaction management | LEADING EDGE | — Steady |
| Invoice processing, matching & exception handling | GOOD PRACTICE | — Steady |
| Supplier & spend analytics | ESTABLISHED | — Steady |
| Tax preparation & optimisation support | GOOD PRACTICE | — Steady |
Read the full technical briefing (1,865 words) →
Where AI Stands in Finance & Accounting
Finance has sorted its AI work into two layers, and the line between them is hardening. The lower layer holds tasks with a fast, unambiguous ground truth and a clear cost for being wrong: transaction fraud scoring, invoice capture, expense categorisation, cash application and bank matching. Here AI has become plumbing. ECB data show 90% of euro-area banks using AI for fraud detection. Visa paid $2.4bn for the behavioural-biometrics vendor BioCatch, and Socure has moved to close the investigation loop across a customer base that includes 18 of the 20 largest US banks. Invoice economics are settled too. Ardent Partners puts best-in-class processing at $2.65 an invoice against a $9.84 average, and nobody serious now argues about whether the capture layer works. The argument is about everything that sits on top of it.
That upper layer is where the domain stalls, and the tools are not the reason. Workday signed 170 customers to Adaptive Decision Intelligence in its first month. BlackLine's Verity agents automate 97% of journal entries at Performance Food Group. Oracle, SAP, Microsoft and Workiva all ship agentic close, payables and reporting features. The trouble is that finance work has to be right every time and defensible to an auditor, and today's models are not built for that. On the APEX-Accounting benchmark, the best model meets 56.4% of grading criteria on average but produces a consistent correct answer across eight runs only 2.6% of the time. A month-end close runs to roughly 40 sequential steps, so even strong per-step accuracy compounds towards failure. The deployments that work therefore share one design: deterministic engines do the arithmetic, AI drafts or proposes, and a named human signs. Cisco's finance team says 80–90% of its MD&A is now machine-drafted, but lawyers review it before it reaches the SEC. What sets finance apart from neighbouring domains is that accountability here is personal and statutory. SOX, PCAOB standards, Circular 230 and fair-lending law all put a human's name on the output, whoever or whatever produced it.
The result is a gap between breadth and depth that shows up in almost every survey. KPMG reports that 75% of finance organisations now use AI, up from 30% in 2024. PEX, surveying 687 finance leaders, finds only 31% using it anywhere in finance and just 9% at a "run" stage of maturity. Deloitte and Trintech find 63% have deployed AI in the close but only 21% see measurable value. The surveys disagree on how many organisations use AI, but they agree on how shallow most use is. Two practices are gaining ground. Audit anomaly detection is being pulled forward by regulators and by the appeal of testing entire populations rather than samples: Cherry Bekaert used MindBridge risk scoring to cut one sample from 384 items to 252. Tax preparation has broadened through specialised, human-assisted platforms. TurboTax Live now makes up 53% of Intuit's tax revenue, and H&R Block's AI Tax Assist handled 4.2m interactions. General-purpose chatbots, by contrast, still fail at tax work. Everywhere else, the data foundations, controls and measurement that should sit under the vendor capability are still being built.
What's New, 2026-09-10 to 2026-09-24
This fortnight's evidence mostly confirmed the trust gap rather than closing it. The PEX survey gave the clearest numbers. Of its 687 finance leaders, 61% are interested in AI-generated financial reports but only 14% use them, and just 28% are comfortable letting AI make routine finance decisions. Accuracy was the top barrier, cited by 36%. Audit brought the sharpest findings. Gartner found that 93% of audit leaders use some AI but only 15% have formal use cases that run routinely, and the FRC found that none of the six largest UK audit firms formally measures how its AI tools affect audit quality. Fieldguide's engineering review of its own control-testing agents helps explain why. It traced 66% of mismatches between agent and auditor to missing firm context, and only 25% to agent errors. In reporting, a peer-reviewed Suffolk University study measured LLM error rates of 7–29% when models retrieve figures from SEC filings, about half of them scaling mistakes, against under 1% on structured XBRL data. A Macabacus survey found 61% of financial-services respondents believe an AI error reached a client or decision-maker in the past year, while only 23% of firms have comprehensive guardrails. IOSCO issued a supervisory toolkit for AI in capital markets that scores human oversight and the risk of generative AI inventing figures.
Two practices look different after this scan. Fraud detection now reads as settled infrastructure, though this fortnight also brought named failures. Australia abandoned a Medicare fraud-detection tool that audit found to be 22% accurate. A vendor case study showed foundation-model embeddings improving card-issuer fraud detection by 24–35%, but only after a first version failed. In India, the window for stopping a fraudulent payment has shrunk to 30–60 seconds. Tax preparation looks more widely adopted, but through assisted platforms such as Black Ore (now at 75 firms), Accrual and Grove's MCP connector, not through autonomous filing. A Saturn benchmark put general chatbots at 43% accuracy on financial questions and 12% on hard multi-part tax scenarios. A survey of 799 Dutch tax advisers found 80% of juniors using AI daily, yet 88% of respondents named output quality as the main barrier. New deployments kept arriving. A Tier-1 bank's voice agents handle 450,000 delinquent-account interactions a month, LA Federal Credit Union cut 30–60-day delinquencies from about $12m to $3.8m in 21 days, AXA is targeting €500–700m a year in AI value by 2029, LKQ automates more than 90% of over 100,000 monthly intercompany transactions on Trintech, and League replaced Workday Adaptive Planning with Pigment. The counterweights were just as concrete. A Forrester study for Airwallex found over a third of AI-using finance leaders pausing expansion. Oliver Wyman found only 10% of tech leaders can trace costs per agent. Ardent found only 11% of CPOs have unified source-to-pay data. Sedgwick found 82% of insurers using AI but 7% scaling it. And a San Francisco Fed study of 1,006 banks linked AI adoption to 0.38 percentage points higher ROA, but also to a larger share of problem loans.
Key Tensions
Probabilistic models in a deterministic ledger. Accounting needs the same inputs to produce the same answer every time, and today's models do not work that way. APEX-Accounting's 2.6% consistent-success rate, Suffolk's 7–29% error rates on retrieval from SEC filings and Saturn's 12% on hard tax scenarios all point to that mismatch. So the working pattern keeps settling in one place. Deterministic engines post and calculate, and AI proposes or drafts. A Canadian distributor's deterministic SKU matching caught a $23,908 extraction error in its first production run, and a five-entity manufacturer cut its close from nine days to three after making every job idempotent.
Interest is broad, authority stays with humans. Finance leaders will let AI draft, flag and recommend, but they will not let it decide. PEX found 28% comfortable with AI deciding routine matters. Deloitte found 95% comfortable with agentic workflows but only 14% supportive of full autonomy on critical decisions. In insurance, ISG found 86% of carriers keeping authority for consequential decisions. Because every automated step still ends at a reviewer, many time savings turn into verification work: Sage has measured about 13 hours a week spent checking AI output.
The bottleneck is the data layer, not the model. Many of the failures blamed on AI start upstream of it. Ardent found only 11% of procurement organisations have a unified source-to-pay data model. Snowflake's internal invoice pipeline found workflow fragmentation, not parsing, to be the bottleneck. AppZen found that 21.9% of AP staff time goes on supplier inquiries, which capture automation moves around rather than removes. In the close, a nine-firm Finnish study named data fragmentation and years of data cleansing as the main constraint, and 54% of companies still process intercompany transactions by hand.
Use is outrunning measurement. Organisations are deploying AI faster than they can prove it works or know what it costs. None of the six largest UK audit firms formally measures AI's effect on audit quality. Only 15% of audit leaders run formal use cases routinely, and only 10% of tech leaders can trace costs per agent. KPMG found 49% of organisations scaling back agent deployments when operating costs outran the value they could see, and payback expectations for finance AI are stretching from one year towards three or four.
Accountability is spreading faster than clear rules. Liability for AI output keeps landing on people and institutions, but the rules are splitting across regulators and jurisdictions. Tax practitioners face binding verification duties under IRS OPR Alert 2026-19, and Tax Court sanctions have followed hallucinated citations in Clinco v. Commissioner. In lending, the federal retreat from disparate-impact liability has made state regimes the fair-lending battleground. In insurance, 13 US states now require human review before AI-only claim denials. The same pressure appears in the labour market: claims-adjuster job postings are down 55% even as regulators insist a licensed human stays responsible for each denial.
Top 10 Evidence Items
- The state of finance: AI and automation in finance operations (adoption-metric) — The PEX survey anchors the breadth-versus-depth gap: 31% use AI at all and trust in accuracy is the leading barrier. https://www.pexcard.com/research/state-of-finance/
- Is AI Financial Analysis More Accurate? What the Research Shows (research-paper) — Peer-reviewed evidence that LLMs misread SEC-filing figures at 7–29% error rates, which is why deterministic engines must do the arithmetic and AI only drafts. https://www.finrep.ai/blog/is-ai-financial-analysis-more-accurate-what-the-research-shows
- AI in Audit Regulation, 2026 (industry-report) — The FRC finding that none of the six largest UK audit firms measures AI's effect on audit quality is the clearest case of use outrunning measurement. https://www.dnl.ai/resources/ai-in-audit-regulation-2026
- Auditing an audit agent (case-study) — Shows that most agent-auditor mismatches come from missing firm context rather than model error, which supports the argument that data and context are the bottleneck. https://www.fieldguide.com/engineering/auditing-an-audit-agent
- DoHDA invented an AI-powered fraud detection tool. It stunk (news-coverage) — A named failure: a government fraud tool that was 22% accurate and abandoned, showing that fraud detection's 'settled' status still has casualties. https://www.medicalrepublic.com.au/dohda-invented-an-ai-powered-fraud-detection-tool-it-stunk/129291
- AI chatbots give wrong financial answers most of the time, study finds (news-coverage) — Puts numbers on why general chatbots fail at tax work (12% on hard scenarios), which is the reason assisted platforms are winning. https://www.investmentnews.com/fintech/ai-chatbots-give-wrong-financial-answers-most-of-the-time-study-finds/268267
- AI-generated errors are reaching clients at most financial firms (adoption-metric) — Only 23% of firms have comprehensive guardrails while 61% believe an AI error reached a client, so accountability is exposed in practice. https://www.wealthprofessional.ca/news/industry-news/ai-generated-errors-are-reaching-clients-at-most-financial-firms/393506
- Opacity, Not Billing, Is the Real AI Risk Under Circular 230 (opinion) — Regulatory move: IRS OPR Alert 2026-19 puts binding verification duties on practitioners, showing personal statutory accountability for AI output. https://www.cpapracticeadvisor.com/2026/09/17/opacity-not-billing-is-the-real-ai-risk-under-circular-230/190326/
- Why CFOs Are Pausing AI Expansion (adoption-metric) — Over a third of AI-using finance leaders are pausing expansion, a concrete counterweight to the vendor deployment announcements. https://www.airwallex.com/global/blog/why-cfos-are-pausing-ai-expansion
- LA Federal Credit Union collections automation: $12M delinquency reduced to $3.8M in 21 days (case-study) — A positive deployment with hard numbers on delinquency reduction, showing where the lower-risk, ground-truth-rich work delivers. https://financeops.ai/blogs/7-signs-your-ar-team-is-losing-money-to-manual-collections