⚖️ Legal, Compliance & Risk
AI for managing contracts, regulation, governance, and organisational risk. Contract review and e-discovery are good practice with proven ROI; regulatory monitoring and due diligence are advancing steadily. Most of the domain sits at leading-edge — adoption is constrained by liability concerns and the need for domain-expert validation rather than by tooling gaps.
The Headline
Almost every legal and compliance team uses AI, but very few run their work on it. Regulators and courts have started demanding proof that a human checked the output.
The Picture
Most organizations are in the same place. The American Corporate Counsel association puts in-house AI adoption at 85%, yet AI actually executes under 5% of contract review work. A small group is pulling ahead where the output can be checked against a known answer. Big Four auditors score entire ledgers, large banks triage sanctions alerts, and litigation teams run document review at a fraction of the attorney hours it used to take. Where the output is judgment, like a brief, a risk score or a policy rewrite, the gains go to firms that spent years building verification discipline, and the liability stays with a named human. The rest are finding that supervisors now ask for evidence that controls worked, not for policies saying they exist.
This Fortnight
OpenAI launched Astra for Law on September 17, selling its most capable AI directly to lawyers. It follows similar legal launches this year from Anthropic and Google. OpenAI's own test shows it answering 54% of legal research questions correctly, and access is limited to selected large firms plus Harvey and Legora, which use it as a building block. Expect sharper pricing pressure on incumbent legal research vendors, and treat the accuracy figures as marketing until someone independent tests them.
State bank supervisors published a common playbook for examining AI at the 3,355 banks they oversee. The Conference of State Bank Supervisors framework includes a 28-page examiner work program and a worksheet that ranks each AI use by risk. Examiners can ask for AI inventories, board reporting, chatbot transcripts and vendor contract terms, and it covers the generative and agentic AI (software that acts on its own without being prompted) that April's federal guidance left out. If you run a state-chartered bank, check now that you can hand over each of those documents.
A federal appeals court proposed requiring lawyers to certify that a human reviewed any AI-assisted filing. The Tenth Circuit acted as Reuters counted AI hallucinations (when an AI tool confidently makes things up) in at least 1,395 US court cases, including fabricated police testimony in a murder appeal. In a separate case against Shell, a court ordered an expert witness to disclose her AI prompts, which means expert preparation done with AI can now be demanded in discovery.
A blind test of seven contract review tools found detection of hidden change-of-control risk falling as low as 49%. The same tools caught plainly worded clauses most of the time, but missed triggers buried in definitions or in merger language. Ask vendors to test on your own hardest contracts before you trust any headline accuracy figure.
The US government's free export-screening list feed was shown to return incomplete results without any warning. In one live test, a query that reported 900 matches returned only 10. Any screening process built on that feed can issue a false "no match" on a restricted party, so find out whether yours depends on it.
Coming Up
Comments on the Tenth Circuit's certification rule close October 18, and the rule would take effect January 1, 2027. More than 140 federal judges already have their own AI orders, a California bill mandating disclosure and citation checks has cleared the legislature, and sanctions are rising on purpose. Before January, build logged, mechanical citation checking into every litigation workflow, including outside counsel's.
Examiners in banking, insurance and capital markets are converging on the same AI evidence requests. Twelve state insurance departments are piloting AI questions inside routine examinations, with wider adoption up for decision at the insurance regulators' fall meeting, and the global securities regulators' body has set similar expectations. Assemble a complete AI inventory with owners, testing records and human-review logs now, because putting it together after the examiner asks is what fails.
The EU's high-risk AI obligations, which cover AI document review and anti-money-laundering screening, now apply from December 2, 2027. Transparency duties have been enforceable since August 2, and the delay does not change what must eventually be documented. Use the extra time to build continuous records of how each high-risk system was tested and overseen, rather than to postpone the work.
What's Hard About This
Checking AI output is the lawyer's own job, and it eats the time savings. In a California case, a judge found eight fabricated quotations in briefs prepared with the publisher's own integrated citation-checking tools, and the state Supreme Court let that opinion stand. Buying better tools does not transfer the duty, so budget for verification as a permanent cost of every AI-assisted filing.
Owning compliance software is not the same as being compliant. One audit found 67% of deployed consent platforms still firing marketing tags after users said no, and GE Aerospace's export penalty cited staff overriding its automated screening system. Regulators now fine broken configurations, so test your controls continuously instead of assuming the dashboard is right.
Compliance teams are shrinking just as governing everyone else's AI becomes their main job. An ISACA survey in June put the median privacy team at five people, down from eight a year earlier, while a OneTrust survey found only 28% of organizations produce AI governance evidence. The function meant to adopt AI is increasingly occupied with policing it, and that trade-off needs an explicit staffing decision.
Practices in this Domain (21)
Read the full technical briefing (1,778 words) →
Where AI Stands in Legal, Compliance & Risk
Legal, compliance and risk has split along a line drawn by verifiability rather than by technology. Where output can be checked against a population, a list or a known answer, AI is in production at scale and the argument is about depth. That covers full-ledger anomaly scoring, sanctions and AML alert triage, clause extraction, document review and data-subject request handling. KPMG Clara, with MindBridge anomaly detection built in, is rolling out to more than 95,000 auditors, and EY Canvas processes 1.4 trillion journal entries a year. Ripjar says its screening platform serves a quarter of the world's systemically important banks. In Redgrave LLP's benchmark, Relativity aiR reached 88% recall on 45,000 documents in 18 attorney hours, against 1,123 for traditional review. Where output is advocacy or judgement, the tooling is just as available and the gains are just as real at firms that have built the checks. That covers bespoke drafts, risk scores that recommend whether to sign, briefs, outcome forecasts and rewritten policies. The difference is that liability stays with a named human who cannot hand it to the vendor. This is what sets the domain apart from its neighbours: its product is evidence that regulators, courts and counterparties test, so a wrong answer is not merely a defect but a sanction, a fine or a lost claim.
The second defining feature is how far breadth of use runs ahead of depth. The ACC finds in-house AI adoption at 85%, up from 23% in 2024, yet AI execution stays below 5% even for contract review. Regnology and Oliver Wyman surveyed 276 practitioners in 22 countries and found 87% exploring, piloting or embedding AI in regulatory reporting, but only 16% at embedded production. Embedded use sits between 8% and 15% at every institution tier, so size does not buy progress. Among financial crime leaders, 4.7% continuously adapt AML controls as risks change, and 76.3% still review alerts manually or with only partial automation. Ethisphere found every ethics and compliance team using AI but only 5% measuring its impact. The binding constraints are fragmented source data, integration work, thin governance evidence and shrinking teams: ISACA puts the median privacy team at five people, down from eight a year earlier. Model capability is rarely the thing in short supply.
The market is reorganising around that gap. Frontier model vendors are now selling to lawyers directly. Anthropic launched Claude for Legal in May, Google launched Gemini Enterprise for Legal in August, and OpenAI added Astra for Law this month. Thomson Reuters has responded with its own legal LLM, trained on Westlaw and Practical Law content. Capital keeps arriving: Harvey lifted its round to US$600m at a US$15.5bn valuation, Norm AI reached a US$1.2bn valuation for agentic compliance, and Bloomberg Law counts more than US$100m of 2026 funding aimed at incumbent contract-management vendors. Meanwhile compliance functions have taken on a second job: governing everyone else's AI. EU AI Act transparency duties have applied since 2 August, supervisors are writing AI examination manuals, and inventories and audit trails for AI are becoming examinable artefacts. On current evidence, that second job is growing faster than the first.
What's New, 2026-09-16 to 2026-09-30
It was a normal fortnight, and no practice changed its position. The news sharpened existing lines rather than redrawing them. The biggest product event was OpenAI's Astra for Law, released on 17 September. It is a GPT-6 configuration with a legal index of more than 230 million URLs, covering over 99.9% of published US precedential case law via the Free Law Project. OpenAI self-reports 54% on 200 Vals AI Legal Research Bench questions, against 38.7% for the base model with web search. Access is limited to a Trusted Access programme, with Harvey and Legora as API customers. Supervisors moved at the same time. On 16 September the Conference of State Bank Supervisors published an AI Supervisory Framework for the 3,355 state-supervised banks. It includes a 28-page examiner work programme and a worksheet that scores each use case Tier 1 to 3, and it covers the generative and agentic models that April's federal model-risk guidance left out. On 18 September the Tenth Circuit proposed requiring certification of human review for AI-assisted filings from 1 January 2027. Reuters counted AI hallucinations in at least 1,395 US court cases, including fabricated police testimony in a murder appeal. In Conservation Law Foundation v. Shell, a court ordered disclosure of an expert witness's AI prompts, which opens a new discovery front. Vendors kept humans at the gate: Leah Contracting launched on 18 September with irreversible actions requiring approval, and OneTrust's new automated policy-violation actions shipped only as a public preview.
The more consequential evidence undercut headline accuracy. The Legal Stack blind-tested seven contract-review platforms on 50 agreements. They detected obvious change-of-control language at 81–96%, but embedded definitional triggers fell to 49–61% for Luminance, Harvey and Kira, and merger-deemed assignment to 22–47%. A live measurement of trade.gov's Consolidated Screening List API found silent truncation: one query reported 900 matches and returned 10, and paging capped out at 1,050 of the Entity List's 3,419 records. Each can produce a false "no match". A Censuswide survey for Feroot found 98.1% of organisations have or plan a consent platform, but only 24% continuously verify that it is enforced. A UK government review of 15 futures leads found those with AI experience almost unanimous that it cannot yet do horizon scanning. Among the counterweights, Alibaba reports AliExpress's AI BrandSafe programme removing infringing listings at about nine times the volume of rights-holder requests, 17 times for Korean brands, with 94% removed before a single sale. Separately, an Australian National Audit Office audit found that an AI Medicare fraud model, which ran from July 2024 to December 2025, contributed 8 of 3,779 identified cases in 2024–25. It is a sober reminder that deployment is not the same as effect.
Key Tensions
Everyone uses it, almost nobody runs on it. Survey after survey shows near-universal use sitting on top of largely manual operations: 85% in-house adoption but under 5% AI execution in contract review, and 16% embedded production in regulatory reporting. Regulatory change management is still 64% manual and policy management 72%, according to the IAPP's report earlier this year. Until data, integration and ownership are fixed, headline adoption tells buyers little about whether work has actually moved.
The verification duty cannot be delegated, and it consumes the savings. Courts treat checking every citation and quotation as the lawyer's own job, and they are raising the price of failure. The Tenth Circuit now proposes certification, and an Illinois appellate court imposed $15,000 in sanctions while signalling that fines will rise. In Quinteros v. Harbor Distributing, eight fabricated quotations got through even with Lexis Protégé and Citation Check in use. The California Supreme Court declined to depublish that opinion, so integrated tooling does not remove the burden.
Headline accuracy hides a long tail of silent misses. Tools score well on stated risk and routine clauses, then fall away on embedded triggers, cross-references, scanned documents and missing protections. The winning system on the DocILE benchmark of real business documents reached only 70.2% average precision. The FCA found screening systems missed one in four minor name variants. Because vendors rarely publish how they measure accuracy, buyers find the tail in production rather than in the proof of concept.
Governing AI is becoming the compliance workload. Examiners now ask for AI inventories, board reporting, human-review logs and vendor terms: the CSBS framework, the IOSCO toolkit and the NAIC's insurance exam pilot all point that way. Yet OneTrust's survey of 1,200 decision-makers found only 28% produce governance evidence or audit trails. In another survey, compliance or risk teams blocked 48% of agent deployments. The function that should be adopting AI is increasingly occupied with policing it.
Owning the tool is not the same as complying. Enforcement targets broken configurations rather than missing software. Feroot's audits found 67% of sampled consent platforms still firing marketing tags after a user rejected them, and 93% of audited sites failing to honour Global Privacy Control. GE Aerospace's export penalty cited manual overrides of automated compliance systems. Regulators increasingly ask whether controls worked, which tooling alone cannot prove.
Top 10 Evidence Items
- The Legal AI 'Change of Control' Clause Audit Report 2026: How AI Contract Review Tools Perform on the Specific Clause Category That Determines Whether Your Entire Agreement Survives an Acquisition (industry-report) — Directly evidences the 'headline accuracy hides a long tail' tension: obvious clauses score well but embedded triggers collapse, exactly the silent-miss pattern the briefing warns buyers find only in production. https://www.thelegalstack.org/research/the-legal-ai-change-of-control-clause-audit-report-2026-how
- Judgment Will Remain Crucial As Lawyers Learn To Embrace AI (opinion) — The ACC's own 85%-adoption-vs-under-5%-execution figure is the briefing's headline stat for breadth running ahead of depth. https://www.forbes.com/sites/rogertrapp/2026/09/28/judgment-will-remain-crucial-as-lawyers-learn-to-embrace-ai/
- OpenAI's Astra for Law: What its AI push for legal sector means for rivals (news-coverage) — Astra for Law is the fortnight's defining product event, showing frontier vendors now selling directly to lawyers rather than through incumbents. https://www.business-standard.com/world-news/openai-astra-for-law-what-its-ai-push-for-legal-sector-means-for-rivals-126092100585_1.html
- State bank examiners get a playbook for inspecting AI (news-coverage) — Shows the compliance function's 'second job' becoming concrete: examiners now have a formal 28-page programme for auditing AI itself. https://www.americanbanker.com/news/state-bank-examiners-get-a-playbook-for-inspecting-ai
- trade.gov's Consolidated Screening List API: Nine Traps in the Search Endpoint | Aervik Labs (tutorial) — A live, measured failure of a screening data source producing false 'no match' results — the kind of silent gap that undermines verifiability claims. https://aerviklabs.com/guides/trade-gov-consolidated-screening-list-api/
- The AI Governance Gap Isn’t a Policy Problem. It’s an Evidence Problem (adoption-metric) — Quantifies the governance-evidence gap (28% keep audit trails despite 87% encouraging agent use) that anchors the 'evidence problem, not policy problem' tension. https://www.kiteworks.com/regulatory-compliance/ai-governance-evidence-gap/
- Federal Appeals Court Moves to Crack Down on AI Filings After Lawyers Cite Fake Cases (news-coverage) — The Tenth Circuit's proposed certification rule is the clearest sign that courts are raising, not softening, the non-delegable verification duty. https://www.lawcommentary.com/articles/tenth-circuit-ai-court-filings-human-review-rule
- AI error-ridden court filings surge despite three years of sanctions (news-coverage) — 1,395 hallucination cases shows the verification burden is scaling with adoption rather than shrinking, consuming the promised savings. https://www.reuters.com/legal/government/ai-error-ridden-court-filings-surge-despite-three-years-court-sanctions-2026-09-17/
- 89% of Banks Fail to Move AI into Regulatory Reporting Production, Regnology Study Finds (press-release) — The Regnology 'agentic gap' finding (87% piloting, 16% embedded) mirrors the ACC stat in a different sector, showing the adoption-execution split is structural, not domain-specific. https://ffnews.com/news/new-research-89-of-financial-institutions-have-yet-to-embed-ai-into-regulatory-r-0f6173ad
- Artificial Intelligence and Medicare Benefits Integrity (industry-report) — The ANAO Medicare audit is the sharpest 'deployment is not the same as effect' data point: a live model contributed to only 8 of 3,779 cases found. https://www.anao.gov.au/work/performance-audit/artificial-intelligence-and-medicare-benefits-integrity