Est.
AI in LegalLong read

Hallucination Risk in AI-Assisted Legal Document Review

Lawyers who trust AI contract review without auditing for hallucination face real sanctions risk.

Features Editor · · 13 min read
Cover illustration for “Hallucination Risk in AI-Assisted Legal Document Review”
AI in Legal · September 21, 2026 · 13 min read · 2,934 words

Hallucination in AI-assisted legal document review is a measured, published, recurring failure mode with a sanctions record attached to it. It's a measured, published, recurring failure mode with a sanctions record attached to it. In-house legal teams that treat it as a solved problem, or worse, as marketing fine print, are exposed in ways the profession is only now starting to price correctly.

Start with the definition, because it gets flattened in casual use. Hallucination is a model generating false content with the same fluency and confidence it uses for true content. It's a model generating false content with the same fluency and confidence it uses for true content: a fabricated citation, an invented clause, a legal standard that doesn't exist, all delivered without any internal flag that something's wrong. There's no hedge, no asterisk, no "confidence score" attached that a reviewer can trust to catch it. That absence of a self-correction signal is what makes hallucination structurally different from ordinary unreliability, and it's why the failure mode is so easy to miss until it's already in a filed document.

The benchmark that reframed this conversation came out of the Journal of Empirical Legal Studies in 2025: the first preregistered empirical evaluation of AI legal research tools. It found that purpose-built products from LexisNexis (Lexis+ AI) and Thomson Reuters (Westlaw AI-Assisted Research and Ask Practical Law AI) hallucinate between 17% and 33% of the time. These are tools built and sold specifically for lawyers, not consumer chatbots repurposed for legal work. They're tools built and sold specifically for lawyers, with vendors claiming that retrieval-augmented generation (RAG) "eliminates" or "avoids" hallucination. A separate Stanford and Yale study (Dahl et al.) found that legal-specific RAG-based tools hallucinate in the 17–34% range, a finding that points to a pattern consistent across benchmark designs. It's a pattern that holds across separate research teams asking the same question in different ways.

RAG helps. It doesn't fix the problem. Retrieval-augmented generation reduces hallucination relative to general-purpose models by grounding outputs in retrieved documents, but the Stanford/Yale researchers found that none of the vendor claims of "eliminating" hallucination came with empirical evidence behind them. And there's a second wrinkle: benchmark performance doesn't predict production performance. When Vectara expanded its hallucination benchmark in November 2025 to 7,700 longer articles spanning law, medicine, finance, and other fields, hallucination rates rose sharply, simply because the tasks got closer to what real enterprise work actually looks like. Clean, short-form benchmarks flatter these tools. Longer, messier, real-world tasks expose them.

If purpose-built legal research tools working against indexed, public caselaw still hallucinate at these rates, contract review deserves its own scrutiny, and a harder one. Contracts are private, unindexed, often unique documents, not public case law sitting in a training set somewhere. They're private, unindexed, often unique documents, and that changes the risk profile substantially.

Where Hallucination Originates and Why Contracts Amplify It

Generative models are built to produce the next statistically plausible chunk of text, not a verified fact. That's the whole mechanism. There's no built-in process for the model to check its own output against ground truth, because the model doesn't have a ground truth to check against, only patterns learned from training data. When a gap in that data occurs, the model doesn't stop and say so. It fills the gap with something that sounds right.

Sterne Kessler's 2025 review of legal AI failures grouped the recurring patterns into three buckets: citations to cases or documents that don't exist, citations to real cases that don't say what's claimed, and quotes pulled from real sources that either fail to support or flatly contradict the legal point being made. Each of these gets worse, not better, when the input is a contract instead of a case.

Here's why. Case law is public and heavily represented in training data; contract language is private, proprietary, and rarely seen by any model during training, so when a model is asked to analyze one, it's often working from thinner priors than it would for caselaw, which pushes it toward filling gaps with plausible but invented language. Commercial contracts also lean hard on precisely defined terms, where a subtle paraphrase can quietly shift the scope of a liability cap or an indemnity carve-out without anyone noticing the substitution happened. And clause interdependency makes this worse still: an indemnity provision might be narrowed by a limitation-of-liability clause three sections later, and catching that kind of cross-reference requires multi-step reasoning, exactly the terrain where non-deterministic outputs are most dangerous.

General-purpose models weren't trained on contract-specific risk logic and don't carry an understanding of what a limitation-of-liability clause is supposed to do in a negotiated deal, and they can produce confident-sounding answers regardless. An independent enterprise AI analysis published in 2025 went further, classifying unsupervised contract review by general-purpose LLMs as "high risk, not suitable," citing non-deterministic outputs and the absence of a native legal audit trail as the primary blockers. The risk isn't spread evenly across every task a lawyer might hand to an AI tool. It concentrates in the tasks that demand exact reproduction of defined terms, synthesis across multiple clauses, and interpretation that depends on jurisdiction.

The Most Consequential Points for Hallucination in Contract Review

Not every task carries the same exposure, and treating a given automated use case as one undifferentiated risk misses where the actual damage happens. Leading AI contract review platforms hit accuracy in the 85-95% range on clause identification for standard agreements like NDAs, MSAs, and vendor contracts. Human reviewers, working the same task, score between 80% and 90%. AI holds its own here, sometimes beats the human baseline, but at enterprise volume, even a small error rate compounds fast, and the margin still matters.

Some moments in the workflow carry outsized weight. Clause extraction feeding into playbook comparison is one: if the AI misreads a damages clause, say, reporting "no consequential damages" when the contract actually contains a carve-out, that error flows straight into a false "acceptable" signal, and nobody downstream has a reason to question it. Obligation and deadline tracking is another, and arguably the sharpest one, because a missed renewal trigger or an invented notice period doesn't become visible as an error until the deadline has already passed and the damage is done. Negotiation summaries handed to business teams are a third pressure point: if a summary misstates where the parties actually landed on a liability cap or an IP ownership term, the business side makes decisions on a false picture of its own exposure. And post-signature, hallucinated metadata sitting in a contract repository poisons everything built on top of it, every search, every report, every risk query run against that data going forward.

Not everything needs the same scrutiny: lower-risk moments exist too. Initial triage and routing, flagging unusual clauses for a human to look at, generating a first-pass redline that a lawyer will review anyway: these are places where an AI error is recoverable, because a human sees the output before it does anything. Risk scales with how far downstream an output travels before a human checks it, and with how much that specific clause matters to the deal's actual economics.

Professional Liability and the Sanctions Record for Failed Verification

Courts are no longer treating this as novel. A manual count from the Charlotin database, tracked by the Drug & Device Law blog, found 51 lawyer hallucination cases in December 2025, 36 in January 2026, and 33 in February 2026 with the month still incomplete. That's a caseload. That's a caseload.

The named cases sketch out the range of what verification failure actually looks like in practice. In Wadsworth v. Walmart, attorneys at Morgan and Morgan and the Goody Law Group filed motions citing nine cases, eight of which turned out to be fake, which should put to rest any assumption that firm size or internal tooling is a reliable safeguard. In Coomer v. Lindell, two attorneys representing MyPillow's CEO filed a brief with nearly 30 defective citations, some to cases that don't exist, and in July 2025 a federal court in Colorado fined each attorney $3,000. Whiting v. City of Athens went further still: in March 2026, a Sixth Circuit panel found fabricated citations across the attorneys' briefs and ordered each to pay $15,000 to the court registry, reimburse the opposing side's full appellate fees across three separate appeals, pay double costs, and face a disciplinary referral. It's the clearest federal appellate statement yet of what the verification duty actually requires.

Kohls v. Ellison shows that the failure isn't confined to litigators drafting briefs under deadline pressure. A Stanford expert's declaration, drafted with help from ChatGPT-4o, cited two academic articles that don't exist and misattributed a third. In January 2025, the federal court struck the declaration entirely, finding that the fabricated citations had wrecked the expert's credibility on everything else in it. And in Idehen v. Stoute-Phillip, heard in the Civil Court of New York, Queens County, an attorney's affirmation citing seven fake cases pushed the court to describe the conduct as escalating "from merely frivolous to egregious misconduct that implicates his honesty, trustworthiness, and fitness to practice law."

None of this required new legislation. Existing rules, FRCP 11(b) chief among them in US federal practice, are doing the work, and Sterne Kessler's 2025 review found that courts come down harder when AI use is obvious but the attorney denies it. The standard is spreading internationally too: the Paris Bar Association's October 2025 White Paper and the French National Bar Association's Ethics and AI guide, issued March 17, 2026, both confirm that using AI-generated content without proper verification exposes a lawyer to discipline, and that liability for the work stays with the lawyer regardless of what tool produced the first draft.

These cases are litigation filings, but the underlying logic transfers cleanly to in-house counsel signing off on an AI-reviewed contract. The duty to verify hasn't shifted an inch. What's changed is the volume of work moving through the pipe and the speed at which it moves, and as Ryan Groff, Director of Learning at DeepJudge, puts it: "Verification is the responsibility of our profession and that has never changed."

Diagram: The Governance Gap: Adoption Far Outpaces Oversight. Visualizes: Visualize the stark gap between AI adoption and actual governance readiness across three paired statistics: (1) 88% of organizations use AI in at least one business function…

Adoption ran ahead of governance, and the gap between the two is wide enough to be the real story here. Aon data puts AI use in at least one business function at 88% of organizations in 2025. Economist Impact research found that only 8% of those same organizations have a comprehensive AI governance framework actually in place. That's most of the industry operating without the guardrails it would need to catch the failure modes already documented above. That's most of the industry operating without the guardrails it would need to catch the failure modes already documented above.

Even the organizations that claim to have governance often don't, in the sense that matters. IBM data shows 87% of organizations say they have clear AI governance frameworks, but fewer than 25% have fully implemented the controls needed to manage bias, transparency, and security risk. Claiming a framework and running one turn out to be two very different things.

In legal specifically, the ACC and Everlaw GenAI Survey found 52% of in-house counsel now actively using generative AI, more than double the 23% reported just a year before, and 84% of legal operations professionals expect GenAI's impact to be transformative. Adoption is accelerating directly into that governance vacuum, not around it.

The shift from AI experimentation to AI accountability is already under way, and organizations that skipped building a governance framework during the experimentation phase are going to find the transition considerably harder than the ones that didn't. The stakes get sharper with agentic AI, where systems move from assisting a lawyer to executing multi-step tasks on their own. Thomson Reuters is bringing CoCounsel Legal's agentic workflows, including autonomous document review, to general availability in August 2026 after a beta starting in April 2026, and LexisNexis' Protégé General AI now runs four specialized agents working in coordination. Deloitte research finds 74% of organizations plan to adopt agentic AI within two years, but only 21% have a governance model in place to manage the risks an autonomous agent introduces. That's a lot of autonomy outrunning a lot of oversight, and errors in that kind of system have more room to travel before a human ever sees them.

None of this is abstract anxiety, either. Among respondents in the 2025 Generative AI in Professional Services Report who said GenAI shouldn't be part of their daily work, 40% named accuracy and reliability as the primary reason, nearly double the next most common concern. The hallucination risk laid out earlier in this piece is an organizational gap, not a technology gap waiting on a better model release. It's an organizational gap, and governance is the tool built to close exactly that kind of gap.

Governance practices that reduce hallucination exposure in contract workflows specifically

Governance frameworks for enterprise AI adoption makes a point that applies directly here: governance works best when it's built into everyone's job, not parked in a compliance team's queue. As AI takes on more of the task volume, humans need to take on more active oversight of the outputs, not less, and Governance frameworks tend to deliver more value when senior leadership actively shapes them rather than delegating the function entirely to a technical team and move on.

A handful of concrete practices follow from everything above.

Risk-stratified human review. Map the workflow the way the earlier section did: mandatory attorney review at the high-stakes points (obligation identification, clause extraction feeding a playbook comparison, negotiation summaries going to the business), lighter spot-checks everywhere else. Treat the AI's output the way a partner treats a junior associate's first draft: useful, often good, never final without a second set of eyes.

Playbook-grounded AI with tight scope. Tools constrained to compare contract language against a defined, pre-approved playbook produce outputs that are easier to audit and less prone to invention than tools left free to generate a position from scratch. That's the real architectural fork in this space: AI that retrieves and compares against known positions, versus AI that generates and, sometimes, invents.

Audit trails that actually trace something. Every AI-assisted review should leave behind a record of what the model was shown, what it returned, and who signed off on it. The absence of a native audit trail is a core reason general-purpose LLMs aren't fit for unsupervised contract review. That same trail becomes the evidence of due diligence if a dispute over a hallucinated clause occurs months later.

Prompt governance. A sound governance framework requires that legal ops teams define what questions can be asked of an AI tool, in what context, and in what output format. Open-ended, unstructured prompting of a general-purpose model is close to the riskiest configuration available for contract work, precisely because it gives the model the most room to fill gaps on its own terms.

Vendor validation before signing anything. Don't take a vendor's hallucination claim at face value; the Stanford/Yale research found marketing language and measured performance pulling apart substantially. Ask vendors to define "hallucination" in their own documentation rather than gesturing at RAG as a blanket fix. Test tools against workloads that resemble what the team actually does, not a scrubbed benchmark dataset built to look good.

Data handling as a hard line, not a preference. Tools that train on customer contract data introduce a separate risk category entirely: confidential terms, counterparty positions, negotiation history, all becoming raw material for a model that might serve other customers later. A governance framework needs an explicit answer to what the AI sees, how long it's kept, and whether it ever feeds back into training.

Summize research found 89% of legal professionals already using AI tools, with 49% using more than one regularly. Governance built for a future rollout misses the point. It has to account for tools already embedded in daily work today.

Purpose-Built Contract AI Versus General-Purpose Tools

The divide comes down to architecture, not marketing copy. Purpose-built contract AI is built around attorney-defined rules and known, pre-approved contract positions, the model's job is to compare and flag deviations against a fixed reference point. General-purpose LLMs work the other way: they generate output probabilistically from broad training data, with no fixed reference point constraining what they produce.

That difference appears exactly where it matters most. A tool built to check a clause against a playbook has a ceiling on how far it can drift, because it's always measured against something concrete. A tool generating a legal position from scratch has no such ceiling, and the sanctions record above is a record of what happens when that lack of a ceiling goes unchecked in a filed document. Contract review sits in a stranger position than legal research, too: the source material is private, proprietary, and rarely resembles anything in a model's training set. Because of that, the retrieval side of the architecture, what the model is grounded against, carries more weight than the generation side.

None of this makes AI unsuitable for contract work. The evidence says the opposite: accuracy in the high 80s to mid 90s on clause identification is a real capability, competitive with human reviewers on the same task. But the way that capability gets deployed, what it's constrained against, what audit trail it leaves, who reviews its output and when, determines whether an organization is managing a documented, measurable risk or simply hoping it doesn't materialize on the contract that matters most.

Sources

  1. The Risks of Hallucinations and Misuse of Generative Artificial Intelligence Before French Courts
  2. AI IP Year in Review - AI Hallucinations in Court Filings and Orders: A 2025 Review of Sanctions Across the Courts and Rule Proposals | Sterne Kessler
  3. GenAI hallucinations are still pervasive in legal filings, but better lawyering is the cure | Thomson Reuters Institute
  4. AI Hallucination Legal Cases: A Sanctions Tracker (2026) — GC AI
  5. developmentcorporate.com
  6. dho.stanford.edu
  7. hai.stanford.edu
Filed underAI in Legal

More in AI in Legal