AI accuracy by field · Legal

Is AI accurate for legal research?

Not on its own. General-purpose AI models fabricate or misstate the law on most legal queries, and even the specialized tools lawyers pay for hallucinate on up to a third of them. AI is useful for drafting and orientation, but nothing it produces is safe to file or cite until a human has checked every case, holding, and quote against the primary source.

Try the free claim checker

Accurate enough to draft, nowhere near accurate enough to cite

In the largest study of its kind, general-purpose large language models produced hallucinated legal answers between 58 percent (GPT-4, the best) and 88 percent (the worst) of the time on specific, verifiable questions about real cases. A hallucination here means the model either invented law that does not exist or cited a real source that does not actually support the claim.

The problem is not limited to consumer chatbots. When Stanford researchers tested the purpose-built, retrieval-backed tools sold by LexisNexis and Thomson Reuters, those tools still hallucinated on 17 to 33 percent of queries, despite being marketed as hallucination-free. Legal research is one of the worst possible tasks for a model that predicts plausible text, because a fabricated citation looks exactly like a real one.

58 to 88%

of legal queries put to general-purpose AI models produced a hallucination: fabricated law, or a real citation that does not support the claim.

Dahl et al., Journal of Legal Analysis (Stanford), 2024

Why AI fails at legal research specifically

A large language model does not retrieve cases. It predicts the next most plausible token, and legal citations are highly structured and pattern-rich, so the model generates citations that look completely real: a plausible case name, a reporter volume, a page number, a court, a year. None of it is checked against an actual corpus of decisions. This is why fabricated citations pass the eye test even for experienced lawyers.

The subtler failure is worse. The model often cites a real case but misstates what it held, or attaches a fabricated quotation to a genuine opinion. Verifying that a case exists is not enough, because the case can be real and the proposition still invented. Models also cannot reliably tell when they have hallucinated, and they routinely accept a user's incorrect legal premise instead of correcting it. Even retrieval-augmented legal tools fail this way, because retrieval can surface the wrong document and the model still overstates what it found.

When lawyers trusted it anyway

Mata v. Avianca (2023) — the landmark

Two New York lawyers filed a brief citing six cases that ChatGPT invented. When asked, ChatGPT insisted the fake cases were real and could be found on Westlaw and Lexis. Judge P. Kevin Castel fined the lawyers and their firm 5,000 dollars and called one AI-generated analysis gibberish.

Wadsworth v. Walmart (2025) — a national firm sanctioned

Attorneys at Morgan and Morgan filed motions citing nine cases, eight of which did not exist, generated by the firm's own internal AI tool. A federal judge revoked one lawyer's admission in the case and fined three attorneys.

An Oregon federal case (2025) — the largest penalty to date

In a dispute involving briefs with fifteen references to nonexistent cases and eight fabricated quotations, a magistrate judge imposed roughly 110,000 dollars in fees and fines, reported as the largest AI-hallucination penalty in US legal history.

The numbers behind it

17 to 33%

hallucination rate for the purpose-built legal AI tools from LexisNexis and Thomson Reuters, the ones marketed as hallucination-free and sold to lawyers.

Magesh et al., Stanford HAI / RegLab, 2024

1,400+

court filings worldwide have now been caught containing AI-fabricated citations, in a database that grows daily.

Charlotin, AI Hallucination Cases Database, 2026

The pattern in every one of these cases is the same: a confident answer, a citation that looked real, and no independent check before it was filed. That is exactly the gap TrueStandard is built to close. Paste the passage, and four to five models check every claim and citation against each other in about a minute, so the fabrications surface before your name is on the brief, not after.

How to verify AI legal-research output

Treat anything an AI gives you as an unverified draft. Before you rely on it, cite it, or file it, run this check.

  1. 01

    Confirm every cited case actually exists. Pull each one independently on Westlaw, Lexis, or a free source like CourtListener or Google Scholar by its real citation. If you cannot find it, it is fabricated.

  2. 02

    Read the actual opinion, not the AI summary. Verify that the case genuinely stands for the proposition the AI attributes to it. A real case cited for a made-up holding is the most common subtle failure.

  3. 03

    Verify every direct quotation at its cited page. Models invent verbatim quotes that do not appear in the opinion.

  4. 04

    Check the case is still good law. Run it through a citator such as KeyCite or Shepard's to confirm it has not been overruled, reversed, vacated, or superseded.

  5. 05

    Confirm jurisdiction and precedential weight. Make sure the authority is actually binding, or at least persuasive, in your forum. Models routinely mix jurisdictions.

  6. 06

    Verify statutes and regulations against the current official code, since training data is stale and provisions get amended or repealed.

  7. 07

    Never file, quote, or advise on AI-generated legal content that a human has not source-checked end to end. Your duties of competence and candor to the court apply regardless of the tool.

How to make AI output reliable: check it across models

The fix is not to hunt for a single more accurate model. Every large language model predicts fluent, plausible text, so each one can be confidently wrong on its own. What changes the odds is agreement. When several independent models are asked the same thing and all land on the same answer, the chance they share the exact same hallucination drops sharply. When they disagree, you have found the precise claim to check by hand before it ships.

That is what TrueStandard does: it runs your draft through four to five frontier models at once and surfaces every disagreement in about a minute, with sources. See the AI fact checker for how the method works, or read why AI cites studies that do not exist for the mechanism behind the failures on this page.

Common questions

Can I use ChatGPT for legal research at all?

For orientation, drafting language, and getting a rough map of an area, yes. As a source of citable authority, no. Studies put the hallucination rate for general-purpose models on verifiable legal questions between 58 and 88 percent, so every case, holding, and quote it produces has to be verified against the primary source before you rely on it.

Are the dedicated legal AI tools like Lexis+ AI and Westlaw more reliable?

More reliable, but not reliable enough to skip verification. Stanford's evaluation found the leading purpose-built tools still hallucinated on 17 to 33 percent of queries, even though they use retrieval over real case databases and are marketed as hallucination-free.

Why does AI make up cases that sound so real?

Because it learned the format of citations, not the underlying set of real decisions. It predicts a plausible-looking case name, reporter, and page number the same way it predicts any other text, so fabricated citations look identical to genuine ones.

What is the safest way to use AI for legal work?

Use it to draft and explore, then verify every factual and legal claim independently before it leaves your desk. Running the output through several independent models and checking where they disagree catches fabrications that any single model, including a legal one, will miss.

Do not publish AI output on trust

Paste your draft. Four to five models check every claim in about 60 seconds, and you see exactly where they disagree before your name is on it.

See pricing
No Training on Your Data · 60-Second Checks · Full Verification Reports