AI accuracy by field · Legal

Is AI accurate for contract review?

Not on its own. On the first benchmark built to test it, the strongest frontier models reached only about 0.64 F1 at identifying clause-level legal risk in commercial contracts, roughly the level of a junior legal assistant. They read plausible-sounding contracts and miss the negation, the carve-out, or the clause that is simply not there. AI is genuinely useful for a first pass and for drafting language, but nothing it flags, or fails to flag, is safe to sign, negotiate, or advise on until a person has checked every clause against the actual document.

Try the free claim checker

Fine for a first pass, not accurate enough to sign on

In ContractEval, the first benchmark built specifically to measure clause-level legal risk identification in commercial contracts, the best models tested, GPT-4.1 and GPT-4.1 mini, reached an F1 of only about 0.64. The paper's own summary is that most large language models perform at the level of a junior legal assistant: useful for a first read, not a substitute for review. An F1 in that range means the model both misses genuinely risky clauses and flags benign ones as risky, often enough that you cannot act on its judgment without checking it.

Older work points the same way. On CUAD, the expert-annotated contract-review dataset of 510 commercial contracts and 13,000-plus lawyer-labeled clauses, the best model reached 44 percent precision when tuned to catch 80 percent of the important clauses, meaning most of what it highlighted was a false positive and it still missed one in five. Newer stress tests are harsher still: when researchers planted subtle flaws in real contracts, models detected omitted or missing terms with F1 scores between roughly 9 and 32 percent, the worst category by far. Contract review is close to the worst possible task for a model that predicts plausible text, because a contract turns on the exact words, and the most expensive error is a clause that is not there at all.

0.64 F1

the best score any frontier model reached at identifying clause-level legal risk in commercial contracts, on the first benchmark built for the task. The paper's own read: most models perform at the level of a junior legal assistant.

Liu et al., ContractEval (arXiv), 2025

Why AI fails at contract review specifically

A large language model does not parse a contract the way a lawyer does. It predicts the next most plausible token, and contract language is dense with the small words that reverse meaning: not, except, notwithstanding, subject to, provided that, solely to the extent. A model reading a clause that says a party shall not indemnify can smoothly summarize it as an indemnity obligation, because the fluent, confident version of the sentence is the one it has seen most. Defined terms, cross-references, and carve-outs compound this. A term defined narrowly in section 1 and used in section 12 can be summarized with its everyday meaning, and the reader never sees the gap.

The more dangerous failure is what the model does not say. Contract review is largely about absence: the indemnity that has no cap, the missing limitation of liability, the termination-for-convenience right that only the other side has, the confidentiality carve-out that was quietly dropped. Detecting what is missing is the single category models handle worst, with detection rates in the single-to-low-double digits on the omission tests above. Models also invent the opposite, asserting that a lopsided clause is standard or market when it is not, or describing an obligation the contract never contains. Retrieval-backed and fine-tuned tools narrow these gaps but do not close them, because the model can still surface the wrong span and then overstate what it found.

When people trusted the output anyway

Moffatt v. Air Canada (2024) — an AI's invented term became binding

Air Canada's website chatbot told a grieving customer he could claim a bereavement fare retroactively, a policy that did not exist and contradicted the airline's actual terms. British Columbia's Civil Resolution Tribunal held Air Canada liable for negligent misrepresentation and flatly rejected the argument that the chatbot was a separate legal entity responsible for its own words. The company was bound by what its AI stated, a direct warning that a model's confident account of a term carries real legal weight.

FTC v. DoNotPay (2024-2025) — the robot lawyer that never tested itself

DoNotPay marketed the world's first robot lawyer that could generate perfectly valid legal documents and substitute for an attorney. The FTC found the company never tested whether its AI produced work equal to a human lawyer and employed no attorneys to check its law-related features. DoNotPay settled for 193,000 dollars and agreed to stop claiming its service performs like a real lawyer without evidence to back it up.

Cursor's support bot (2025) — a fabricated policy people acted on

Cursor's AI support agent, presenting as a person named Sam, told users that a single-device sign-in limit was company policy. No such policy existed; the model invented it. Developers canceled subscriptions before a co-founder confirmed the rule was never real and blamed AI hallucination. The same mechanism, a model stating a rule or term with total confidence and no basis, is exactly what makes an AI contract summary unsafe to rely on unchecked.

The numbers behind it

44%

precision the best model reached on the expert-annotated CUAD contract-review benchmark when tuned to catch 80 percent of the important clauses. Most of what it highlighted was a false positive, and it still missed one in five.

Hendrycks et al., CUAD (NeurIPS), 2021

9 to 32%

F1 range for detecting omitted or missing terms on the Better Call CLAUSE benchmark of flaws planted in real contracts, the worst-handled category. The leading model matched the correct governing law to a flaw only 13.7 percent of the time.

Choudhury et al., Better Call CLAUSE (arXiv), 2025

The pattern under every one of these numbers is the same: a fluent summary, a confident risk call, and no independent check of whether the clause actually says that, or exists at all. That is the gap TrueStandard is built to close. Paste the clause, the redline, or the risk memo, and four to five frontier models review it against each other in about a minute, so a misread negation or a missing indemnity cap surfaces as a disagreement to resolve before you sign, not after.

How to verify AI contract-review output

Treat anything an AI gives you as an unverified first pass. Before you rely on it, redline from it, or advise on it, run this check against the actual document.

  1. 01

    Read every clause the AI describes against the real contract text. Confirm the words it attributes to a clause are actually there, in that section, not a fluent paraphrase that drifted from the original.

  2. 02

    Check the small reversing words. Verify each not, except, notwithstanding, subject to, and provided that, because a single one flips an obligation, and models routinely summarize the confident version instead of the literal one.

  3. 03

    Confirm defined terms against their definitions. A term defined narrowly earlier in the document can be summarized with its everyday meaning. Follow it to section 1 or the definitions block.

  4. 04

    Resolve every cross-reference. When a clause points to another section or schedule, open it and confirm the AI actually followed the chain rather than guessing what it says.

  5. 05

    Check for what is missing, not only what is present. Ask specifically about absent protections: liability cap, indemnity limits, termination for convenience, governing law, assignment. Missing terms are the category models catch least often.

  6. 06

    Do not trust standard or market labels. Verify any claim that a clause is typical against your own playbook or precedent, since models assert market-standard confidently and wrongly.

  7. 07

    Never sign, send, or advise on an AI redline or risk memo that a person has not checked clause by clause against the source. The duty of care applies to you regardless of the tool that drafted it.

How to make AI output reliable: check it across models

The fix is not to hunt for a single more accurate model. Every large language model predicts fluent, plausible text, so each one can be confidently wrong on its own. What changes the odds is agreement. When several independent models are asked the same thing and all land on the same answer, the chance they share the exact same hallucination drops sharply. When they disagree, you have found the precise claim to check by hand before it ships.

That is what TrueStandard does: it runs your draft through four to five frontier models at once and surfaces every disagreement in about a minute, with sources. See the AI fact checker for how the method works, or read why AI cites studies that do not exist for the mechanism behind the failures on this page.

Common questions

Can I use ChatGPT to review a contract?

For a first read, a plain-language explanation, and a list of points to look into, yes. As the basis for signing, negotiating, or advising, no. On the benchmark built for this task the strongest models scored only about 0.64 F1 at identifying clause-level risk, so every clause it summarizes and every risk it flags or misses has to be checked against the actual document before you rely on it.

Are dedicated contract-review AI tools more accurate than a general chatbot?

Somewhat, but not accurate enough to skip verification. Purpose-built tools use retrieval and fine-tuning over contract data, which helps, yet even the best models on the CUAD contract-review benchmark reached only 44 percent precision at 80 percent recall, and specialized stress tests still show detection of missing terms in the single-to-low-double digits. The failure mode does not disappear, it just gets quieter.

What is the most dangerous kind of contract-review error AI makes?

The clause that is not there. Contract risk is often about absence, a missing liability cap or a one-sided termination right, and detecting omissions is the single category models handle worst. Right behind it is the misread negation, where a model turns shall not into an obligation to do the thing. Both look completely fine in a fluent summary, which is why they slip through.

What is the safest way to use AI for contract review?

Use it to speed up the first pass and to draft language, then verify every clause, every negation, and every omission independently before it leaves your desk. Running the same contract through several independent models and checking where they disagree catches the misreads and missing terms that any single model, including a specialized one, will confidently pass over.

Do not publish AI output on trust

Paste your draft. Four to five models check every claim in about 60 seconds, and you see exactly where they disagree before your name is on it.

See pricing
No Training on Your Data · 60-Second Checks · Full Verification Reports