Original measurement · updated 2026-08-27

Which AI Hallucinates the Least?

We measured it instead of guessing: 100 questions, four models, three runs, every DOI they produced resolved against the registries that issue them. The ranking flips depending on which column you read, and that is the answer.

The model with the lowest fabrication rate answered only 68% of the questions and finished third on how often it produced a source that exists.

— TrueStandard benchmark v8, three runs, 2026-08-26

Try the free claim checker

The short answer

The question has no single-number answer, and our latest measurement shows why. Across 100 questions and three runs, the model that fabricated least, Grok 4.3, finished third on how often it produced a source that actually exists. It earned the low rate by declining about a third of the questions. Change which column you sort on and the order changes with it.

Any reliability score built on one column can be topped by staying quiet, so read the answer rate beside the error rate. Even then only the best and worst of the four separate cleanly; the middle two overlap and cannot be ordered. The useful question was never which vendor to trust. It is what to do given that every one of them invents sources at rates in the tens of percent on hard questions, which is to check the citation, not the brand.

30.9% → 3rd

Grok 4.3 posted the lowest fabrication rate of the four and came third on what a reader actually receives, because it declined a third of the questions. One number cannot rank truthfulness.

Method, intervals and per-release measurements

The measurement

100 questions sampled from Artificial Analysis's public AA-Omniscience set, run through four fast-tier models three times each on 2026-08-26, giving 1,200 answers, no web access, temperature 0. Each model was asked for up to three peer-reviewed sources with DOIs, and every DOI it produced was resolved against Crossref and then DataCite. Two columns, not one, because a model that stays quiet can top any single-column board.

Model Answered Fabricated (95% CI) Answered with a real source
GPT-5.6 Luna 98% 32.8% (29.0–36.8) 49%
Gemini 3.7 Flash 98% 37.5% (34.2–41.0) 42%
Grok 4.3 68% 30.9% (27.3–34.8) 30%
Claude Sonnet 5 69% 39.3% (35.0–43.7) 27%

Benchmark v8, three runs on 2026-08-26. 1,200 cells, 0 errors, 0 truncations. Fabricated = the DOI exists in neither Crossref nor DataCite. Intervals are Wilson 95% over pooled runs. Cost to run: $2.13. A control model re-run on our older claim set the next day scored 10.8%, inside its established 7.1–11.4% band, so the higher rates here come from harder questions and not from our tooling. We re-scored every run behind this table on 3 September 2026, after finding a defect in our DOI extractor. See the correction below.

Frontier models are measured separately, three times each against the version they replaced: the model release pages

Read the two columns together or neither means much. Grok 4.3 has the lowest fabrication rate on this table and comes third on what you actually get, because it declined about a third of the questions while Gemini and Luna answered nearly all of them. Rank on fabrication alone and you crown the model that says least. Three further limits. The four are closer than the ordering suggests: Grok 4.3 and GPT-5.6 Luna are statistically indistinguishable, as are Gemini 3.7 Flash and Claude Sonnet 5, and neither pair can be ordered. Every rate here is pooled over three runs on the same day, and the runs still move: Gemini ranged 33.7% to 40.9% and Grok 4.3 29.2% to 32.0% across them, so read each figure as a band and not a verdict. We ran the same 100 questions three times for exactly this reason. The fabricated column is a lower bound. A further 33 to 44% of citations resolved to a real paper whose title did not match, and those sit in a review queue instead of counting as failures. And these are fast-tier models on hard health and science questions, so the rates are not comparable to the low single digits this page carries for frontier models on general claims. Correction, 3 September 2026. This page ran GPT-5.6 Luna at 54.2% fabricated and 32% useful, which put it last on both. Both numbers were wrong and the cause was ours: our DOI extractor treated the asterisks of markdown bold as part of the identifier, and Luna is the only model in this field that bolds its citations, so the scorer turned 179 real papers into inventions. Re-scored it sits at 32.8% and 49%, second on fabrication and first on the column that matters. The other three are unchanged to the decimal. This is the second scorer defect we have found and published; the first, in August, counted real arXiv preprints as fabrications. Both inflated the model that already looked worst, which is why neither was caught by reading the numbers.

Models do not fabricate when they know nothing

Half the claims in the set were constructed for the test and have no literature behind them at all — a coefficient that does not exist, a trial that was never run. Models were given no permission to decline and no suggested wording for it. They declined anyway, on 88 of 90 occasions, in their own words: None. No peer-reviewed sources found. An empty list. This is the one result that has held steady across every run we have done: 97.3% in the previous run, 97.8% in this one. It also corrected an earlier caution of ours — we had assumed the near-perfect refusal rate was an artifact of having granted permission to abstain, and it was not. Where a model has no knowledge, it says so.

The danger is partial knowledge, not absent knowledge

Twenty-three of the 24 fabrications in this run landed on a claim with real literature behind it. That is the failure shape that matters, because the output looks correct: the journal is right, the author is right, the year is plausible, and only the identifier is invented. A model that knows roughly which paper answers your question, but not its identifier, will construct one that fits the pattern. Nothing about the result looks wrong until you try to resolve it. The single exception — GPT-5.5 inventing a neuroimaging source for a finding we made up — is the rarer and more alarming failure, and its successor did not repeat it.

Two generations invented the same fake identifier

On one claim about minimum-wage effects, GPT-5.5 and GPT-5.6 Sol produced the identical non-existent DOI: 10.2307/2118030. A generation change did not disturb it. In the previous run, Gemini invented that same string for that same claim, and Grok 4.5 invented a different one that was equally plausible. These are not random slips — they are what a correct identifier for that literature would look like, and models complete the pattern rather than recall the record. This sets a real limit on naive cross-checking: two models agreeing is not corroboration when both are pattern-completing the same format. The check that works is resolving the identifier, not polling the models.

What this means if you publish

Choosing a model is not a control. We cannot show that any current frontier model is safer than another, and when we re-ran the identical test a week later the rows moved more than they differ from each other. Switching vendors buys you nothing you can measure. What does work is mechanical and boring: resolve every DOI before it ships. A citation either exists or it does not, and that question has an objective answer that does not require trusting any model — including ours.

97.8%

Cells where a model declined rather than invent a source for a claim with no literature — with no permission to decline given. Held at 97.3% in the previous run

Benchmark v3 · 88 of 90

23 of 24

Fabrications that landed on a claim with real literature behind it. Partial knowledge is the failure mode, not ignorance

Benchmark v3

5 → 3

Fabrications by GPT-5.5 versus GPT-5.6 Sol on identical prompts in one arm — directionally better, but not significant at this sample size (p = 0.719)

Benchmark v3 · like-for-like arm

The fix is not to hunt for a single more accurate model. Every large language model predicts fluent, plausible text, so each one can be confidently wrong on its own. What changes the odds is agreement. When several independent models are asked the same thing and all land on the same answer, the chance they share the exact same hallucination drops sharply. When they disagree, you have found the precise claim to check by hand before it ships.

Questions about this benchmark

So which AI hallucinates the least?

We do not think this question has a trustworthy answer yet, and our own data is why. Gemini 3.1 Pro is lowest in the most recent run at 0.0% — but the same model measured 4.7% a week earlier, and Claude Sonnet 5 went the other way, 7.0% to 13.3%, with no version change in either case. The movement between runs is bigger than the gaps between models. If you want a defensible one-line answer: current frontier models fabricate somewhere under one citation in ten, and no single run can order them.

Why measure DOIs instead of asking whether the claim is true?

Because a DOI either resolves or it does not, which means no judge is needed. Every benchmark scored by a language model inherits that model's self-preference. A company whose thesis is that models cannot reliably grade themselves cannot credibly publish a benchmark graded by a model, so there is no LLM anywhere in our scoring path — only the Crossref and DataCite registries.

Were the models allowed to say they did not know?

In this run, no. An earlier version told them that declining was acceptable and preferred; this one removed that entirely to measure default behaviour. They declined anyway, 88 times out of 90 — and 73 out of 75 in the run before it. Every run and every prompt set is published so the difference can be checked.

You changed these numbers. Why should I trust the new ones?

Because we found the error ourselves, published what it was, and kept the old files. Our scorer checked DOIs against Crossref only. arXiv registers with DataCite, so genuine preprints were being counted as fabrications and every rate we published was too high — Grok 4.3 by 11 points. The scorer now checks both registries, every run has been re-scored, and the original Crossref-only results are still in the repository so you can diff them. The conclusion this page has always led with did not change: the movement between runs is larger than the gaps between models.

What would make this benchmark better?

Repeated sampling, and it is now clearly the top priority. Running the identical claim set twice showed us that a model's rate can move ten points across a week with nothing changed, which means more claims measured once would only give us a more precise number we still could not reproduce. Three runs on three days would tell us more than one run of ninety claims. After that, more claims.

Do not pick a model. Check the citation.

TrueStandard runs your draft past four models and resolves what they cite, so a fabricated source gets caught before your name is on it.

See pricing
No Training on Your Data · 60-Second Checks · Full Verification Reports