Which AI Hallucinates the Least?
We measured it instead of guessing: five frontier models, 30 claims, every DOI they produced resolved against the registry that issues them. The honest answer is not the one a leaderboard wants to give.
73 of 75 times, a model asked to support a claim with no real literature refused rather than invent one. — TrueStandard benchmark v2, 2026-08-03
The short answer
No current frontier model hallucinates measurably less than the others. GPT-5.5 is nominally lowest at 11.1%, Claude Sonnet 5 highest of the current generation at 14.0% — but every pairwise difference between the top four is statistically indistinguishable, and their confidence intervals overlap almost completely.
The one clear signal in the data is generational, not competitive: the previous-generation Grok 4.3 fabricated at 24.4%, and its successor cut that to 13.3%. So the useful question is not which vendor to trust. It is what to do given that all of them fabricate at roughly the same rate — which is to check the citation, not the brand.
Fabricated-DOI rate across the four current frontier models — a spread too small to separate them at this sample size. Any list that ranks them confidently is over-reading its own data.
The measurement
30 claims, ungrounded, low reasoning effort, temperature 0 where supported. Each model was asked for up to three peer-reviewed sources with DOIs. Every DOI it produced was resolved against Crossref and doi.org. The rate is fabricated DOIs divided by DOIs the model chose to emit, so a model that declines more is not penalised for the caution.
| Model | DOIs given | Fabricated | Rate (95% CI) |
|---|---|---|---|
| GPT-5.5 | 45 | 5 | 11.1% (4.8–23.5) |
| Gemini 3.1 Pro | 43 | 5 | 11.6% (5.1–24.5) |
| Grok 4.5 | 45 | 6 | 13.3% (6.3–26.2) |
| Claude Sonnet 5 | 43 | 6 | 14.0% (6.6–27.3) |
| Grok 4.3 (previous gen) | 45 | 11 | 24.4% (14.2–38.7) |
Benchmark v2, run 2026-08-03. 150 cells, 0 errors. Fabricated = the DOI resolves in neither Crossref nor doi.org. Intervals are Wilson 95%. Cost to run: $1.14.
Read the intervals before quoting a row. At 30 claims this test can tell you that frontier models fabricate roughly one citation in eight — it cannot tell you which vendor is safest. Fisher exact tests put every pairwise comparison among the top four between p = 0.755 and p = 1.000, which is another way of saying we found no difference. Separating an 11% rate from a 14% one would need several hundred claims.
Models do not fabricate when they know nothing
Half the claims in the set were constructed for the test and have no literature behind them at all — a coefficient that does not exist, a trial that was never run. Models were given no permission to decline and no suggested wording for it. They declined anyway, on 73 of 75 occasions, in their own words: None. No peer-reviewed sources found. An empty list. This corrected our own earlier caution: a previous run had granted explicit permission to abstain, and we assumed the near-perfect refusal rate was an artifact of that permission. It was not. Where a model has no knowledge, it says so.
The danger is partial knowledge, not absent knowledge
Every fabrication in the run landed on a claim with real literature behind it. That is the failure shape that matters, because the output looks correct: the journal is right, the author is right, the year is plausible, and only the identifier is invented. A model that knows roughly which paper answers your question, but not its identifier, will construct one that fits the pattern. Nothing about the result looks wrong until you try to resolve it.
Two vendors invented the same fake identifier
On one claim about minimum-wage effects, GPT-5.5 and Grok 4.5 independently produced the identical non-existent DOI: 10.1257/aer.84.4.772. It is exactly what a correct identifier for that literature would look like — the right journal prefix, then volume 84, issue 4, page 772. Both models completed the same structural pattern and arrived at the same invention. This is worth sitting with, because it sets a real limit on naive cross-checking: two models agreeing is not corroboration when both are pattern-completing the same format. The check that works is resolving the identifier, not polling the models.
What this means if you publish
Choosing a model is not a control. The gap between the best and worst current frontier model here is three percentage points and we cannot show it is real, so switching vendors buys you nothing you can measure. What does work is mechanical and boring: resolve every DOI before it ships. A citation either exists or it does not, and that question has an objective answer that does not require trusting any model — including ours.
Cells where a model declined rather than invent a source for a claim with no literature — with no permission to decline given
Fabrications on claims where the model had no knowledge. Every single one landed on a claim with real literature behind it
Fabrications by Grok 4.3 versus Grok 4.5 on identical prompts — directionally better, but not significant at this sample size (p = 0.281)
The fix is not to hunt for a single more accurate model. Every large language model predicts fluent, plausible text, so each one can be confidently wrong on its own. What changes the odds is agreement. When several independent models are asked the same thing and all land on the same answer, the chance they share the exact same hallucination drops sharply. When they disagree, you have found the precise claim to check by hand before it ships.
Questions about this benchmark
So which AI hallucinates the least?
On this test, GPT-5.5 — but only nominally. Its 11.1% and Claude Sonnet 5's 14.0% are not distinguishable at 30 claims (p = 0.755), and the same is true of every other pair among the current four. If you want a defensible one-line answer: current frontier models fabricate roughly one citation in eight, and the differences between them are smaller than this test can resolve.
Why measure DOIs instead of asking whether the claim is true?
Because a DOI either resolves or it does not, which means no judge is needed. Every benchmark scored by a language model inherits that model's self-preference. A company whose thesis is that models cannot reliably grade themselves cannot credibly publish a benchmark graded by a model, so there is no LLM anywhere in our scoring path — only the Crossref and doi.org registries.
Were the models allowed to say they did not know?
In this run, no. An earlier version told them that declining was acceptable and preferred; this one removed that entirely to measure default behaviour. They declined anyway, 73 times out of 75. Both runs and both prompt sets are published so the difference can be checked.
How often is this updated?
On every frontier model release. Grok 4.5 shipped on 2026-07-08 and is in the table above; the previous generation stays in for comparison. The prompt set is frozen per version, so numbers published under one version remain reproducible.
What would make this benchmark better?
More claims, mainly. Thirty is enough to show that frontier models fabricate at a meaningful rate and that the effect concentrates on partially-known topics. It is not enough to rank vendors, and we would rather say so than publish a ranking the data does not support. Repeated sampling for run-to-run variance is the other gap.
Do not pick a model. Check the citation.
TrueStandard runs your draft past four models and resolves what they cite, so a fabricated source gets caught before your name is on it.