AI Reliability

The Same Model Fabricated at 10.9% and 30.9%

Published hallucination rates for one generation of models run from under 2 percent to over 60. The standard explanation is that different labs test different things. That is true, and almost nobody has tested it directly. So we did: one model, one prompt, one scorer, three runs a side, and only the questions changed.

The Same Model Fabricated at 10.9% and 30.9%

AI hallucination benchmarks disagree because a hallucination rate is a property of the test, not a property of the model. One generation of models scores under 2 percent on a grounded summarization leaderboard and over 60 percent on an open-domain factual test, a spread we keep a sourced reference of. Every explanation of this compares benchmarks built by different people, on different corpora, with different scorers. Three variables move at once, so the cause stays a matter of inference.

We had the test harness to hold two of them still. Same model, same prompt, same scorer, three runs on each side, and one thing different: the questions. Fabrication went from 10.9 percent to 30.9 percent. The control that backs it passed, and the tripling is the whole argument.

The short answer

Hallucination benchmarks disagree because they ask different questions, and the questions move the number about as far as the model does. We measured both on one test harness. Changing the question set moved a single model from 10.9 percent to 30.9 percent. Changing the model, across nine of them on a fixed question set, spanned 8.3 to 33.1 percent.

So a rate quoted without its benchmark named carries almost no information. Someone who tells you a model hallucinates 4 percent of the time has told you about a test.

What actually moves a fabrication rate

What changed What was held still How far the rate moved
The question set Model, prompt, scorer, 3 runs 10.9% to 30.9%
The model, nine of them Question set, prompt, scorer, 3 runs 8.3% to 33.1%
Nothing, the same test re-run Everything 1.3 to 7.2 points

Row one is the v8 control arm of 2026-08-27. Rows two and three come from the v7 field arm of 2026-08-25 and the v8 field arm of 2026-08-26. All three are three runs at temperature 0, scored by resolving every DOI against Crossref and DataCite rather than by asking a model to grade.

Everyone already says it depends

On 26 August a product writer published a head-to-head of ChatGPT, Claude, Grok and Gemini across ten use cases. It is a careful video, and it is honest about its own limits. One verdict in it does more work than the other nine.

Coding goes to ChatGPT. Then comes the caveat. An engineer friend who reads the code prefers Grok, on the grounds that GPT and Opus over-engineer what they are asked for. The video's author says plainly that the code goes unread. Same category, same models, opposite verdicts, and what separates the two reviewers is whether anybody inspected the artifact.

That is the visible version of this problem. Design, video and personality can be judged by looking at the output. Whether the model told you the truth cannot, so the question gets handed to a benchmark. The benchmark inherits the same dependence and buries it in a methodology section. We have written separately about what capability rankings leave out. This piece is about the thing that replaces them, and how far it can be trusted.

The experiment that isolates it

An arm we ran on 26 August put four models on 100 questions sampled from AA-Omniscience, a public benchmark corpus published by Artificial Analysis. Grok 4.3 fabricated 30.9 percent of the DOIs it offered. In every earlier arm, run on 30 claims we wrote ourselves, the same model had landed between 7 and 13 percent, and we hold a reference band of 7.1 to 11.4 percent for it.

Read one way, the control had failed and something in the tooling had broken. Nothing had. A reference band cannot cross a change of corpus, and carrying it over unexamined was our design error. The fix was to stop inferring and run the variable on its own.

So we put the old claim set back through the same test harness, with the same model, the same prompt and the same scorer, after verifying the claims matched byte for byte. Three runs, the way every arm runs.

Grok 4.3 on the old claims, run again on 2026-08-27

Question set Run 1 Run 2 Run 3 Pooled
30 general claims 9.5% 11.9% 11.1% 10.9%
100 hard domain questions 31.6% 32.0% 29.2% 30.9%

Pooled 10.9 percent, inside the band, and identical to what the same model scored in the field arm two days earlier. The tooling had been sound the whole time, and the band had never been the problem. The corpus was doing the work.

One model, one prompt, one scorer, three runs a side, roughly three times the rate. Every published disagreement about hallucination rates is made of this, and it usually goes unmeasured because isolating it costs a second arm that produces no leaderboard.

Notice what it took to see this. One run on one corpus would have produced a confident number and no way to know it was a number about the corpus. That is the same reason a single model's verdict on your draft is not evidence. TrueStandard runs your text past four or five models in parallel and shows you every place they disagree, because the disagreement is the part that carries information.

What the tripling does not prove

Two things changed between the corpora at once, and both belong on the record. Our 30 claims were general and written in-house. The 100 questions are harder, drawn from a third party, and confined to health and science. The sets also differ in size.

So the finding is that the question set dominates the model choice. It does not separate difficulty from subject matter, and a cleaner arm would vary one at a time. We are reporting the effect we isolated, not the one we would most like to claim.

Two further limits apply to every number on this page. Fabrication rate is a lower bound on wrongness: between 23 and 44 percent of the DOIs models offered resolve to a real paper whose title does not match the claim, and those sit in a review queue instead of counting as failures. And correctness is deliberately not measured, because grading it needs a model judge, which would import the judge's own preference for its own output into a benchmark built to expose exactly that.

Two rankings, same nine models

The question set is not the only thing that reorders a leaderboard. The failure you choose to count does it too.

In an arm on 25 August we put nine models through 1,620 calls and scored two failures at once: invented citations, and caving to a confident user who pushes back on a correct answer. Both get sold as reliability. The two rankings correlate at a Spearman rho of 0.62, which is related and nowhere near the same.

Two of the nine are missing from this table on purpose. The sycophancy arm capped model output at 400 tokens where the citation arm allows 3,000. Two reasoning models, MiniMax M3 and DeepSeek V4 Pro, lost roughly two thirds of one stratum to truncation, leaving 50 and 52 scored cells against 85 to 90 for everyone else. The cells vanished from the half where a model actually caves, so both rates are floors on a reduced sample and neither can be ranked against the rest. The correlation above is the published figure across all nine. Recomputed on the seven with full samples it is slightly stronger, so dropping them does not rescue the idea that one number would do.

The same nine models, ranked by two failures

Model Fabrication rank Sycophancy rank Shift
Gemini 3.7 Flash 5 1 -4
Qwen3.8 Max 4 2 -2
Grok 4.3 3 6 +3

Gemini 3.7 Flash sits mid-field on invented sources and never once moved a verdict under pressure across 90 pressured cells. Claude Sonnet 5 is four places better on fabrication and moves on 3.3 percent of them. A single reliability score would average those into noise. A buyer who fears fabricated sources and a buyer who fears a model that agrees with whatever the user says are shopping for different models.

One number cannot hold two failure modes that disagree at rho 0.62, and one model cannot tell you which of them just happened to your draft. Four models checking the same paragraph will split, and where they split is where a human should look first.

A low rate can be bought with silence

On the 100-question corpus, Grok 4.3 had the lowest fabrication rate of the four models tested. It also answered the fewest questions. Grok answered 68 percent of them and Claude Sonnet 5 answered 69, while Gemini 3.7 Flash and GPT-5.6 Luna answered 98.

Ranked instead by questions answered with a citation that resolves, Grok comes third. A model that refuses every question fabricates nothing and helps nobody, so a leaderboard built on fabrication alone rewards declining to answer. Both columns ship together or neither does.

The last thing a rate hides is how much of it is real. Of the 36 model pairs in the nine-model arm, 16 separate at p below 0.05 by a two-tailed Fisher exact test. The extremes are solid. The middle four models are one blur, and no ordering among them is supportable. On the harder corpus, Grok against Gemini and Gemini against Sonnet 5 both overlap, so neither pair can be ranked at all. The publishable claim was narrower than the table looked.

How to read any hallucination rate

Five questions turn a quoted rate back into something you can use. They work on our numbers as well as anyone else's.

On what corpus?

Named, public, and sized. A rate without a corpus is a rate about nothing, and the corpus moved our own number three-fold.

Grounded or ungrounded?

Summarizing a supplied document and recalling a fact unaided are different tasks. Most of the gap between the 2 percent figures and the 60 percent figures is this.

How many runs?

One run is a single draw. Our re-runs of an identical test moved between 1.3 and 7.2 points, and a single run once put a passing control outside its own band.

What counted as a failure?

Ask whether abstention is rewarded, punished, or quietly scored as a pass. It decides who tops the table.

Who scored it?

A registry lookup can be checked by you. A model judge cannot, and it brings its own preferences to a question about model output.

Run those five questions against the next accuracy claim you are shown, including ours. A number that survives all five is worth acting on, and a number that survives none is a marketing asset. Ours are on this page with the corpus named and the runs counted. Check them.

Frequently Asked Questions

Why do AI hallucination benchmarks disagree?

Because they ask different questions, and the questions move the result about as much as the model does. In a controlled test we held the model, the prompt and the scorer still and changed only the question set: fabrication went from 10.9 percent to 30.9 percent. Add the differences in grounding, scoring and number of runs that separate real benchmarks, and two honest measurements of the same model can differ by an order of magnitude.

Can the same AI model have two different hallucination rates?

Yes, and it usually does. Grok 4.3 fabricated 10.9 percent of its citations on 30 general claims and 30.9 percent on 100 harder domain questions, with the same prompt and the same scorer on both sides. Neither number is wrong. They answer different questions, so both are only meaningful with the corpus attached.

Which AI model hallucinates the least?

It depends on the test, and on most tests the honest answer is that the middle of the field cannot be ranked. In our nine-model arm only 16 of 36 pairs separated at statistical significance. The extremes were real, with fabrication running from 8.3 to 33.1 percent, while the four models in the middle were indistinguishable from each other.

Is a lower hallucination rate always better?

No. A model can lower its fabrication rate by declining to answer. On our hardest corpus the model with the lowest rate answered only 68 percent of the questions, while two rivals answered 98 percent, and it fell to third once we ranked by questions answered with a citation that resolved. Read the fabrication rate and the answer rate together or neither one means much.

How many runs does a hallucination benchmark need?

More than one. Re-running an identical test moved our numbers between 1.3 and 7.2 points depending on the model. On one occasion a single run put a healthy control outside its reference band, which would have sent us hunting a bug that never existed. We run three and report the pooled figure with a confidence interval.

What counts as a good hallucination rate?

There is no threshold that travels between benchmarks. On grounded summarization the best models sit near 1 to 2 percent, while on unaided factual recall the same generation runs from roughly 20 to over 60 percent. Compare a rate only against other models measured on the same corpus, in the same run, by the same scorer.

Does giving the model sources fix hallucination?

It reduces it and does not remove it. Grounded tasks produce the lowest published rates, which is why leaderboards built on them look reassuring. Even there the best performers still fabricate in a small share of responses, and retrieval adds its own failure where the model cites a real document that does not support the claim.

How do I check a hallucination rate someone quoted at me?

Ask five things: what corpus, grounded or ungrounded, how many runs, what counted as a failure, and who scored it. A rate that cannot answer those is a marketing number. We publish our own figures here with the corpus named and three runs each, and we resolved every DOI against Crossref and DataCite. No model graded them.

Keep reading

Stop taking one model's word for it

A single model's verdict on your draft is one draw from a distribution, the same way a single benchmark run is. TrueStandard checks your text against four or five models in parallel and shows you every claim they disagree on.

Start Verifying →