AI Reliability

Does GPT-5.6 Fabricate Fewer Citations Than GPT-5.5?

We asked both generations for peer-reviewed sources on the same 30 claims, on the same day, and resolved every DOI they produced against a real registry. GPT-5.6 came out lower. Then we noticed something in the data that matters more than which model won.

A frontier model ships and the internet grades it on feel. Someone says the new one is sharper, someone says it is lazier, and almost nobody re-runs anything measurable. We keep a frozen set of 30 claims and a script that resolves every DOI a model emits against Crossref, so a release becomes a one-command data event rather than a round of impressions.

There is a second reason we ran this one. GPT-5.6 Sol has been callable on OpenRouter, the gateway our benchmark runs through, since 9 July. Our benchmark was rebuilt on 3 August and still measured GPT-5.5 as OpenAI's current model, 25 days after its successor shipped. We published a fabrication rate for a superseded model, again. The same thing happened with Grok 4.3 in our first study. Fixing it 32 days late is not a fix worth being quiet about, so it is the first thing in this post rather than a footnote.

The result

Both GPT generations, the same 30 claims, the same settings, the same scoring pass, run inside a single arm on 10 August. Nothing differs between these two rows except the model version.

GPT-5.5 vs GPT-5.6 Sol, 30 claims, ungrounded, 2026-08-10

Model DOIs given Resolved Fabricated Rate
GPT-5.5 45 39 5 11.1%
GPT-5.6 Sol 39 33 3 7.7%

GPT-5.6 offered six fewer citations and invented two fewer of them. Its fabrication rate is about a third lower. Read no further and you would conclude the new model is meaningfully safer, which is roughly what every launch-week write-up will tell you.

Fabricated means the DOI exists in neither Crossref nor DataCite. Full method, raw responses and every individual verdict: the original four-model benchmark.

Why this is not an improvement yet

Five versus three sounds like progress. A Fisher exact test on 5/45 against 3/39 returns p = 0.719. That is not a near miss. It is the number you would expect if the two models were behaving identically and we ran the test twice.

The honest sentence is: GPT-5.6 Sol fabricated fewer citations than GPT-5.5 in this run, and this test cannot show the difference is real.

When we published the Grok 4.5 result we leaned on a second argument: 4.5's failures were a strict subset of 4.3's. We checked whether the same held here. It did not — and on re-scoring in August it turned out not to have held for Grok either, so that post has been corrected too.

It does not.

Which claims each generation fabricated on

Model Fabricated on
GPT-5.5 S03, S11, S13, S15, U09
GPT-5.6 Sol S10, S11, S15

GPT-5.6 fixed three claims and broke one its predecessor got right. That is a trade, not a nested improvement, and a trade is exactly the pattern you would see from a model that had simply gotten luckier on the day. We are not reusing the Grok framing here, because the data does not support it.

This is the structural problem with release-week coverage. The sample that fits in a fast turnaround is far too small to separate models that are genuinely close, so the coverage reports noise as signal and everyone repeats it. The fix is not to stop measuring. It is to publish the interval beside the number and let the answer be boring when it is boring.

What did get better

One thing in the run is a difference in kind rather than degree, and it is the one that actually protects you.

Half our claim set is constructed: specific, plausible-sounding statements written for this benchmark with no literature behind them. The correct behaviour is to decline. On all 15 of those, GPT-5.6 emitted no DOI at all. GPT-5.5 leaked on two.

GPT-5.5: invented a source for a claim we made up
U09 (fabricated 2023 neuroimaging finding) 10.1523/JNEUROSCI.2291-22.2023 — does not exist
GPT-5.6 Sol: same 15 claims
No DOI offered on any of them. 15 / 15 declined.

Two cells out of fifteen cannot carry a significance claim either. But "stops inventing sources when it has none" is a categorically different behaviour from "invents slightly fewer," and it is the behaviour worth watching across the next few releases.

None of this changes what you have to do with the output. A 7.7% rate means roughly one in thirteen citations GPT-5.6 hands you still does not exist, and its interval runs to 20%. That is not a number you can publish against without checking. It is the entire reason we built a tool that checks every claim against four models from different labs rather than trusting any one of them.

The finding that outranks the headline

Because we kept every v2 model in this arm, the run doubles as something we had never done: a repeat of the exact same measurement. Same 30 claims, same prompt, same settings, same scorer, seven days apart, with no model version changing in between.

If the metric is stable, every row should barely move.

Identical claims, identical settings, one week apart

Model 3 Aug 10 Aug
GPT-5.5 11.1% 11.1%
Grok 4.3 11.1% 13.3%
Grok 4.5 4.4% 8.9%
Claude Sonnet 5 7.0% 13.3%
Gemini 3.1 Pro 4.7% 0.0%

GPT-5.5 reproduced exactly: 5 fabrications out of 45, both times. Claude nearly doubled, 3 to 6. Gemini went from 2 to zero. Grok 4.3 moved 5 to 6 and Grok 4.5 moved 2 to 4. Nothing about any of those models changed. Nothing about the questions changed.

The obvious objection is that this says more about our benchmark than about the models. The swing is not confined to citation tests, or to us. A talk on building production voice agents, published the same week we ran this, gives a whole chapter to model drift: a text-to-speech model that scores well one week underperforms the next, which it attributes to providers shipping updates on their own schedule.

That explanation does not cover our case. Both runs called pinned model versions through the same gateway, so there was no update to absorb. Either providers change something behind a version string that does not move, or a 30-claim sample is this noisy on its own. Neither reading lets one run stand as a measurement.

What that invalidates

This run produces a pairwise gap that clears the usual significance bar: Gemini against Claude at p = 0.027, and the same against Grok 4.3. Run the identical comparison on last week's data and it does not exist. It depends entirely on Claude sitting at a one-week high and Gemini at a one-week low on the day we happened to measure.

So we are not publishing them as vendor findings, and neither should anyone quoting this page. A p-value calculated inside a single run does not know about the variance between runs, which means it is more confident than the evidence deserves. Our own earlier caution was that 30 claims cannot rank vendors. The stronger version, which this run forced on us, is that one run cannot rank vendors at any sample size.

This is also the cleanest argument we have for why picking the safest model is the wrong move. The gap between a model's good week and its bad week was larger than the gap between most of the models.

Two generations, one invented DOI

GPT-5.6 did not inherit its predecessor's architecture wholesale, but it inherited two of its three fabrications exactly.

10.5555/3295222.3295349 · 10.2307/2118030
Both absent from Crossref and DataCite, verified 2026-08-10 and re-confirmed 2026-08-21. Emitted by GPT-5.5 and GPT-5.6 Sol alike.

A generation change did not disturb either one. These are not random slips; they are stable artifacts of how the family completes a citation pattern. The second identifier is the exact string Gemini invented for the same claim in our previous run. Three model instances, two vendors, two generations, one non-existent JSTOR record.

The claim that minimum-wage increases have small short-run disemployment effects broke all four model families again. Each invented something different and each looked correct.

What four models invented for one claim

Model Fabricated identifier
GPT-5.5 10.2307/2118030
GPT-5.6 Sol 10.2307/2118030
Grok 4.3 10.2307/2117956
Grok 4.5 10.1257/aer.84.4.772

Look at the two JSTOR identifiers: 2118030 and 2117956 are neighbours in a real range. This is pattern completion over a citation format, which is why a fabricated DOI looks so convincing and why "does it look right?" is not a check. It is also a limit worth naming on our own approach: cross-checking works because independent models fail independently, and when two of them complete the same pattern they can fail identically. What settles it is resolving the identifier, which is why our benchmark scores DOIs against a registry instead of asking a model to grade another model.

How we ran it

Thirty claims from a frozen set: 15 with substantial peer-reviewed literature behind them, 15 constructed for this benchmark with none. Each model was asked for up to three peer-reviewed sources with title, first author, year and DOI. No web search, no grounding, no retrieval: the target is what a model produces from memory, which is the failure a verification layer exists to catch.

Temperature 0 where the model accepts it, low reasoning effort, 3,000 max tokens, uniform across all six. GPT-5.5 and GPT-5.6 Sol ran inside the same arm on the same day, so nothing but the version number separates their rows.

Every DOI was extracted and resolved against the Crossref API. There is no language model anywhere in the scoring path. A benchmark about models inventing detail cannot be graded by a model without inheriting the exact bias it claims to measure.

One defect we found while writing this up: our scorer's documentation claimed it checks a second registry, but the code only queried Crossref. We wrote at the time that we had manually resolved the eight fabricated DOIs quoted in this post and all returned 404, so the numbers stood. That reasoning was wrong, and we corrected it on 2026-08-21. The spot-check covered the DOIs this post happened to quote as examples, not the full set the rates were computed from — and the full set was full of arXiv identifiers, which Crossref does not index because arXiv registers with DataCite. Real preprints were being counted as fabrications. The scorer now checks both registries and every run has been re-scored. GPT-5.5 and GPT-5.6's own rates in this post are unchanged, because neither cited arXiv on the claims they failed. Several other models in the comparison table above moved by as much as 11 points, and the leaderboard has been updated. Checking the examples you chose to write about is not checking your data.

180 calls across six models, zero errors, $1.53 in API spend. The prompt set, the raw responses and every individual verdict are committed to the repository.

FAQ

Does GPT-5.6 hallucinate less than GPT-5.5?

In our run it fabricated 3 of the 39 DOIs it offered, against GPT-5.5's 5 of 45: 7.7% versus 11.1%. But Fisher's exact test returns p = 0.719, so this test cannot show the difference is real. The clearer improvement is behavioural: GPT-5.6 declined on all 15 claims we invented, while GPT-5.5 fabricated a source for one of them.

Which AI model fabricates the fewest citations?

We do not think that question has a trustworthy answer yet, and our own data is why. Re-running the identical 30 claims a week apart moved Claude Sonnet 5 from 7.0% to 13.3% and Gemini 3.1 Pro from 4.7% to 0.0%, with no model change. The gap between one model's good week and its bad week was bigger than the gap between most of the models, so any single-run leaderboard is measuring the day as much as the model.

Why publish a result that is not statistically significant?

Because the alternative is publishing only the flattering half of what we measure. A null result is still information: it tells you a launch-week comparison at this sample size cannot answer the question people are asking. The number, the interval, the p-value and the raw responses are all published so anyone can re-run it at a larger n.

How many citations from GPT-5.6 should I expect to be fake?

Roughly one in thirteen of the ones it volunteers, based on this run, with a 95% interval running from 2.7% up to 20.3%. That interval is wide because 30 claims is a small test. Either end of it is too high to publish against unchecked.

Will you re-run this on the next model release?

Yes, and faster than this one. The claim set is frozen and the run is a single command, so every frontier release should be a data point rather than a debate. We were 32 days late here, which is the second time we have published a rate for a model its vendor had already replaced. The next priority for the benchmark is repeated sampling across several days, since that is now clearly the bigger source of error.

Keep reading

What we measured on this model

Each release page carries our own fabrication data for one model version, measured against the model it replaced in the same vendor's line.

One in thirteen citations still does not exist. Check yours.

GPT-5.6 is the best model we have measured on this test and it still invented three sources out of thirty-nine, including two its predecessor invented before it. No model release removes the need to check, and picking a different vendor does not either. Paste your draft into TrueStandard and four models from different labs check every claim in parallel, flagging what none of them can corroborate in about 60 seconds.

Check Your Draft →