A model release is normally an occasion for vibes. Someone posts that the new one feels sharper, someone else posts that it feels worse, and nobody re-runs anything. We keep a frozen set of 30 claims, and a script that resolves each DOI a model gives us against Crossref. So a release is a one-command data event, not an opinion.
This one had a second motive. Grok 4.5 shipped on 8 July, our four-model benchmark ran on 30 July, and it tested Grok 4.3. So we published a fabrication rate for a model that had already been replaced. This study's whole value is being checkable, so that is a real defect. The fix: run both generations side by side, in the same arm, on the same day.
The result
We ran both Grok generations on the same 30 claims, same settings, same scoring pass. Each emitted exactly 45 DOIs, which makes the comparison very clean.
Grok 4.3 vs Grok 4.5, 30 claims, ungrounded, 2026-08-03
| Model | DOIs given | Resolved | Fabricated | Rate |
|---|---|---|---|---|
| Grok 4.3 | 45 | 28 | 5 | 11.1% |
| Grok 4.5 | 45 | 43 | 2 | 4.4% |
Fabrications fell from 5 to 2, and resolved citations rose from 28 to 43. The manual-review queue went from 12 to zero; that queue holds DOIs that exist, but whose titles do not match what the model claimed. On every column, 4.5 looks better than 4.3.
Fabricated means the DOI exists in neither Crossref nor DataCite. Corrected 2026-08-21: this post originally reported 24.4% and 13.3%, scored against Crossref alone. But arXiv registers with DataCite, so real preprints were being counted as fabrications. The scorer now checks both, and each run has been re-scored. Full method, raw responses and each verdict: the original four-model benchmark.
Why we are not calling this an improvement
Five versus two sounds decisive. It is not. A Fisher exact test on 5/45 against 2/45 returns p = 0.434, which is well outside any usual threshold. With 30 claims, a difference this size turns up often enough, even if the two models are identical.
Here is the honest sentence. Grok 4.5's point estimate is well under half its predecessor's, and this test cannot confirm the difference is real.
We could have written the other headline. "Grok 4.5 halves fabrication rate" is true as arithmetic, and it would travel further, too. It would also be the exact failure this benchmark exists to catch: a confident number with nothing underneath it. We would deserve to be quoted back at ourselves the first time someone checked.
This is the structural problem with model-release coverage. A launch-week turnaround only fits small sample sizes, far too small to separate models that are really close. So the coverage reports noise as signal. The fix is not to stop measuring, but to publish the interval next to the number. Let it be boring when it is boring.
The claim we had to withdraw
This section first argued that the shape of the change was more convincing than the rate. The reason: 4.5's failures were a strict subset of 4.3's. On the corrected scoring that is no longer true, and the argument does not survive it.
Correctly scored, Grok 4.3 fabricated on three claims (S07, S12 and S15), and Grok 4.5 on two (S11 and S15). It fixed S07 and S12, and broke S11, which its predecessor got right. That is a trade, not a nested improvement, and just what a model that had merely gotten luckier would give. The first version of this post said 4.5 broke nothing. That was an artifact of our scorer, which counted a real arXiv preprint as invented. We were wrong.
What fixing a claim looked like
Take ocean acidification and shell formation in marine calcifiers. On that claim, 4.3 gave one real DOI and two invented ones, while 4.5 gave three that all resolve.
The same pattern held on vitamin D and respiratory infection. There, 4.3 gave two fabrications, while 4.5 gave three clean, resolving sources. Those two fixes are real, and they survive the re-scoring. What does not survive is the claim that nothing broke in the other direction. Then there is S11, where 4.5 invented an identifier its predecessor got right. Set against that, here is the fair summary: the newer model is probably better, but the evidence for it is weaker than we first said.
None of this changes the practical advice. A 4.4% fabrication rate means roughly one in twenty-three citations Grok 4.5 gives you does not exist. That is better than one in nine, but still not a number you can publish against without checking. And you cannot tell which one is the fake by looking at it.
Two models, one invented DOI
The most interesting thing in the run was not about Grok at all. One claim says minimum-wage increases have small short-run disemployment effects, and it broke all four models in our first study. On that claim, GPT-5.5 and Grok 4.5 each gave the same fabricated identifier, on their own.
Look at the structure: it is the American Economic Review prefix, then volume 84, issue 4, page 772. That is just what a real identifier for that literature would look like. Two models from different vendors completed the same pattern and landed on the same invention.
This is a limit worth naming on our own approach. Cross-checking works because independent models fail independently. But here both were completing the same citation format from the same widely-documented literature. They can fail identically, so agreement between models is not corroboration in that case. What settles it is resolving the identifier. That is why our benchmark scores DOIs against a registry rather than asking a model to grade another model. The pattern turned out to be more stable than we expected. We re-ran this claim set for the GPT-5.6 release, and both GPT generations invented the same two identifiers as each other. One of them was a fake Gemini had already produced here.
How we ran it
Thirty claims from a frozen set: 15 with substantial peer-reviewed literature, 15 constructed for the benchmark with none. Each model was asked for up to three peer-reviewed sources, each needing a title, first author, year and DOI. No web search, no grounding, no retrieval. The target is what the model produces from memory, the failure a verification layer exists to catch.
Temperature 0, low reasoning effort, 3,000 max tokens, uniform across both models. Both Grok generations ran inside the same arm on the same day. So nothing but the model version differs between the two rows.
Each DOI was extracted and resolved against the Crossref API. When Crossref does not hold the record, the check falls through to DataCite. There is no language model anywhere in the scoring path. This benchmark is about models fabricating detail: grade it with a model and it inherits exactly the bias it claims to measure. The DataCite half of that check was missing until 2026-08-21, which is why the numbers in this post changed.
150 calls across five models, zero errors, $1.14 in API spend. The prompt set, the raw responses and each verdict are committed to the repository.
FAQ
Is Grok 4.5 better than Grok 4.3 at citations?
Probably, but this test cannot prove it. Fabrications fell from 5 to 2 out of 45 DOIs, at p = 0.434, which is not significant at 30 claims. It fixed the two claims its predecessor broke worst, and broke one its predecessor got right. That is a trade, not a clean improvement.
How does Grok 4.5 compare to GPT-5.5, Claude and Gemini?
It sits in a pack that is statistically indistinguishable. In the same arm GPT-5.5 measured 8.9%, Claude Sonnet 5 7.0%, Gemini 3.1 Pro 4.7% and Grok 4.5 4.4%. The intervals overlap throughout, and there is no safest current model to pick. A later re-run of the identical test moved several of these rows by more than they differ from each other.
Why publish a result that is not significant?
Because the alternative is publishing only the flattering half of what we measure. And a null result is still information: it tells you that a launch-week comparison at this sample size cannot answer the question people are asking. The number, the interval and the raw data are all published. Anyone can re-run it at a larger n.
Will you re-run this on the next model release?
Yes. The claim set is frozen and the run is one command. So each frontier release is a data point, not a discussion. The comparison table lives on our benchmark page, and we update it in place.
Keep reading
Is There a Most Accurate AI Model?
The honest answer is no. The ranking changes with the task, the benchmark, and the month, and even the leader still hallucinates.
Claude Fable 5 Tops the Hallucination Benchmark. It Does Not Hallucinate Least.
It leads AA-Omniscience with a score of 40 — the highest recorded. Artificial Analysis says that score comes from accuracy, not from low hallucination. And 9% of its answers came from a different model.
Does Claude Hallucinate Less at Higher Effort?
We asked Sonnet 5 and Opus 5.5 for sources on the same 30 claims at three effort levels each. Then we checked every DOI they gave us against the real registries.
Does GPT Hallucinate Less at Higher Reasoning Effort?
We asked GPT-5.6 Luna for sources on the same 30 claims at all five of its effort levels. Then we checked every DOI it gave us against the real registries.
AI Detectors Ask the Wrong Question
They flag honest writers. They mark famous human documents as AI, and OpenAI quietly killed its own detector. But the deeper problem is not that detection is unreliable. It is that 'was this written by AI?' was never the question that protects you.
What we measured on this model
Each release page carries our own fabrication data for one model version, measured against the model it replaced in the same vendor's line.
One in eight citations still does not exist. Check yours.
Grok 4.5 fabricated 6 of the 45 DOIs it gave us. The current frontier models all land within a few points of each other. There is no model you can pick that removes the need to check. Paste your draft into TrueStandard. Four models from different labs check each claim in parallel, and flag what they cannot corroborate in about 60 seconds.
Check Your Draft →