A model release is normally an occasion for vibes. Someone posts that the new one feels sharper, someone else posts that it feels worse, and nobody re-runs anything. We keep a frozen set of 30 claims and a script that resolves every DOI a model produces against Crossref, so a release is a one-command data event instead of an opinion.
This one had a second motive. Grok 4.5 shipped on 8 July. Our four-model benchmark ran on 30 July and tested Grok 4.3. We published a fabrication rate for a model that had already been superseded. That is a real defect in a study whose whole value is being checkable, so the fix was to run both generations side by side, in the same arm, on the same day.
The result
Both Grok generations, same 30 claims, same settings, same scoring pass. Each emitted exactly 45 DOIs, which makes the comparison unusually clean.
Grok 4.3 vs Grok 4.5, 30 claims, ungrounded, 2026-08-03
| Model | DOIs given | Resolved | Fabricated | Rate |
|---|---|---|---|---|
| Grok 4.3 | 45 | 28 | 5 | 11.1% |
| Grok 4.5 | 45 | 43 | 2 | 4.4% |
Fabrications fell from 5 to 2. Resolved citations rose from 28 to 43. The manual-review queue (DOIs that exist but whose titles do not match what the model claimed) went from 12 to zero. On every column, 4.5 looks better than 4.3.
Fabricated means the DOI exists in neither Crossref nor DataCite. Corrected 2026-08-21: this post originally reported 24.4% and 13.3%, scored against Crossref alone. arXiv registers with DataCite, so real preprints were being counted as fabrications. The scorer now checks both and every run has been re-scored. Full method, raw responses and every verdict: the original four-model benchmark.
Why we are not calling this an improvement
Five versus two sounds decisive. It is not. A Fisher exact test on 5/45 against 2/45 returns p = 0.434, well outside any conventional threshold. With 30 claims, a difference this size is what you would expect to see reasonably often even if the two models were identical.
The honest sentence is: Grok 4.5's point estimate is well under half its predecessor's, and this test cannot confirm the difference is real.
We could have written the other headline. "Grok 4.5 halves fabrication rate" is true as arithmetic and would travel further. It would also be the exact failure this benchmark exists to catch: a confident number with nothing underneath it. We would deserve to be quoted back at ourselves the first time someone checked.
This is the structural problem with model-release coverage generally. The sample sizes that fit in a launch-week turnaround are far too small to separate models that are genuinely close, so the coverage reports noise as signal. The fix is not to stop measuring. It is to publish the interval next to the number and let it be boring when it is boring.
The claim we had to withdraw
This section originally argued that the shape of the change was more convincing than the rate, because 4.5's failures were a strict subset of 4.3's. On the corrected scoring that is no longer true, and the argument does not survive it.
Correctly scored, Grok 4.3 fabricated on three claims: S07, S12 and S15. Grok 4.5 fabricated on two: S11 and S15. It fixed S07 and S12 — and broke S11, which its predecessor got right. That is a trade, not a nested improvement, and a trade is exactly what a model that had merely gotten luckier would produce. The original version of this post said 4.5 broke nothing. That was an artifact of our scorer counting a real arXiv preprint as invented, and we were wrong.
What fixing a claim looked like
On ocean acidification and shell formation in marine calcifiers, 4.3 gave one real DOI and two invented ones. 4.5 gave three that all resolve.
The same pattern held on vitamin D and respiratory infection, where 4.3 produced two fabrications and 4.5 produced three clean, resolving sources. Those two fixes are real and they survive the re-scoring. What does not survive is the claim that nothing broke in the other direction. Set against S11, where 4.5 invented an identifier its predecessor got right, the fair summary is that the newer model is probably better and the evidence for it is weaker than we first said.
None of this changes the practical advice. A 4.4% fabrication rate means roughly one in twenty-three citations Grok 4.5 gives you does not exist. That is better than one in nine. It is not a number you can publish against without checking, and you cannot tell which one is the fake by looking at it.
Two models, one invented DOI
The most interesting thing in the run was not about Grok specifically. On the claim that minimum-wage increases have small short-run disemployment effects, the one that broke all four models in our first study, GPT-5.5 and Grok 4.5 independently produced the same fabricated identifier.
Look at the structure. It is the American Economic Review prefix, then volume 84, issue 4, page 772: precisely what a real identifier for that literature would look like. Two models from different vendors completed the same pattern and landed on the same invention.
This is a limit worth naming on our own approach. Cross-checking works because independent models fail independently, but when both are completing the same citation format from the same widely-documented literature, they can fail identically. Agreement between models is not corroboration in that case. What settles it is resolving the identifier, which is why our benchmark scores DOIs against a registry rather than asking a model to grade another model. The pattern turned out to be more stable than we expected: when we re-ran this claim set for the GPT-5.6 release, both GPT generations invented the same two identifiers as each other, and one of them was a fake Gemini had already produced here.
How we ran it
Thirty claims from a frozen set: 15 with substantial peer-reviewed literature, 15 constructed for the benchmark with none. Each model was asked for up to three peer-reviewed sources with title, first author, year and DOI. No web search, no grounding, no retrieval: the target is what the model produces from memory, which is the failure a verification layer exists to catch.
Temperature 0, low reasoning effort, 3,000 max tokens, uniform across both models. Both Grok generations ran inside the same arm on the same day, so nothing but the model version differs between the two rows.
Every DOI was extracted and resolved against the Crossref API, falling through to DataCite when Crossref does not hold the record. There is no language model anywhere in the scoring path. A benchmark about models fabricating detail cannot be graded by a model without inheriting exactly the bias it claims to measure. The DataCite half of that check was missing until 2026-08-21, which is why the numbers in this post changed.
150 calls across five models, zero errors, $1.14 in API spend. The prompt set, the raw responses and every individual verdict are committed to the repository.
FAQ
Is Grok 4.5 better than Grok 4.3 at citations?
Probably, but this test cannot prove it. Fabrications fell from 5 to 2 out of 45 DOIs, at p = 0.434, not significant at 30 claims. It fixed the two claims its predecessor broke worst and broke one its predecessor got right, which is a trade rather than a clean improvement.
How does Grok 4.5 compare to GPT-5.5, Claude and Gemini?
It sits in a pack that is statistically indistinguishable. In the same arm GPT-5.5 measured 8.9%, Claude Sonnet 5 7.0%, Gemini 3.1 Pro 4.7% and Grok 4.5 4.4%, with overlapping intervals throughout. There is no safest current model to pick, and a later re-run of the identical test moved several of these rows by more than they differ from each other.
Why publish a result that is not significant?
Because the alternative is publishing only the flattering half of what we measure, and because a null result is still information: it tells you that a launch-week comparison at this sample size cannot answer the question people are asking. The number, the interval and the raw data are all published so anyone can re-run it at a larger n.
Will you re-run this on the next model release?
Yes. The claim set is frozen and the run is one command, so every frontier release is a data point rather than a discussion. The comparison table lives on our benchmark page and gets updated in place.
Keep reading
Is There a Most Accurate AI Model?
The honest answer is no. The ranking changes with the task, the benchmark, and the month, and even the leader still hallucinates.
ChatGPT vs Claude vs Gemini: Which Is Most Accurate?
The three big plans cost the same twenty dollars, and every comparison of them reaches a verdict on accuracy without running a test. We ran one on all three, twice, a week apart. The winner changed.
We Benchmarked the Models Behind Our Own Product
Six jobs in our product each pick an AI model. Exactly one of those choices had ever been tested. We built a deterministic benchmark, ran 227 graded calls against a slate of four to five candidates, and changed five of the six. Two of the tables decided nothing, one of our own test labels turned out to be wrong, and the exercise found three bugs it was not looking for.
How Accurate Is ChatGPT?
Accurate enough to trust for everyday questions, and wrong often enough to get you sued if you publish it unchecked. Here is what the measurements actually say, and what to do about it.
AI Detector vs Fact Checker
One asks who wrote this. The other asks is this true. Before you publish, only one of those questions protects your reputation — and most teams are watching the wrong one.
One in eight citations still does not exist. Check yours.
Grok 4.5 fabricated 6 of the 45 DOIs it gave us, and the current frontier models all land within a few points of each other. There is no model you can pick that removes the need to check. Paste your draft into TrueStandard and four models from different labs check every claim in parallel, flagging what they cannot corroborate in about 60 seconds.
Check Your Draft →