Does Claude Haiku 5.5 fabricate fewer citations than Haiku 4.5?
Yes, on this test. We gave both models the same thirty claims, three times each. Haiku 5.5 invented 7.8 percent of the DOIs it offered, and Haiku 4.5 invented 19.0 percent. A DOI is the permanent ID that links a citation to its paper, so an invented one points at nothing. If the two models were equally good, chance would open a gap this wide about 2 times in 100 (p = 0.017). Haiku 5.5 still invented one ID in thirteen, and it gave the same fake ID for one trial in all three runs.
Lower in all three runs
Haiku 5.5 invented 6.7 percent of its DOIs, then 9.5, then 7.1. Haiku 4.5 ran the same thirty claims on the same day and invented 22.2 percent, then 22.2, then 12.1.
Pooled, that is 7.8 percent against 19.0. A Fisher exact test gives p = 0.017. That test asks how often chance alone would open a gap this size, and the answer is rarely. No one run was enough to show it alone, but all three pointed the same way, so the pooled result is the finding. The gap grows when we also count real IDs attached to the wrong paper. By that count, Haiku 4.5 got 45.7 percent of its citations wrong, and Haiku 5.5 got 15.5 percent wrong (p < 0.001).
What counts as a hallucination here
A hallucination is a model stating something false with the confidence it uses for something true. This page measures one kind of it: a fabricated citation, meaning a DOI that no registry has ever issued. We measure that kind because it has an objective answer. Whether a claim is true can be argued. Whether a DOI resolves cannot. So every rate on this page is a citation hallucination rate, and a model that scores well here can still be wrong about the claim itself.
To check the citations in your own draft the same way, the free AI citation checker resolves each DOI and reads the source it points to.
What we measured
Thirty claims, no internet access, the same prompt and settings, run three times on 8 October 2026. Haiku 5.5 is the subject. Haiku 4.5 is the model it replaces, measured again here on the same day. There was no Haiku 5. Grok 4.3 is a control that has run in every test since the second one. It sits in another line and is never compared with the pair. It scored 10.6 percent here, inside its usual band.
| Model | Run 1One complete pass of all thirty claims through the model. Same prompt, same settings, same day; only the hour changed. | Run 2 | Run 3 | RangeThe lowest and highest of the three runs. A wide range means one run on its own would have told you a different story. | Pooled (95% CI)All three runs added together, about 135 citations per model, with a 95 percent confidence interval: the span the true rate most likely sits in. When two models' spans overlap, this test cannot tell them apart. |
|---|---|---|---|---|---|
| Haiku 5.5 subject | 6.7% | 9.5% | 7.1% | 6.7–9.5% | 7.8% (4.3–13.7) |
| Haiku 4.5 predecessor | 22.2% | 22.2% | 12.1% | 12.1–22.2% | 19.0% (12.7–27.6) |
| Grok 4.3 control | 13.3% | 13.3% | 4.8% | 4.8–13.3% | 10.6% (6.4–17.0) |
How to read this
- Run
- One complete pass of all thirty claims through the model. Same prompt, same settings, same day; only the hour changed.
- Range
- The lowest and highest of the three runs. A wide range means one run on its own would have told you a different story.
- Pooled (95% CI)
- All three runs added together, about 135 citations per model, with a 95 percent confidence interval: the span the true rate most likely sits in. When two models' spans overlap, this test cannot tell them apart.
- The p-value
- A Fisher exact test asks how often chance alone would produce a gap this large at this sample size. Anything above 0.05 is treated as noise.
- Why each rate is a floor
- A DOI that points at a real paper the model described wrongly sits in a manual-review queue and is not counted either way, so the true rate can only be higher than the one shown.
The rate divides invented DOIs by DOIs offered. A DOI counts as invented only when no registry has issued it. Pooled over three runs, Haiku 5.5 invented 10 of 129, with a 95 percent interval of 4.3 to 13.7 percent. Haiku 4.5 invented 20 of 105, from 12.7 to 27.6 percent, and Grok 4.3 invented 14 of 132, from 6.4 to 17.0 percent. Fisher exact on the pair gives p = 0.017, and run by run it gives 0.054, 0.21 and 0.69. Another 42 DOIs from the pair led to a real paper whose title did not match what the model wrote. We graded each one against the paper's registered first author and year, then read each one by hand. Haiku 5.5 had 10 miscitations and Haiku 4.5 had 28. Counting both kinds of error, Haiku 5.5 got 20 of 129 wrong and Haiku 4.5 got 48 of 105. The whole test cost $0.44.
What this data cannot tell you
- It cannot rank vendors. Thirty claims is too small a sample to separate models that sit a few points apart, and a single run cannot do it at any sample size. Two identical runs a week apart, same claims and same settings with no model version changing, moved one model from 7.0 to 13.3 percent and another from 4.7 to 0.0. Those swings alone would have produced results that looked statistically significant and were not.
- It only compares like with like. Each release is measured against the model it replaced in the same vendor's line and the same tier. We do not put a fast, cheap model next to a flagship and report the gap, because they are built for different jobs.
- It measures ungrounded recall, not the product you use. No web search, no retrieval, no tools. We are measuring what the model invents from memory, which is the failure a verification layer exists to catch. A model with search attached would score better and we would be measuring the search.
- A result that shows nothing still gets published. Most of these comparisons come back inconclusive, and we say so.
Ten fake IDs, and one it gave three times
The thirty claims split into two groups. Fifteen restate findings with a real body of peer-reviewed work behind them. Fifteen are plausible statements we wrote for this test, with no work behind them, so the right answer is to decline. Haiku 5.5 declined all of them. Once, it then named a nearby real paper and said plainly that it did not support the claim. That ID resolves.
All ten of Haiku 5.5's fakes came on real claims, attached to real, well-known papers. The model knew the paper but not the ID, and it wrote one anyway. Haiku 4.5 did the same thing twice as often, and it attached real IDs to the wrong paper 28 times.
In all three runs, Haiku 5.5 cited Murdoch's VIDARIS vitamin D trial against 10.1001/jama.2016.8284. No registry has issued that ID. The real trial ran in JAMA in 2012, not 2016. In run 3 it also renamed the trial VITAL. A fake that comes back on every run at the same setting is built into the model, so asking again will not catch it.
Haiku 5.5 cited Card and Krueger's 1994 study twice as 10.2307/2118030, and once as 10.1257/aer.84.4.772. Neither exists. The first string has now turned up in thirteen of our tests. Haiku 4.5 gave it twice as well.
Haiku 4.5 cited Riemann's review of insomnia against 10.1016/S0140-6736(17)30129-0. That ID is real, but it belongs to a Lancet paper on influenza. A reader who checks only that the link opens would pass it.
One real paper, three made-up IDs
Gazeau's 2007 paper on shellfish and ocean acid is real, and its ID is 10.1029/2006GL028554. Haiku 5.5 gave it 10.1029/2007GL030427. In our last test, Opus 5 gave it an ID one digit away from the real one. When two models name the same paper, the paper is probably real. Whether its ID exists is a separate question, and a registry lookup is the only way to answer it.
Fabrication is one failure. Sycophancy is the other.
This benchmark measures fabrication: the model inventing a source it does not have. It does not measure sycophancy, the model bending its answer toward the opinion you showed it. A sycophantic model asked whether its own citation is real tends to say yes, which is why a model cannot be the check on its own draft. We test for both, with different instruments. Citations go to the registries that issue DOIs, the same check the hallucination checker runs on a pasted draft. Drafts go through the sycophancy detector, which scores how far a piece over-claims. If the word is new to you, what AI sycophancy is covers the mechanism.
Haiku 5.5 makes up fewer IDs than Haiku 4.5, and one in thirteen is still made up. Some of its fakes come back on every run, so asking twice will not catch them. TrueStandard does the checking. Paste a draft, and four models from different labs check each claim and citation against each other and against the registries that issue DOIs.
How current this is
Anthropic released Claude Haiku 5.5 on 7 October 2026. We measured it and put this page up on 8 October, one day later. That is inside the 48-hour target we set ourselves.
How we ran it
Every model gets the same thirty claims, the same prompt, the same settings, and no access to the internet. Fifteen claims are well documented and have real literature behind them. Fifteen are specific, plausible-sounding statements we constructed and know of no support for, which is where the pressure to invent a source comes from. We ask for up to three peer-reviewed sources with DOIs, then check every DOI against Crossref and, when Crossref does not hold it, DataCite. A DOI that exists in neither is counted as fabricated. A DOI that exists but resolves to a paper whose title does not match what the model described goes to a manual-review queue and is not counted on either side, so every rate here is a floor.
Three caveats. First, all three runs happened on one day. Earlier tests have shown a model's rate move between days, which is why we measure the old model again next to the new one. A second day for this pair is the open work on this page. Second, the gap is large, but thirty claims still give wide intervals. Haiku 5.5's true rate could sit anywhere from about 4 to 14 percent. Third, both models ran at low reasoning effort with no web access. That is close to the setting most chat users get. Our scorer checks DOIs against Crossref and DataCite. We then checked all 37 dead IDs from this test at doi.org, which covers every agency. None of them resolve.
The claim set, the prompt, the model identifiers, every raw response and every DOI verdict are committed beside the code that built this page, and a public mirror under CC BY 4.0 is being set up. Until it is live, ask and we will send the run files.
Related reading
Common questions
Does Claude Haiku 5.5 hallucinate citations?
Yes, less often than Haiku 4.5. With no web access, it invented 7.8 percent of the DOIs it offered across three runs. Haiku 4.5 invented 19.0 percent on the same claims that day (p = 0.017).
Is Claude Haiku 5.5 better than Haiku 4.5 at citations?
On this test, yes: it was lower in every run. Some of its IDs were made up, and some were real but pointed at the wrong paper. Counting both, Haiku 5.5 got 15.5 percent wrong and Haiku 4.5 got 45.7 percent. We have one day of data, so a second day is still to come.
Can I trust Haiku 5.5's citations without checking them?
No, because about one ID in thirteen was invented. It gave the same fake ID for one trial in all three runs, so asking again does not help. Check each DOI against a registry before you cite it.
Why compare it to Haiku 4.5 and not to Sonnet 5.5 or GPT?
Thirty claims cannot rank labs, and Sonnet is a different tier. We measure each release against the model it replaced, in the same line and tier. Haiku 4.5 is that model, since there was no Haiku 5. Sonnet 5.5 has its own page.
Why does another benchmark give Haiku 5.5 a different hallucination rate?
Because it measures something else. Most hallucination scores count wrong answers to questions. Ours counts citation IDs that no registry has issued. A DOI either resolves or it does not, so anyone can check our number without trusting a model to grade it.
Can I reproduce this?
Yes. We keep every input and output: the claims, the prompt, the model IDs, all three runs of raw answers, every DOI verdict and the graded miscitation file. The test cost $0.44. Ask and we will send the run files.
Is this the Claude Haiku 5.5 hallucination rate?
It is its citation hallucination rate: the share of DOIs it gave that no registry has issued, with no web access. A general accuracy score would need a different test. We measure citations because the answer is a fact anyone can check.
Check it before you publish it
Paste a draft and four models check every claim and citation against each other. You see where they disagree, which is where the errors are.