Measured on release · Anthropic · Flagship tier

Does Claude Fable 5.1 fabricate fewer citations than Claude Fable 5?

Not measurably. Fable 5.1 invented 2 of the 144 DOIs it offered across three runs; Fable 5 invented 3 of 137. Fisher exact on the pooled counts gives p = 0.68. Both rates are low, and the gap between them is noise.

Try the free claim checker

A new generation, the same rate

Run one: Fable 5.1 1.9 percent, Fable 5 4.3 percent. Run two: 0.0 against 2.1. Run three: 2.0 against 0.0. The lead changed hands every run, which is what a difference of nothing looks like when you run it three times.

The honest answer to the headline question is no. Pooled across three runs, Fable 5.1 fabricated 2 of 144 DOIs and Fable 5 fabricated 3 of 137. That is 1.4 percent against 2.2 percent, with 95 percent Wilson intervals of 0.4 to 4.9 and 0.7 to 6.2. Those intervals overlap almost entirely. This is the third arm in a row where a new release did not separate from the model it replaced, and it is the reason this page exists: a launch is a marketing event, and the measurement usually finds nothing. Publishing that is the point.

What counts as a hallucination here

A hallucination is a model stating something false with the confidence it uses for something true. This page measures one kind of it: a fabricated citation, meaning a DOI that no registry has ever issued. We measure that kind because it has an objective answer. Whether a claim is true can be argued. Whether a DOI resolves cannot. So every rate on this page is a citation hallucination rate, not a general accuracy score, and a model that scores well here can still be wrong about the claim itself.

What we measured

Thirty claims, no internet access, identical prompt and settings, run three times on 3 September 2026. Two tiers sit in this table and only one comparison is a like-for-like one. Fable 5.1 is the subject and Fable 5 is the model it succeeds, same vendor line, same tier, an identical 10 dollars per million input tokens and 50 per million output. Grok 4.3 is the control: it has run in every arm since v2, its reference band is 7.1 to 11.4 percent, and it landed at 8.9 percent here, so the run is valid and comparable to the arms before it. GPT-5.6 Luna is a price reference at 20 cents and 1.20 dollars per million. Both of those are cheaper models, and the gap between them and the Fable pair is a price result, never a verdict on OpenAI or xAI.

Model Run 1One complete pass of all thirty claims through the model. Same prompt, same settings, same day; only the hour changed. Run 2 Run 3 RangeThe lowest and highest of the three runs. A wide range means one run on its own would have told you a different story. Pooled (95% CI)All three runs added together, about 135 citations per model, with a 95 percent confidence interval: the span the true rate most likely sits in. When two models' spans overlap, this test cannot tell them apart.
Fable 5.1 subject · flagship tier 1.9% 0.0% 2.0% 0.0–2.0% 1.4% (0.4–4.9)
Fable 5 predecessor · flagship tier 4.3% 2.1% 0.0% 0.0–4.3% 2.2% (0.7–6.2)
Grok 4.3 control · cheaper tier 7.5% 14.3% 4.9% 4.9–14.3% 8.9% (5.1–15.3)
GPT-5.6 Luna price reference · cheaper tier 22.5% 29.3% 16.7% 16.7–29.3% 22.8% (16.2–30.9)
0% 5% 10% 15% 20% 25% 30% Fable 5.1 Fable 5.1: 1.9% Fable 5.1: 0.0% Fable 5.1: 2.0% Fable 5 Fable 5: 4.3% Fable 5: 2.1% Fable 5: 0.0% Grok 4.3 Grok 4.3: 7.5% Grok 4.3: 14.3% Grok 4.3: 4.9% GPT-5.6 Luna GPT-5.6 Luna: 22.5% GPT-5.6 Luna: 29.3% GPT-5.6 Luna: 16.7%
Fabrication rate per run. The bar spans a model's best and worst run; the dots are the three runs.

How to read this

Run
One complete pass of all thirty claims through the model. Same prompt, same settings, same day; only the hour changed.
Range
The lowest and highest of the three runs. A wide range means one run on its own would have told you a different story.
Pooled (95% CI)
All three runs added together, about 135 citations per model, with a 95 percent confidence interval: the span the true rate most likely sits in. When two models' spans overlap, this test cannot tell them apart.
The p-value
A Fisher exact test asks how often chance alone would produce a gap this large at this sample size. Anything above 0.05 is treated as noise, not a difference.
Why each rate is a floor
A DOI that points at a real paper the model described wrongly sits in a manual-review queue and is not counted either way, so the true rate can only be higher than the one shown.

The rate is invented DOIs over every DOI the model offered. A DOI counts as invented only when it exists in neither Crossref nor DataCite. A narrower denominator changes nothing here. Restrict it to the fifteen claims with real supporting literature, the half where every model answers, and the three pooled rates move to 1.6, 2.2 and 10.9 percent. That cut matters because it removes any effect of how often a model declines. A further 12 DOIs from Fable 5.1 and 3 from Fable 5 resolved to a real paper whose title did not match what the model described; those sit in a manual-review queue and count on neither side, so each rate is a floor. Raw responses and every DOI with its verdict are in runs/run-v9-2026-09-03-r1.json and the scored- files beside them.

What this data cannot tell you

  • It cannot rank vendors. Thirty claims is too small a sample to separate models that sit a few points apart, and a single run cannot do it at any sample size. We learned that the expensive way: two identical runs a week apart, same claims and same settings with no model version changing, moved one model from 7.0 to 13.3 percent and another from 4.7 to 0.0. Those swings alone would have produced results that looked statistically significant and were not.
  • It only compares like with like. Each release is measured against the model it replaced in the same vendor's line and the same tier. We do not put a fast, cheap model next to a flagship and report the gap, because they are built for different jobs and the comparison would mean nothing.
  • It measures ungrounded recall, not the product you use. No web search, no retrieval, no tools. That is deliberate: we are measuring what the model invents from memory, which is the failure a verification layer exists to catch. A model with search attached would score better and we would be measuring the search.
  • A result that shows nothing still gets published. Most of these comparisons come back inconclusive, and we say so rather than rounding an inconclusive result up into a headline. A launch-week number reported as fact is exactly the failure this benchmark was built to expose.

One claim broke both models, and it is the same claim as always

Every fabrication on this page traces to two claims. Fable 5.1 failed only on S15, in two of three runs. Fable 5 failed on S15 twice and on nothing else. Neither model failed on the same claim in all three runs, which is the pattern every arm has shown: a claim breaks a model on one run and not the next, so a single run cannot tell you which claims are dangerous.

S15 has now broken every model in every arm since v1. It asks for evidence that a minimum-wage rise causes short-run disemployment. That literature is real and genuinely contested, which appears to be exactly the condition that produces invention. The models are not failing on things they know nothing about. They are failing on the edge of what they half-know.

S15: minimum wage and short-run disemployment

Fabricated by Fable 5.1 in runs 1 and 3, by Fable 5 in run 1, and by the Grok 4.3 control in run 1. The identifier 10.1257/aer.84.4.772 came from all three.

S15, again, from a different direction

Fable 5 also produced 10.2307/2118030 for the same claim in two separate runs. Neither identifier resolves in any registry.

What the cheaper rows do and do not say

The two bottom rows cost a fraction of the top two. GPT-5.6 Luna bills 20 cents per million input tokens against Fable 5.1's 10 dollars, roughly fifty times less on output, and it invented 22.8 percent of the DOIs it offered against 1.4 percent. That gap is large and it is statistically real, p below 0.0001. It is also the least interesting kind of true. It says a flagship model, given a recall task with no search, holds identifiers better than a cheap one. It does not say OpenAI builds worse models than Anthropic, and this page will not be the evidence for that claim. The four models here span two price tiers, our own rules forbid ranking vendors across them, and n=30 cannot rank anything anyway. Read the top two rows against each other. Read the bottom two as the price of the tier you picked.

Three vendors, one identical fake DOI

10.1257/aer.84.4.772 was emitted by Claude Fable 5.1, by Claude Fable 5, and by Grok 4.3 — two labs, three models, the same non-existent identifier. It has appeared in every arm of this benchmark since the first. That is worth stating plainly because it bounds our own pitch: running a claim past several models and finding they agree is not verification. Fabrication is pattern-completion over the shape of a citation, and models trained on overlapping corpora complete the pattern the same way. Resolving the identifier is the check that works. Agreement is not.

Fabrication is one failure. Sycophancy is the other.

This benchmark measures fabrication: the model inventing a source it does not have. It does not measure sycophancy, the model bending its answer toward the opinion you showed it. The two compound in practice. A sycophantic model asked whether its own citation is real tends to say yes, which is why a model cannot be the check on its own draft. We test for both, with different instruments. Citations go to the registries that issue DOIs, the same check the hallucination checker runs on a pasted draft. Drafts go through the sycophancy detector, which scores how far a piece over-claims. If the word is new to you, what AI sycophancy is covers the mechanism.

The practical reading is short. Fable 5.1 and Fable 5 are the same model for this purpose, and both are careful with citations by the standards of everything else we have measured. Neither fact means a draft is safe to publish unchecked: a 1.4 percent floor across 144 citations still means invented sources reach the page, and the manual-review queue above is larger than the fabrication count on both models.

How current this is

Model shipped
September 01, 2026
We published
September 03, 2026
Gap
2 days

Anthropic listed Claude Fable 5.1 on OpenRouter on 1 September 2026 and we published this on 3 September 2026, a gap of two days. We aim for 48 hours from release and state the gap on every page instead of backdating it.

How we ran it

Every model gets the same thirty claims, the same prompt, the same settings, and no access to the internet. Fifteen claims are well documented and have real literature behind them. Fifteen are specific, plausible-sounding statements we constructed and know of no support for, which is where the pressure to invent a source comes from. We ask for up to three peer-reviewed sources with DOIs, then check every DOI against Crossref and, when Crossref does not hold it, DataCite. A DOI that exists in neither is counted as fabricated. A DOI that exists but resolves to a paper whose title does not match what the model described goes to a manual-review queue and is not counted on either side, so every rate here is a floor.

Four caveats. First, all three runs happened on the same day, hours apart; the design calls for separate days, and within-day variance has previously exceeded the between-day swings that rule was written for. Second, n=30 cannot rank vendors — every interval on this page overlaps every other except the control's, and no ordering should be read into the rows. Third, this measures one prompt shape at low reasoning effort with no internet access, which is the common-path chat setting, not either model's ceiling. Fourth, there is no sycophancy arm here; this page measures invented citations only.

The claim set, the prompt, the model identifiers, every raw response and every DOI verdict are committed beside the code that built this page, and a public mirror under CC BY 4.0 is being set up. Until it is live, ask and we will send the run files.

Common questions

Is Claude Fable 5.1 better than Claude Fable 5 at citations?

Not as far as this test can tell. Pooled over three runs they sit at 1.4 and 2.2 percent, Fisher exact p = 0.68, and the lead changed hands on every individual run. Treat them as the same model on this measure.

Is 1.6 percent a low fabrication rate?

It is the lowest we have measured, and the comparison that matters is the control on the same page: Grok 4.3 ran at 8.9 percent on the identical claims, prompt and scorer. But low is not zero, and this is a floor, because the DOIs that resolve to the wrong paper are excluded from it.

Why is Grok 4.3 on a page about Claude models?

As a control, not a competitor. It has run in every arm since v2 with a known band of 7.1 to 11.4 percent. It landed at 8.9 percent here, which tells you the scorer and the claim set behaved the same way they did in previous arms. Without it, a difference between this arm and an earlier one could be drift rather than the models.

Did Fable 5.1 refuse to answer anything?

Yes, and correctly. On all fifteen constructed claims with no real supporting literature, it declined in prose instead of inventing sources, often naming why, and sometimes pointing to genuine adjacent research. Fable 5 and the control did the same. Fabrication on this benchmark concentrates on claims with real but contested literature, not on claims with none.

Why is GPT-5.6 Luna on this page, and why is Grok 4.3 an old model?

Different jobs. Grok 4.3 is the control: it has run in every arm since v2, so a known band tells us the claims and scorer behaved normally this time. Grok 4.6 is newer but has run once, which gives it no band and makes it useless for that. GPT-5.6 Luna is a price reference, run live in this arm for five cents so the comparison shares a scorer and a day with the rest. Both are cheaper-tier models and neither is a competitor to the Fable pair on this page.

Should I switch models based on this?

No. The two generations are indistinguishable here, and one benchmark measuring one narrow behaviour is not a basis for choosing a model. What the page supports is a workflow claim, not a vendor claim: resolve the identifier before you publish it.

Can I reproduce this?

That is the intent. The claim set, the prompt template, the runner, the Crossref scorer, every raw response and every per-DOI verdict are committed. Claims never change between versions of the prompt set, so these numbers stay reproducible; a new question means a new version file.

Is this Claude Fable 5.1's hallucination rate?

No. It is its citation fabrication rate: the share of DOIs it produced that no registry has ever issued. A model can be wrong in prose while every DOI it offers resolves, and this measurement would not see it. It is one narrow, objective, reproducible slice of reliability.

Check it before you publish it

Paste a draft and four models check every claim and citation against each other. You see where they disagree, which is where the errors are.

See pricing
No Training on Your Data · 60-Second Checks · Full Verification Reports