AI Reliability

Claude Fable 5 Tops the Hallucination Benchmark. It Does Not Hallucinate Least.

It leads AA-Omniscience with a score of 40 — the highest recorded. Artificial Analysis says that score comes from accuracy, not from low hallucination, and 9% of its answers came from a different model.

Claude Fable 5 Tops the Hallucination Benchmark. It Does Not Hallucinate Least.

Claude Fable 5 leads Artificial Analysis's knowledge and hallucination benchmark with a score of 40, the highest recorded. That does not mean it hallucinates less than other models. Artificial Analysis says so themselves: accuracy drives that score, not low hallucination.

We have not benchmarked Fable 5. Everything here about it comes from Artificial Analysis's published write-up, linked so you can check us. What we can add is our own measurement of the problem their caveat points at — one that shows up in every single-number reliability ranking, including ours.

What the benchmark actually says

AA-Omniscience asks 6,000 questions across 42 topics. It awards points for a correct answer, subtracts them for a confident wrong one, and treats an honest "I don't know" as neutral. A model earns nothing for guessing.

Fable 5 scores 40, seven points above the previous leader, Gemini 3.1 Pro Preview.

“This score is driven by leading accuracy (percentage questions correct), rather than low hallucinations.”

Artificial Analysis, on Claude Fable 5's AA-Omniscience result

A model can climb a hallucination benchmark by knowing more, not by inventing less. Those are different virtues, and only one of them is what you want when you are about to publish something it told you. The 600-question public subset is on Hugging Face if you want the raw set.

Three things the number hides

None of these are criticisms of Artificial Analysis. They publish every one of them. They are criticisms of reading one number and stopping there.

01

The index adds accuracy and hallucination together

The index rewards correct answers and penalises wrong ones inside a single score. A knowledgeable model that also fabricates can land in the same place as a cautious model that knows less. You cannot recover the split from the total.

02

Nine percent of the answers came from a different model

Artificial Analysis notes that Fable 5 falls back to Opus 4.8 on 9% of AA-Omniscience questions. Roughly one answer in eleven did not come from the model named on the leaderboard row.

03

The configuration changes the score

The same model reports different results under different reasoning-effort and fallback settings. A number quoted without its configuration is not reproducible, and reproducibility is the property that makes a benchmark a benchmark.

A single reliability score cannot tell you whether a model is safe to publish from. That is why we run every claim past four models from different vendors instead of trusting one.

The same confound in our own data

On 26 August 2026 we ran 100 questions sampled from that same public AA-Omniscience set through four fast-tier models, three times each — 1,200 answers. We did not test Fable 5. We asked a narrower question than Artificial Analysis does: when a model cites a source, does the source exist?

We checked every citation against Crossref and DataCite. A DOI either resolves or it does not, so there is no judge model and no opinion anywhere in the scoring.

Fabricated citations across three runs, 100 questions

Model Answered Fabricated 95% CI Answered with a real source
Gemini 3.7 Flash 98% 37.5% 34.2–41.0 42%
GPT-5.6 Luna 98% 54.2% 50.1–58.3 32%
Grok 4.3 68% 30.9% 27.3–34.8 30%
Claude Sonnet 5 69% 39.3% 35.0–43.7 27%

Grok 4.3 has the lowest fabrication rate in that table and comes third on usefulness. It earns the low rate by refusing: it declined roughly a third of the questions, while Gemini and Luna answered almost everything. Rank on fabrication alone and you crown the model that says least.

The refusals are explicit, not blank responses. Claude Sonnet 5 returns "I don't have verified sources for this specific claim." Grok 4.3 returns "No sources found." Both are correct behaviour, and neither appears in a fabrication percentage.

The two numbers you need

Any honest reliability figure needs a companion. One number tells you how often the model speaks; the other tells you whether what it said was real. Apart, each is easy to game.

  • How often does it answer? A model that abstains constantly looks flawless on error rate and is useless in practice.
  • When it answers, how often is that answer checkable? This is the number that decides whether you can publish without re-verifying by hand.

Artificial Analysis built this into their design by treating abstention as neutral instead of free. Vectara publishes an Answer Rate column beside its hallucination rate for the same reason. Any leaderboard without that second column rewards staying quiet.

This is why our verification runs report what each model refused to support, not only what it got wrong.

A surprise in the fabrications

Across all 1,200 answers our models emitted 975 dead DOIs, 946 of them distinct. Only two were invented by more than one model. We expected substantial overlap, on the theory that models trained on similar corpora would converge on the same plausible-looking fakes. They mostly do not. Each model invents its own.

That matters for anyone hoping to catch fabrications with a shared blocklist. There is no common set of fake citations to filter against. The overlap that did occur was on well-formed identifiers in high-prestige journals, which is the pattern worth watching.

How we measured

  • 100 questions sampled with a fixed seed from the 600-question public AA-Omniscience subset, stratified by topic. Four models, three runs each, temperature 0, no web access.
  • Health and science topics only. Our scorer resolves DOIs, and questions in law or finance are properly sourced to statutes and standards that carry no DOI, so scoring those would measure our tooling, not the model.
  • The fabrication rate is a lower bound. A further 23–44% of citations resolved to a real paper whose title did not match the claim. Those sit in a manual review queue and are not counted as failures, so real miscitation is higher than the table shows.
  • A control model re-run on our older claim set landed at 10.9%, inside its established 7.1–11.4% band, which is how we know the higher rates here come from harder questions rather than a bug in our tooling.
  • These figures are not comparable to our earlier published rates of 6–11%. Those used 30 general claims; these are 100 hard domain questions. The jump is the corpus, not an industry-wide decline.
  • We tested fast-tier models. Fable 5 is a flagship and is not in this table. Nothing here is a measurement of Fable 5.

FAQ

Does Claude Fable 5 hallucinate less than other models?

Not according to the benchmark it leads. Fable 5 scores 40 on AA-Omniscience, the highest recorded, but Artificial Analysis credits higher accuracy, not lower hallucination. The index folds both into one figure, so a top score does not establish that it fabricates less often.

What is Claude Fable 5's AA-Omniscience score?

40, which is seven points above the previous leader, Gemini 3.1 Pro Preview. Artificial Analysis also notes that Fable 5 falls back to Opus 4.8 on about 9% of the benchmark's questions, so part of that score reflects a different model's answers.

Why can a model top a hallucination benchmark without hallucinating less?

Because most reliability indexes combine accuracy and fabrication into a single number. Points earned for knowing more offset points lost for inventing things. Two models with very different failure profiles can land on the same score, and the total does not let you separate them.

Is a low hallucination rate always good?

No. A model that refuses most questions posts an excellent error rate while being useless. In our own testing the model with the lowest fabrication rate declined roughly a third of the questions and ranked third on how often it produced a source that actually existed. Read the answer rate alongside the error rate.

Can I check these numbers myself?

Yes. The Fable 5 figures come from Artificial Analysis's published article, linked above, and the 600-question corpus we sampled is public on Hugging Face. Our own run files and scoring output are committed in our benchmark directory.

Keep reading

Model Selection | 11 min read

Is There a Most Accurate AI Model?

The honest answer is no. The ranking changes with the task, the benchmark, and the month, and even the leader still hallucinates.

Model Selection | 11 min read

We Benchmarked the Models Behind Our Own Product

Six jobs in our product each pick an AI model. Exactly one of those choices had ever been tested. We built a deterministic benchmark, ran 227 graded calls against a slate of four to five candidates, and changed five of the six. Two of the tables decided nothing, one of our own test labels turned out to be wrong, and the exercise found three bugs it was not looking for.

AI Reliability | 12 min read

Every AI Tier List Measures the Same Six Things. Truth Isn't One.

A developer who built a $200k business on AI ranked 11 companies and 17 models across six categories. It is a good ranking. In 27 minutes the words hallucination, accurate and sycophancy never come up, and one model gets marked down, in so many words, for checking its work. So we ran the missing column ourselves: nine models, both failure modes, three times each. It separates, and it does not agree with any capability ranking.

AI Reliability | 12 min read

How Accurate Is ChatGPT?

Accurate enough to trust for everyday questions, and wrong often enough to get you sued if you publish it unchecked. Here is what the measurements actually say, and what to do about it.

AI Reliability | 9 min read

Does Claude Watermark My Writing?

Claude marks the words it chooses. If you wrote them, there is almost nothing to mark.

Check the claim, not the leaderboard

Benchmarks tell you how a model behaves on average across thousands of questions. They cannot tell you whether the paragraph in front of you is true. TrueStandard runs your draft past four models from different vendors and shows you every point they disagree on.

Verify a draft