AI Reliability

Which AI Fabricates Citations? We Tested Four Models

We asked four frontier models for peer-reviewed sources on 30 claims, then resolved every DOI they gave us against Crossref. 24 of 146 did not exist. The fabrication rate varies four-fold between vendors, and every single fake landed on a claim that has real literature behind it.

A fabricated citation is the hardest AI error to catch, because it does not look like an error. The journal is real, the author is real, the year is plausible. Only the identifier is invented, and almost nobody checks the identifier. So we ran the test: 30 factual claims, four frontier models, one instruction to supply peer-reviewed sources with DOIs. Then we resolved all 146 DOIs they returned against Crossref, the registry that actually issues them.

24 were fake. That is 16.4 percent of every DOI the four models offered, and the spread between the best and worst vendor is more than four-fold. This is our own data, not a survey of published benchmarks. For the sourced roundup of what other 2026 benchmarks measured, see AI hallucination rates in 2026. The prompt set, the raw responses, the grading script and every verdict are published so anyone can re-run it.

The results

Each model was asked for up to three sources per claim and was explicitly told that finding none was an acceptable answer. The fabrication rate below is the share of DOIs a model chose to emit that do not resolve, so a model that declines more offers fewer and is not penalised for the caution.

Fabricated DOIs by model, 30 claims, ungrounded

Model DOIs given Resolved Fabricated Fabrication rate
Gemini 3.1 Pro 40 35 3 7.5%
GPT-5.5 39 32 4 10.3%
Grok 4.3 33 17 7 21.2%
Claude Sonnet 5 34 19 10 29.4%

The ranking does not track general capability. These are all current flagship models, and on this one axis they are nowhere near each other. The best result is still 7.5 percent, which means roughly one in thirteen sources the strongest model handed us was invented. No model here clears the bar you would set for a research assistant.

A fifth column exists in the raw data and we deliberately do not score it: DOIs that resolve to a real paper whose title does not match what the model claimed. That category is genuinely ambiguous, and treating it as failure would inflate every number on this page. We explain why in how we tested it.

Update, 3 August 2026: a second run under default conditions — no permission to abstain, and with Grok 4.5 added — found the four current models sitting between 11.1% and 14.0%, close enough that every pairwise difference is statistically indistinguishable. The four-fold spread on this page did not reproduce, so treat the ranking above as one sample rather than a verdict. Current numbers and intervals: which AI hallucinates the least.

Given permission, every model declined

Half the claims were constructed to have no supporting literature at all. Plausible-sounding, specific, and invented for this test. A 2019 trial finding that Baroque music in the operating theatre cut infection rates by 23 percent, a Kellerman-Voss coefficient relating ceiling height to revenue per employee, a 2018 Helsinki Consensus Statement on algorithmic decision fatigue.

On all 15 of those claims, all four models declined. Sixty cells, and not one DOI was offered. Every model answered NO SOURCES FOUND, every time.

That result comes with a condition attached, and the condition matters as much as the number. The prompt told each model, in plain terms, that replying NO SOURCES FOUND was acceptable and preferred over an uncertain citation. Real prompts do not say that. So this is a ceiling, not a default: it shows what these models will do when granted permission to admit a gap, not what they do when a user simply asks for sources. Running the same test without that sentence is the obvious next arm, and until it exists nobody should read this as evidence that models decline on their own.

It does suggest something useful in the meantime, though. If a model will reliably say it has nothing when told that is allowed, then telling it that is allowed costs one sentence in your prompt.

The fabrications land where the model half-knows

All 24 fabrications occurred on the other 15 claims, the ones with substantial real literature behind them. Handwashing and diarrhoeal disease. Fine particulate matter and cardiovascular mortality. Statins. Transformer architectures. Topics with hundreds of citable papers.

Zero fabrications where the model knew nothing. Twenty-four where it knew something. The danger zone is partial knowledge, not absent knowledge, and that inverts the intuition most people work from. A model asked about an obscure topic tends to be visibly uncertain. A model asked about a well-studied one is fluent, specific, and occasionally inventing the one component nobody checks.

What a fabricated citation actually looks like

On the handwashing claim, three models returned the same real Cochrane review and the same real Lancet trial. Claude returned the correct journal, the correct year and the correct study, and a DOI that does not exist:

What Claude gave us. Resolves to nothing.
10.1016/S0140-6736(05)71199-4
The real Lancet DOI, from Gemini and Grok
10.1016/S0140-6736(05)66912-7

Same journal prefix. Same year in the suffix. Same paper being described. Four digits different, and the citation is worthless. Nothing about the first one reads as wrong. You cannot catch this by looking at it. That is why it gets published.

This is the failure a second pair of eyes catches and a re-read does not. Asking the same model whether its citation is real gets you the same confidence that produced it. The reason is structural, and we cover it in why AI can't check its own work. TrueStandard runs your draft past four models from different labs at once and surfaces the claims they disagree on, which is where the invented ones show up.

One claim broke all four models

On a single claim, every model fabricated: that minimum-wage increases have small or negligible short-run disemployment effects in some labour markets. A real, heavily studied, genuinely contested empirical literature, and all four invented an identifier rather than reach for one of the many real papers.

Four models, one claim, four different fake DOIs

Model Fabricated DOI Looks like
GPT-5.5 10.1257/aer.84.4.772 American Economic Review
Gemini 3.1 Pro 10.2307/2118030 a JSTOR archive record
Grok 4.3 10.1257/aer.94.2.138 American Economic Review
Claude Sonnet 5 10.1177/0019793918806645 ILR Review

Look at what they did not do. They did not converge on one wrong answer. Each one built a different plausible identifier from the right publisher prefix. Two reached for the same journal and still produced different numbers. That is the signature of generation rather than recall, and it is also the opening: you do not need to know the correct DOI to notice that four independent models produced four different ones. The disagreement is detectable without the answer.

For comparison, four of the fifteen well-supported claims got clean recall from every model: spaced repetition and long-term retention, statins and cardiovascular risk, CBT for insomnia, ocean acidification and shell formation. So this is not a model that cannot cite. It is a model that cites correctly most of the time and invents the rest with no change in tone.

How we tested it

The design choices matter more than usual here, because the obvious way to run this test is wrong in a way that would undermine the result.

No AI grades the output

The score is DOI resolution against Crossref, checked again on doi.org. A DOI either exists or it does not, so there is no judgement call and no rubric. We could have used a model to grade the citations, and it would have been faster and much weaker: a company arguing that models cannot reliably grade themselves cannot then publish a benchmark scored by a model. Every number here is machine-checkable by anyone.

No web search

The models ran without retrieval, browsing or grounding. That is deliberate. Turn on search and you are measuring the search tool, not the model. What we wanted was what the model produces from memory, which is what happens when someone asks for sources in a normal chat window and gets a fluent answer back.

Two strata, one honest limit

Fifteen claims with substantial real literature, fifteen constructed to have none. We do not claim the second group is provably unsupported, because you cannot prove a negative. We do not need to. The headline metric is whether an emitted DOI resolves, which does not depend on the claim being true. The strata exist to vary the pressure, not to serve as a truth label.

The category we refuse to score

Some DOIs resolve to a real paper whose title does not obviously match the claim. It is tempting to count those as failures, and we do not, because the heuristic proved unreliable in both directions. One GPT citation flagged this way resolves to a landmark air-pollution mortality study, exactly the right paper for its claim, flagged only because the model phrased the title differently. Another resolves to a paper on mollusc phylogenomics, cited in support of antibiotic-resistance gene transfer, which is a genuine miscitation. Separating those needs a human, so the tooling surfaces them and declines to grade them. Publishing that column as a failure rate would be the same sin this study exists to expose.

The frozen prompt set, the runner, the Crossref scorer, the raw responses and every per-DOI verdict are in the repository, along with the full methodology and its limitations. The whole run cost 75 cents and produced 120 responses with no errors. Changing a claim means a new version of the prompt set, never an edit to this one, so the numbers on this page stay reproducible.

What this study does not show

Four limits worth stating plainly, because a benchmark quoted without them stops being evidence.

Thirty claims is directional, not definitive

120 cells, one run, no repeated sampling, so no confidence intervals. Treat the four-fold spread as a real signal and the individual percentages as approximate.

Reasoning effort was pinned low for all four

Uniformly, which is what makes the models comparable, and because at default effort one model spent its entire token budget reasoning and returned nothing. These results describe low-effort behaviour, the common path from a chat surface, not each model's ceiling.

The permission to decline was explicit

The 60-out-of-60 abstention result depends on a prompt that told the model declining was preferred. Without that sentence the numbers could look very different, and we have not run that arm yet.

One model per vendor, one point in time

Rates move with every release. A 2026 measurement of these four checkpoints is not a permanent property of the vendors, so any citation of these numbers needs the date attached.

How to check a citation in 30 seconds

The specific failure this study documents is cheap to catch once you know it concentrates in the identifier rather than the surrounding text.

1

Resolve the DOI, not the title

Paste it into doi.org. A title is easy to generate plausibly; an identifier either resolves or it does not. Most fabrications die here, in a few seconds.

2

If there is no DOI, ask for one

A model that invented the paper will invent the identifier too, and an invented identifier is faster to disprove than an invented title. Absence of a DOI is itself a signal on anything post-2000.

3

Search the author plus the year

A recurring pattern is a real researcher attached to a paper they never wrote. Their actual publication list settles it, and this one looks correct until you check.

Doing that by hand across a full draft is the tax that makes AI-assisted research feel slower than writing from scratch. That is the job TrueStandard automates: paste the draft, four models from different labs check every claim in parallel, and you get back the specific ones they cannot corroborate in about 60 seconds. Try it on a single claim with the free claim checker, or read the method behind checking whether an AI citation is fake by hand.

Frequently Asked Questions

Which AI fabricates the most citations?

In our test, Claude Sonnet 5 had the highest fabrication rate at 29.4 percent of the DOIs it offered, followed by Grok 4.3 at 21.2 percent, GPT-5.5 at 10.3 percent and Gemini 3.1 Pro at 7.5 percent. The important caveat is that this measures one specific behaviour, ungrounded citation recall at low reasoning effort, on 30 claims in July 2026. It is not a general accuracy ranking, and the ordering here does not match how these models rank on general capability.

Does ChatGPT make up sources?

Yes, and our measurement puts a number on it: 4 of the 39 DOIs GPT-5.5 gave us did not exist, a rate of 10.3 percent. Every one of those fabrications was on a claim that has real published literature, which is the pattern that makes them hard to catch. The journal, author and year are usually right and only the identifier is invented. Independent 2026 work found a comparable effect in legal filings, where GPT-5.1 fabricated 6.57 percent of citations against 1.23 percent for GPT-4o a year earlier.

Why does AI invent DOIs instead of saying it doesn't know?

Because it is generating the citation rather than retrieving it. A DOI is a short, highly patterned string, so a model that has learned what Lancet and American Economic Review identifiers look like can produce a convincing one without having the specific paper memorised. In our test the models did this only where they had partial knowledge of the topic; on claims they knew nothing about, when explicitly told declining was acceptable, all four declined every time.

How do I check whether an AI citation is real?

Resolve the DOI at doi.org rather than searching the title, since titles are easy to generate plausibly and identifiers are not. If no DOI was given, ask for one; a model that invented the paper will invent the identifier, which is quicker to disprove. Then search the first author plus the year, because a common failure is a real researcher credited with a paper they never wrote. Those three checks take about 30 seconds and caught every fabrication in this study.

Which AI model is safest for finding sources?

Gemini 3.1 Pro had the lowest fabrication rate in our test at 7.5 percent, but that still means about one in thirteen of the sources it offered was invented, which is not a standard any publication would accept. The practical answer is not to pick a model but to verify the identifier of any citation you intend to publish, whichever model produced it. Where models from different labs disagree is a reliable pointer to the claims worth checking.

How was this benchmark run?

30 factual claims in two strata, 15 with substantial peer-reviewed literature and 15 constructed to have none, each sent to GPT-5.5, Claude Sonnet 5, Gemini 3.1 Pro and Grok 4.3 via OpenRouter with no web search or retrieval. Each model was asked for up to three sources with DOIs and told that replying NO SOURCES FOUND was acceptable and preferred. Every DOI was then resolved against the Crossref API and re-checked on doi.org. No language model grades any output. The score is identifier resolution, which is objective and reproducible. 120 responses, no errors, 75 cents.

Keep reading

One in six sources was invented. Check yours.

Across four frontier models, 24 of 146 citations did not exist, and none of them looked wrong. Paste your draft into TrueStandard and four models from different labs check every claim in parallel, flagging the ones they can't corroborate in about 60 seconds.

Check Your Draft →