No AI model we tested could check a citation by reading it. We gave four of them, from Anthropic, Google, OpenAI and xAI, the same short answer three ways: with no sources, with two DOIs that lead to real papers, and with two that lead nowhere. None could look anything up. All four scored the answer about three points higher out of ten for carrying sources, and none could tell the live links from the dead ones.
So if your fact-check step is pasting a draft back into ChatGPT and asking whether the sources hold up, the chatbot is grading how a citation looks.
The short answer
Across 168 paired tests, two dead DOIs raised the average score by 2.81 points out of 10. Swapping them for live ones moved it by 0.10. The judges paid for the citation and could not see whether it existed.
Average score out of 10 for the same answer, three versions
| Judge | No sources | Two live DOIs | Two dead DOIs | Dead DOIs vs none |
|---|---|---|---|---|
| Claude Sonnet 5 | 3.52 | 4.71 | 4.64 | +1.12 |
| Gemini 3.8 Flash | 5.52 | 8.43 | 8.67 | +3.14 |
| GPT-5.6 Luna | 3.38 | 7.00 | 6.95 | +3.57 |
| Grok 4.3 | 2.69 | 6.60 | 6.10 | +3.40 |
Each figure is the judge's own rating, averaged over 14 claims and 3 runs on 16 September 2026. Each judge uses its own scale, so read a row across rather than a column down.
Pooled over all four judges, the dead-DOI version beat the bare answer in 145 of the 147 cases where the score moved. Live and dead got the same score in 86 of 168 cases. Where they differed, live won 46 times and dead 36, a split a coin could produce: a sign test puts it at p = 0.32.
How we tested it
We wrote one fixed two-sentence answer for each of 14 well-documented scientific claims, such as that washing hands with soap cuts diarrheal disease in children. Each answer came in three versions that differed only in their last lines: no sources, two live DOIs or two dead ones. The sources were bare DOI links with no titles, because a title gives a judge something to reason about besides whether the link works.
Each version went to each judge in a separate call, so no judge could compare versions. The judge rated how well the evidence supported the answer from 1 to 10 and gave one sentence of reasoning. Four judges, 14 claims, three versions and three runs came to 504 calls, with no errors, for 52 cents.
Every DOI in the test came from our earlier citation benchmark. The live ones are papers that models cited correctly there. The dead ones are fabrications the same models produced, and we checked each again against Crossref and DataCite, the two registries that issue scientific DOIs, before using it. That makes them hard fakes: well-formed, in the right journal's format, and in one case a single character away from a real paper.
A 2024 paper, Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge, named this authority bias and found judges swayed by book, web and quote citations that GPT-4-Turbo invented for the study. We asked 2026 models a narrower question: shown a live identifier or a dead one of the same kind, can the judge tell which is which?
The judges vouch for dead links
Each judge also gave a one-sentence reason, and we read all 168 it gave for the dead-DOI version. In 106, the judge said something about the sources it had no way of knowing: that they were peer-reviewed, which journal or year they came from, what type of study they were, sometimes a trial's name. Four of the reasons, verbatim:
“The claim is accurate and cites real, relevant landmark trials (4S and ASCOT-LLA)”
Claude Sonnet 5, on statins
“The provided sources are a 2005 cluster RCT and a 2003 systematic review”
Grok 4.3, on handwashing
“provides DOIs to landmark randomized controlled trials that directly substantiate the claim”
Gemini 3.8 Flash, on statins
“the cited JAMA and Lancet studies provide relevant evidence”
GPT-5.6 Luna, on air pollution
Claude was looking at two dead DOIs. The first is the identifier of 4S, a real 1994 statin trial, with one character missing, so the identifier does not exist: 10.1016/0140-6736(94)90566-5, where the real one reads 10.1016/S0140-6736(94)90566-5. Follow it and you get an error page. Claude saw a 1994 identifier from The Lancet in the right shape, named the trial it resembled and called it real. It named 4S in all three runs.
A model that makes up a DOI builds it from the parts a real one has: the publisher prefix, the journal code, the year. A model grading that DOI checks for the same parts, finds them and passes it. Asking a model to check a fabricated DOI repeats the reading that produced it.
Two judges sometimes noticed. Grok 4.3 called the dead DOIs unrelated to the claim in 4 of its 42 cases and never said so of the live ones. Claude Sonnet 5 doubted it could verify the dead DOIs 14 times and the live ones 6 times, and said 4 dead pairs might be fabricated or unrelated. Its scores did not follow its words: it gave the live and dead versions 4.71 and 4.64 on average. GPT-5.6 Luna and Gemini 3.8 Flash never called a DOI fake, unrelated or unverifiable, live or dead.
Those doubts came from two labs' models, Grok and Claude, on different claims, so any single judge would have raised half of those flags or none. TrueStandard runs your draft past four models from different labs and shows you where they disagree, and its citation check opens each DOI instead of reading it.
A real source can score lower than a fake one
One answer said intermittent fasting lowers C-reactive protein, a marker of inflammation. Its first live DOI is a 2017 trial in JAMA Internal Medicine, and Gemini 3.8 Flash recognized it. In all three runs Gemini marked the answer down because, it said, that trial found no significant reduction in C-reactive protein. It scored the answer 4, 4 and 3.
Gemini had the trial right in substance: its abstract reports no significant difference in C-reactive protein between alternate-day fasting and daily calorie restriction. The same answer with two dead DOIs scored 8, 7 and 8. With nothing to check, Gemini said the literature backed the claim.
A checker that works this way penalizes a real source that complicates your claim and passes a fake one, because an invented source never disagrees with the sentence it was invented for.
This already happened in court
In 2023 two New York lawyers submitted a court filing that cited cases ChatGPT had made up. When the court questioned the citations, the lawyer who had done the research asked ChatGPT directly, typing "Is Varghese a real case" and "Are the other cases you provided fake". According to the judge's sanctions opinion in Mata v. Avianca, ChatGPT answered that it had supplied real cases that could be found through Westlaw, LexisNexis and the Federal Reporter.
The court fined the two lawyers and their firm $5,000. ChatGPT made up the cases, and when asked, it vouched for them. Our test measures that second failure, three years and several model generations later. For more cases like it, see our list of fake-citation disasters. For why self-review fails in general, see why AI can't check its own work.
This result holds from run to run
Our fabrication rates are noisy. When we re-ran the same citation test a week apart, one model's rate nearly doubled with nothing changed, which is why we won't rank vendors on a single run (why hallucination benchmarks disagree). This result did not move: across three runs, no judge's bonus for dead DOIs changed by more than 0.43 points, and every run gave a sourced answer the benefit of the doubt.
The bonus for dead DOIs, lowest and highest run
| Judge | Lowest run | Highest run |
|---|---|---|
| Claude Sonnet 5 | +1.07 | +1.21 |
| Gemini 3.8 Flash | +2.93 | +3.29 |
| GPT-5.6 Luna | +3.43 | +3.71 |
| Grok 4.3 | +3.14 | +3.57 |
What this test does not show
It does not rank the judges. Claude Sonnet 5 had the smallest bonus, but it also scored in the narrowest range, giving live-DOI answers 4.71 on average, and at telling live from dead it did no better than the others.
Grok 4.3 leaned toward the live DOIs, preferring them in 17 of the 26 cases where its score differed. At that sample size the lean can't be told apart from chance (p = 0.17), so we don't claim it.
It does not cover a judge that can browse. ChatGPT with search on can open a link, and a link that fails to load tells it something. We didn't test that. If you rely on it, check that it quotes the page it opened rather than describing the source.
It covers one answer shape on one scale: a short assertion, a rating from 1 to 10 and 14 claims with real literature behind them.
It measures a different bias from judges preferring their own answers. That one is self-preference, covered in can one AI fact-check another.
What works: resolve the identifier
A DOI leads to a paper or it doesn't, and you can find out in seconds. Paste it into the resolver at doi.org and see whether a paper loads. If one does, check that the title and authors match what your draft says, because a live DOI can still point to the wrong paper.
For a whole draft, our four-check method covers the rest. Or paste the draft into the AI citation checker, which resolves each DOI and reads the page it points to.
Across this test, a single model graded what a citation looked like, and a dead one looked fine. TrueStandard runs your draft past four models from different labs and flags every claim they can't agree on, so your checking time goes to the sentences that need it.
FAQ
Can ChatGPT check if my citations are real?
Not by reading them. In our September 2026 test, four AI models with no browsing, including OpenAI's GPT-5.6 Luna, scored an answer carrying two dead DOIs almost the same as the same answer carrying two real ones. Open each DOI yourself, or use a tool that resolves it.
Can AI tell a fake citation from a real one?
Not from the citation alone. Across 168 paired tests, four judges from four labs scored live DOIs 0.10 points higher than dead ones on a 10-point scale, no better than chance (p = 0.32). A fabricated DOI is built to look like a real one, and its look is all the judge has.
What is authority bias in LLM-as-a-judge?
It's when a language model grading an answer scores it higher because the answer cites sources, whether or not the sources hold up. A 2024 study, Justice or Prejudice?, named it. In our 2026 test, two dead DOIs raised an answer's score by 2.81 points out of 10 on average.
Why do AI judges give higher scores to answers with citations?
Citations look like evidence, and a judge without a tool can't check them, so it scores how much the answer resembles a well-sourced one. In our test, judges described the cited papers in 106 of 168 cases where the DOIs led nowhere.
Would ChatGPT catch a fake citation if it can search the web?
Possibly, if it opens the link. We didn't test browsing. A model with search can still describe a source without fetching it, so check that it quotes the page it opened. Resolving the DOI yourself takes seconds and settles it.
How do I check whether an AI-generated citation is real?
Paste the DOI into the resolver at doi.org and confirm a paper loads, then check that the title, authors and year match your draft. For a book or a court case, search the title in a library catalog or a court database. Do not ask the model that wrote the citation to confirm it.
Which AI model is best at checking citations?
Our test cannot rank them. Every judge gave dead DOIs a large bonus over no sources, and none told dead from live better than chance. A smaller bonus on a narrower scoring scale does not mean a judge is harder to fool.
Is it safe to use an LLM as a judge for fact-checking?
For checking whether sources exist, no: our judges couldn't tell live DOIs from dead ones. If your pipeline uses a model to grade answers, resolve every identifier in code before the model sees the answer, and never let a citation raise a score on its own.
Keep reading
Why AI Hallucinations Are Structural
DELEGATE 52, GPT-5.5, and a Purdue impossibility proof. Three April 2026 results that move 'hallucinations are structural' from take to documented fact.
AI Fake-Citation Disasters: A 2026 Reference
A running catalogue of documented cases where AI fabricated citations, across law, academia, and media. Every entry is sourced. It is here to be linked, cited, and updated as new cases surface.
How to Check If AI Citations Are Fake
Four checks catch a fabricated reference before your readers do. One of them is new: in 2026, a DOI that resolves no longer means the citation is real.
Can One AI Reliably Fact-Check Another AI?
If ChatGPT wrote the draft, can Claude safely verify it? Sometimes helpful, not sufficient by default — and the reason is what these models share, not what they don't.
Does Claude Watermark My Writing?
Claude marks the words it chooses. If you wrote them, there is almost nothing to mark.
Check the sources before your readers do
One model grades what a citation looks like. TrueStandard checks your draft against four models from different labs in parallel and shows you every claim they disagree on, in about 60 seconds.
Start Verifying →