AI Reliability

When AI Cites Studies That Don't Exist

AI does not just get facts wrong. It invents whole sources: cases, studies, DOIs. Then it cites them with the same confidence it uses for real ones. Here is why it happens, and the disasters it has already caused. And here is how to catch a fake citation before your name is on it.

When AI Cites Studies That Don't Exist

Yes, AI often cites sources that do not exist. Language models invent studies, court cases, and DOIs that look completely real. Then they hand them to you in the same confident voice they use for real ones. Independent tests put the fake-source rate anywhere from 18 percent to more than 50 percent, depending on the model. Verify every citation before you publish it.

This is not a rare glitch that a better model will iron out. It follows directly from how these systems generate text. It has already sanctioned lawyers, slipped past peer review, and forced newsrooms into mass corrections. This guide walks through the recorded cases and the mechanism underneath them. It covers the subtler failure that most citation checkers miss. And it gives you a practical way to catch a fake source before it ships with your name on it.

The disasters: when invented sources got published

The fastest way to understand the problem is to look at the people it caught. They were professionals with every reason to check. They trusted the output because it looked right.

The original example is Mata v. Avianca. In 2023, New York attorney Steven Schwartz used ChatGPT to write a brief. It cited six court cases that did not exist, including a made-up Varghese v. China Southern Airlines. ChatGPT even assured him the cases could be found in Westlaw and LexisNexis. Judge P. Kevin Castel sanctioned the lawyers and their firm 5,000 dollars. The case became a warning headline worldwide (opinion and order). It was not a one-off from a careless solo lawyer. In April 2026, Sullivan and Cromwell admitted to a federal bankruptcy judge that one of its filings contained AI-invented citations. This is one of the most prestigious firms in the world. There were more than 40 errors in all, including quotes attributed to the court that were never said. Opposing counsel caught it, not the firm's own review (reported by Bloomberg Law and CNN, April 2026).

These are not one-off headlines. Researcher Damien Charlotin keeps a public database of court decisions involving AI-hallucinated content. By mid-2026 it had logged more than 1,600 cases worldwide, and the count climbs almost daily. And the problem is not confined to law. A Columbia University School of Nursing audit was published in The Lancet in May 2026. It verified 97 million references across 2.5 million biomedical papers. It found 4,046 fake citations, sources that do not exist in any scientific database. They sat across 2,810 published, peer-reviewed papers (Columbia). A separate Nature investigation the same year had its own estimate. More than 110,000 papers published in 2025 may carry invalid, AI-generated references (Nature).

A few of the recorded cases

Case What the AI invented What it cost Year
Mata v. Avianca Six court cases, including a fictitious Varghese v. China Southern Airlines $5,000 sanction and worldwide headlines 2023
Sullivan & Cromwell filing ~28 fabricated or misquoted citations in a court brief Public apology to a federal judge 2026
Columbia / The Lancet audit 4,046 fake references across 2,810 papers Invented sources inside peer-reviewed science 2026
Nature investigation An estimated 110,000+ papers with invalid AI references Polluted scientific literature 2026

The through-line is not carelessness. A fake citation looks exactly like a real one. So normal review slides right over it. For the full running list across law, academia, and media, see our reference of AI fake-citation disasters.

Why a language model invents a citation

To catch the failure reliably, it helps to know why it happens. The reason tells you exactly where to look.

A language model does not pull a citation from a database. It generates the most probable next token, one after another, from patterns in its training data. A citation has a highly predictable shape: an author, a year, a journal, a volume, a page range. So the model can produce a flawlessly formatted reference with nothing real behind it. The format is easy to fake. The object it points to is not something the model ever checked. That is the whole trick, and it traces back to the structural cause of AI hallucination.

The obvious fix is to give the model a real database to draw from. It helps. It does not solve the problem. A 2024 Stanford study tested the leading grounded legal research tools, the ones sold to end hallucination. Lexis+ AI and Westlaw's AI-Assisted Research still invented between 17 and 33 percent of the time. Westlaw's tool was wrong about a third of the time (Stanford HAI). Retrieval narrows the gap. It does not close it.

The raw rates are worse than most people assume. In controlled studies, ChatGPT-3.5 invented roughly half of its sources. That was 47 percent in one Cureus analysis and 55 percent in Scientific Reports. GPT-4 cut that to about 18 percent, which is real progress. But among the citations that were genuine, it still got substantive details wrong about a quarter of the time (Walters and Wilder). And newer does not mean safe. A 2026 benchmark of legal citations found GPT-5.1 invented sources at 6.57 percent. That is up from 1.23 percent for the GPT-4o release a year earlier (Who Checks the Citations?). The number moves with the model and the topic. It never reaches zero.

The harder failure: real source, wrong claim

Most people picture one kind of fake citation: a source that simply does not exist. That one is easy enough to catch. The more dangerous failure is a real source, cited correctly, that does not say what the AI claims it says.

The same benchmark that counted fake legal citations sorts the failures into five distinct types. Only the first is caught by checking whether a link resolves. The model can point you to a genuine paper that argues the opposite of the claim. It can attach a real, working DOI that opens an unrelated study. It can quote a source that never contained the quote. Each of these passes the shallow test, does the link work. Each is still wrong.

Five ways an AI citation fails

Failure type What it looks like Caught by a link check?
Non-existent source The paper or case does not exist at all Yes
Wrong source A real, working DOI that opens an unrelated paper No
Wrong location Real source, but the wrong page or section No
Misquote A quote attributed to a real source that never said it No
Misrepresentation Real source that does not support the claim No

This is not a rare edge case. Researchers checked ChatGPT-5's citations against the papers it was citing. Roughly a third of its claims carried errors against the real, correctly cited source (Journal of Technology in Behavioral Science, 2026). A separate study looked at fake citations that came with a DOI. Of those, 64 percent resolved to a real but unrelated published paper. Someone clicking the link lands on a real article (EurekAlert release on the JMIR Mental Health study, 2025). This is why a citation checker that only confirms the URL loads gives you false confidence. The link working is not the question. Whether the source says what the AI says it says is the question.

Are peer reviewers using AI to flag papers for hallucinated references?

Yes, and the failure runs in both directions. In April 2026, a researcher posted on r/LanguageTechnology about a peer reviewer. The reviewer had used an LLM to flag real PubMed-indexed citations in their paper as fake. The reviewer comment: "Seems to be a hallucinated reference, duplicate or erroneous references," followed by a list of supposedly faked citations. The references were real. (r/LanguageTechnology thread) A preprint at the same time documents 17 similar cases. In each, LLM-based peer-review systems wrongly accused real PubMed citations of being fake. (clawrxiv preprint)

This is the loop that makes single-vendor self-checking untrustworthy. The same model architecture that produces hallucinated references is sometimes used to check references. It then produces hallucinated rejections of real ones. The error mode is symmetric.

What did the Lancet study on fabricated medical references actually find?

On May 7, 2026, a peer-reviewed letter in The Lancet reported a first. It was the first large-scale, methodologically grounded estimate of AI-driven fake citations in scientific literature. An audit of 2.5 million biomedical papers found a 12-fold rise in fake sources since 2023. About 3,000 published medical papers were found to contain fake citations. (Nature coverage · STAT coverage · EurekAlert summary)

Two implications matter for anyone publishing under their own name.

The rise tracks the LLM-assisted writing timeline

The likely cause is researchers drafting with AI and citing what the model returns, with no independent check. The 12-fold growth curve maps onto the rollout curve of ChatGPT, Claude, and Gemini.

Automated screening will produce false positives

The same letter notes that screening will produce false positives. The way out is to use a different method from the one that introduced the errors. Otherwise the peer-review false-flag pattern repeats at industrial scale.

For non-medical writers, the implication is the same. When you publish a claim, the cost of an unchecked AI-generated source is now measurable. Retroactive screening is becoming standard. The error is no longer your private secret.

Why asking the AI to check itself fails

Almost everyone reaches for the same instinct: ask the same AI whether its citation is real. It does not work. Knowing why points straight at what does.

A model asked to check its own output runs the same probability distribution that produced the citation. Say it invented a source because the pattern looked right. Asking is this reference real returns the same confident yes, for the same reason. Nothing independent has entered the loop. Re-pasting into a fresh window of the same model is weaker than it feels, too. Same architecture, same training, same blind spots. Real checking needs independence. It needs a source of judgment that was trained differently and does not share the original error. This is the core of why AI cannot check its own work.

This is the exact gap TrueStandard is built to close. Instead of trusting one model to grade itself, it runs your draft across four to five frontier models from different labs at once. Say only one model knows a source, because only that model hallucinated it. The others fail to back it up, and that disagreement is the flag. Independence is the mechanism, not more careful reading.

How to catch a fake citation before you publish

You do not need to become a forensic librarian. You need a habit and, for anything you publish at volume, a faster version of it.

Start with the mindset shift. Treat every AI-produced citation as a claim to verify, not a fact you already have. Then go to the primary source and ask two questions, not one. Does this source exist? And does it actually say what the draft claims? Checking only the first, does the link resolve, is what lets misrepresented and misquoted sources through. For a single high-stakes citation, tracing it by hand is the gold standard, and worth the few minutes. The problem is that hand-checking every citation in every draft under a deadline does not scale. That is why most people skip it and hope.

The efficient version keeps the independence and drops the manual grind. Ask several models trained by different labs the same question. Then read where they disagree. A source one model invented cannot be backed up by models that never shared the hallucination. So a fake citation shows up as a split rather than a consensus. When independent models land together, that agreement is a real signal the source is sound. When they split, they have pointed you at the exact citation a human needs to confirm. You are not left rereading the whole draft hunting for the one that is wrong.

Ways to check an AI citation, compared

Method What it catches Time cost
Ask the same AI to check itself Little; it shares the original blind spot Seconds
Click the link to see if it resolves Only fully-invented sources; misses misrepresentation Seconds
Trace each citation to its primary source by hand Nearly everything, if you have the time High, per citation
Cross-check across independent models Fabrications, plus the exact claims to verify About 60 seconds

That last row is what TrueStandard does for you. Paste your draft. Four to five frontier models from different vendors check every claim and citation in parallel. In about 60 seconds you get back the ones they cannot agree on, ranked by how far apart they are. Want the manual method first? Our guide to checking whether AI citations are fake walks through it step by step. And how accurate is ChatGPT puts the wider error rate in context.

Frequently Asked Questions

Does ChatGPT make up citations and sources?

Yes, routinely. A language model generates the most probable text rather than pulling a real record. So it can produce a perfectly formatted citation, with author, journal, year and DOI, and nothing real behind it. In controlled studies, fake-source rates ran from about 18 percent for GPT-4 to over 50 percent for GPT-3.5. Even retrieval-grounded tools still invented 17 to 33 percent of the time. Treat any AI citation as unproven until you check it.

Why does AI invent citations?

A citation has a very predictable structure. So it is easy for a model to generate a convincing one from pattern alone. The model was never checking a database. It was predicting what a plausible source looks like. The format is trivial to fake, and the model never consulted the underlying source. That is why the source can look flawless and still point to nothing.

How common are fake AI citations?

Common enough to assume it will happen. Lab studies measured fake-source rates of roughly 47 to 55 percent for GPT-3.5 and about 18 percent for GPT-4. A 2024 Stanford study found grounded legal AI tools still invented 17 to 33 percent of the time. In the wild, a 2026 audit found fake citations in nearly 3,000 peer-reviewed medical papers. So the problem is reaching published work, not just draft output.

If the citation link works, is the source real?

Not necessarily. A working link only tells you a page exists, not that it supports the claim. In one 2025 study, 64 percent of fake citations that included a DOI resolved to a real but unrelated paper. The link opened a real article that had nothing to do with the claim. You have to check whether the source actually says what the AI says it says, not just whether it loads.

Can AI check whether its own citations are real?

No, not reliably. Asking a model to check its own output runs the same process that produced the error. So it tends to confirm a fake citation with the same confidence it invented it. Re-pasting into a fresh window of the same model has the same blind spot. Reliable checking needs independence. That means a human tracing the primary source, or several models trained by different labs cross-checking each other.

How do I check if an AI citation is real?

Go to the primary source and ask two things. Does it exist? And does it actually support the claim? For a single important citation, trace it by hand. For a whole draft, run the claims across several independent models and look for disagreement. A fake source cannot be backed up by models that did not share the hallucination. That points you at exactly which citation to confirm.

Is it safe to use AI for academic research if I check the citations myself?

Yes, if checking means two things. Verify each citation against its primary source, not just that it exists. And verify the quoted content matches what is in the source. Most researchers check only that the citation appears to exist. The Lancet study finding suggests that version has stopped being enough. Multi-vendor verification cuts this to seconds per claim.

Which AI hallucinates the fewest citations?

It varies by model and topic. And newer does not reliably mean safer. One 2026 benchmark found a later GPT model invented legal citations more often than an earlier one. Picking a single best model is not a fix. Whichever one you choose still invents on the specific claims that matter most. The durable answer is to verify with independence rather than to trust one model.

Keep reading

Is AI accurate for your field?

The failure modes change by profession. These break down what AI gets wrong in specific fields, with the incidents and the checks that catch them.

Catch the Fake Citation Before Your Readers Do.

AI invents sources that look identical to real ones. A working link does not mean the source is real. Paste your draft into TrueStandard. Four to five frontier models check every claim and citation in about 60 seconds. They flag the ones they cannot back up, before your readers or your editor find them first.

Check Your Draft →