AI Reliability

AI Hallucination Rates in 2026: What the Data Actually Shows

A sourced reference of the 2026 hallucination-rate numbers: what each benchmark measured, which models did best and worst, and where the widely quoted figures get misread. Built to be linked and kept current as new data lands.

How often does AI hallucinate? It depends on what you ask it to do. On deliberately hard factual questions, hallucination rates across 26 top models range from 22 to 94 percent, per the Stanford AI Index 2026. Hand the model the source text and ask it to summarize, and the best performer still fabricates in 13.6 percent of responses. In professional tools, grounded legal research products hallucinated on 17 to 33 percent of queries in Stanford's testing. There is no single hallucination rate, and anyone quoting one number without naming the benchmark is telling you less than they appear to.

This page collects the 2026 numbers that hold up, each with the benchmark it came from and what that benchmark actually measured. The short version: rates are high where models must recall facts unaided, lower but not low when they are handed sources, and worst on exactly the material published work depends on, specific statistics, citations, and niche detail. For the mechanism behind the failure, see why AI hallucinates. This is the data.

The headline number, read correctly

The most quoted hallucination statistic of 2026 comes from the Stanford AI Index, and most people quoting it have not read what it measures.

The Stanford HAI 2026 AI Index reports, in its Responsible AI chapter: hallucination rates across 26 top models range from 22 to 94 percent on a new accuracy benchmark (Stanford HAI). The benchmark behind that sentence is AA-Omniscience, from the evaluation firm Artificial Analysis: 6,000 questions across 42 topics in six domains, including law, health, and software engineering, written to be hard enough that models frequently do not know the answer (arXiv).

The definition is the part that matters. On AA-Omniscience, the hallucination rate is the share of questions a model got wrong out of the questions it could not answer correctly, in other words, how often it guessed instead of saying it did not know. A rate of 94 percent does not mean the model is wrong 94 percent of the time you use it. It means that when the model lacks the knowledge, it fabricates an answer 94 times out of 100 rather than admitting the gap. That is a calibration failure, and for anyone who publishes it is the dangerous one, because a model that never says I don't know delivers its fabrications with the same fluency as its facts, a trap we unpack in why AI is confidently wrong.

The finding underneath the range is starker than the range itself. When Artificial Analysis launched the benchmark, 33 of the 36 models tested were more likely to hallucinate than to answer correctly on these hard questions; only three scored above zero on the accompanying Omniscience Index, which rewards correct answers, penalizes wrong ones, and does not penalize declining to answer. GPT-5 (high) posted a hallucination rate of 81 percent. The best rate in the field, Claude 4.5 Haiku's, was still 26 percent, roughly one confident fabrication for every four gaps in its knowledge.

The benchmarks behind the numbers

Four 2026 measurements come up constantly. They test different things, which is why their numbers differ, and why quoting one without its context misleads.

The 2026 hallucination benchmarks

Benchmark What it measures Headline finding Source
AA-Omniscience Guessing vs. abstaining on 6,000 hard factual questions Rates of 22-94% across 26 top models; 33 of 36 more likely to hallucinate than answer correctly Artificial Analysis / Stanford AI Index, 2026
Vectara Hallucination Leaderboard Fabrication when summarizing a document the model is handed Best model (Gemini-3-Pro) at 13.6%; GPT-5, Claude Sonnet 4.5, and Grok-4 all above 10% Vectara, 2026
KaBLE Whether models can tell a user's false belief from a fact, 13,000 questions, 24 models GPT-4o fell from 98.2% to 64.4% accuracy on false first-person beliefs; DeepSeek R1 from over 90% to 14.4% Suzgun et al., Nature Machine Intelligence, 2025
OpenAI PersonQA OpenAI's own hallucination eval on questions about people o1: 16% → o3: 33% → o4-mini: 48%, newer reasoning models hallucinated more OpenAI o3/o4-mini system card, 2025

Sources: Stanford HAI 2026 AI Index; AA-Omniscience; Vectara Hallucination Leaderboard; Suzgun et al., Nature Machine Intelligence; OpenAI system card.

Notice what changes between rows. AA-Omniscience tests unaided recall under pressure, so its numbers are the ceiling of the problem. Vectara hands the model the source and asks only for faithful summarization, about the easiest factual task there is, and the field still cannot get under 10 percent. KaBLE is different in kind: every model tested had the knowledge to answer, and accuracy collapsed anyway the moment a false claim arrived framed as the user's own belief, which is how false claims actually arrive in drafts, briefs, and newsletters. And PersonQA is OpenAI grading itself, which is what makes its direction notable: each successive reasoning model hallucinated more than the last, not less.

Which AI hallucinates the least?

On 2026 data there is no single safest model. Claude 4.5 Haiku posted the lowest hallucination rate on hard factual recall (26 percent on AA-Omniscience). Gemini-3-Pro leads grounded summarization (13.6 percent on Vectara's harder benchmark). And Claude 4.1 Opus topped the Omniscience Index overall, largely by declining to guess when it did not know. The leader changes with the benchmark, and every leader still fabricates at rates no publication would accept from a human researcher.

The Opus result is worth sitting with, because it exposes what low hallucination actually costs. The Omniscience Index does not penalize a model for saying I don't know, and Opus won it by abstaining, trading answers for reliability. That is the right trade for a careful assistant and a useless one for anyone hoping a low rate means the output can ship unchecked: the rate is low because the model answered less, not because its answers became true. A fuller treatment of why the model-ranking question is the wrong question is in is there a most accurate AI model, and the ChatGPT-specific numbers live in how accurate is ChatGPT.

There is a practical use for the fact that models fail differently, though. Trained by different labs on different data, they rarely invent the same fact the same way, so where several independent models agree, the claim is usually sound, and where they disagree, you have found exactly what needs a human check. That disagreement signal is the mechanism TrueStandard runs on: a draft goes through four to five frontier models in parallel, and the claims they cannot all corroborate come back flagged, in about 60 seconds.

Hallucination rates in the wild, by field

Benchmarks are controlled conditions. The documented rates from professional use are the numbers that show what the benchmarks predict.

Documented rates by field, 2023-2026

Field Documented rate Source
Legal research Grounded legal AI tools hallucinated on 17-33% of queries; GPT-5.1 fabricated 6.57% of legal citations, up from 1.23% for GPT-4o Stanford RegLab; Who Checks the Citations? (2026)
Medical & academic citations 47% of ChatGPT-generated medical references fully fabricated; 4,046 fake citations found across 2,810 published papers Cureus (2023); Columbia / The Lancet audit (2026)
News & current events 45% of AI answers about news carried at least one significant issue; 8 AI search engines misattributed sources on 60%+ of queries BBC / EBU (2025); CJR Tow Center (2025)
Software packages 19.7% of AI-suggested package names did not exist; 43% of the fabrications recurred in every one of ten repeated runs USENIX Security (2025)
Personal finance 7 major AI platforms gave inconsistent advice to identical prompts, with recommendations varying by the hypothetical user's race or gender Journal of Financial Planning (June 2026)

Sources: Stanford RegLab; Who Checks the Citations?; Cureus; Columbia Nursing; BBC / EBU study; CJR Tow Center; USENIX Security; CNBC on the JFP study.

The rates land highest on exactly the material that carries the most weight: the citation that proves the claim, the statistic that anchors the argument, the niche detail that makes a piece worth reading. And the failures are no longer near-misses caught in review; they are reaching the permanent record. By mid-2026, researcher Damien Charlotin's public database had logged more than 1,600 court cases worldwide involving AI-hallucinated content, and the Columbia audit found fabricated references sitting uncorrected in peer-reviewed journals. The full catalogue of documented cases, from Mata v. Avianca to CNET's corrections, is in our fake-citation disasters reference.

Four ways these numbers get misread

Hallucination statistics travel badly. The same four misreadings appear almost everywhere the numbers are quoted.

Reading 94% as "AI is wrong 94% of the time"

The AA-Omniscience figures count guesses on questions the model could not answer, questions written to be hard. On everyday, well-documented topics, frontier models are right far more often than not. The honest statement is narrower and more useful: when a model does not know, it usually fabricates rather than admitting it, and nothing in the output tells you which mode you are in.

Quoting the best benchmark number as if it transfers

Gemini-3-Pro's 13.6 percent on Vectara measures summarizing a document it was handed. That number says nothing about open factual recall, where the same generation of models posts rates from 22 to 94 percent. A rate only means something attached to the task it was measured on, and vendors understandably quote the task they score best at.

Assuming newer models hallucinate less

The trend is not monotonic, and OpenAI's own evaluations are the cleanest evidence: on PersonQA, o3 hallucinated at double the rate of o1, and o4-mini reached 48 percent. A 2026 legal-citation benchmark found GPT-5.1 fabricating references at five times the rate of the year-older GPT-4o. Each release is better at many things; being finished with hallucination has not been one of them.

Comparing rates across different definitions

One study counts fully fabricated citations, another counts any significant issue, a third counts guesses on unanswerable questions. The BBC study's 45 percent and AA-Omniscience's 94 percent and Vectara's 13.6 percent are not points on one scale. Before comparing two rates, check whether they count the same failure, and most quoted comparisons do not survive that check.

What the rates mean if you publish

For anyone shipping work with their name on it, the practical readings of this data are stable across every benchmark and field.

First, the floor is not zero anywhere. The best model on the friendliest task still fabricates more than once in ten responses, so verify before you publish is not caution, it is what the data says. Second, confidence carries no signal, since the same fluency wraps the facts and the fabrications, and a model reviewing its own output shares the blind spots that produced it, which is why AI can't reliably check its own work. Third, the errors concentrate in specifics, numbers, names, citations, quotes, so verification effort should concentrate there too. The step-by-step workflow is in how to fact-check AI writing before publishing.

The one property of this data that works in a publisher's favor: independent models almost never share a fabrication. TrueStandard turns that into a check. Paste a draft, and four to five frontier models from different labs verify every claim in parallel; in about 60 seconds you get back the specific claims they disagree on, which are the ones worth your attention before anyone else reads them. Try a single claim free with the claim checker, or see the full AI hallucination checker.

Frequently Asked Questions

How often does AI hallucinate?

It depends on the task. On hard factual questions answered from memory, hallucination rates across 26 top models ranged from 22 to 94 percent in the Stanford AI Index 2026 data. When summarizing a document the model is given, the best 2026 models still fabricate in 10 to 14 percent of responses. In professional use, grounded legal research tools hallucinated on 17 to 33 percent of queries in Stanford's testing. On easy, well-documented topics the rate is far lower, but it reliably rises on specifics: statistics, citations, quotes, and niche subjects.

What is ChatGPT's hallucination rate?

There is no single number, and the spread is the story. GPT-5 (high) posted an 81 percent hallucination rate on AA-Omniscience's hard factual questions, meaning it guessed rather than admitted uncertainty four times out of five when it lacked the answer. On Vectara's grounded summarization benchmark, GPT-5 exceeded 10 percent. OpenAI's own PersonQA evaluation measured o3 at 33 percent and o4-mini at 48 percent, and a 2026 benchmark found GPT-5.1 fabricating 6.57 percent of legal citations. The rate depends entirely on the task; none of these rates is safe to publish unverified.

Which AI model hallucinates the least?

The leader depends on the benchmark. Claude 4.5 Haiku had the lowest hallucination rate on AA-Omniscience's hard factual questions at 26 percent. Gemini-3-Pro led Vectara's document-grounded benchmark at 13.6 percent. Claude 4.1 Opus topped the overall Omniscience Index, mostly by declining to answer when unsure rather than by knowing more. No 2026 model posts a rate a publication would accept from a human researcher, so the practical question is not which model to trust but how to verify whichever one you use.

Are AI hallucination rates going down?

Not reliably. OpenAI's own evaluations showed o3 hallucinating at twice the rate of the earlier o1, with o4-mini higher still, and a 2026 benchmark found GPT-5.1 fabricating legal citations at five times the rate of GPT-4o a year earlier. Average accuracy on easy questions keeps improving, which is a different thing: the remaining errors get more dangerous as the output feels more trustworthy and people check less.

What does the Stanford 22-94% hallucination figure actually mean?

It comes from the AA-Omniscience benchmark: 6,000 deliberately hard questions across law, health, software engineering, and three other domains. The hallucination rate is the share of questions a model answered wrongly out of those it could not answer correctly, in other words, how often it guessed instead of saying it did not know. A 94 percent rate means that model fabricated an answer 94 times out of 100 when it lacked the knowledge. It does not mean the model is wrong 94 percent of the time in everyday use.

How do I protect my work from AI hallucinations?

Treat every specific claim in an AI-assisted draft as unverified until checked: statistics, citations, quotes, names, and dates, because that is where the documented failures concentrate. Trace high-stakes claims to a primary source by hand. For a full draft, use independence: run the claims across several models from different labs and check what they disagree on, since independent models rarely invent the same fact the same way. Agreement is not proof, but disagreement is a reliable pointer to what needs a human eye.

Keep reading

The rates say verify. This is the fast way.

Every benchmark on this page points the same direction: no 2026 model is safe to publish unverified. Paste your draft into TrueStandard and four to five frontier models check every claim in parallel, flagging the ones they can't corroborate in about 60 seconds.

Check Your Draft →