Every ChatGPT vs Claude vs Gemini comparison reaches a verdict on accuracy, and almost none of them measured anything. Ask Google which of the three is most accurate and its own summary answers Claude, drawn from pages that ran no test. We run a citation-fabrication benchmark across all three, so we can tell you what turns up when somebody does measure it.
On the run we finished on 10 August, the model the consensus calls the most accurate produced the most fake sources of the three. A week earlier, on the same claims and the same prompt, it sat in the middle. Both numbers are ours and both are real.
The short answer
All three plans cost about twenty dollars. What the reviews agree on is roughly right. The accuracy ranking they add on top is the one thing nobody tested.
The three plans, and the one column that is actually measured
| Plan | What the reviews agree it is for | Model we tested | Fabricated sources, 10 Aug |
|---|---|---|---|
| Google AI Pro | Google Workspace, video, the widest bundle | Gemini 3.1 Pro | 2.4% |
| ChatGPT Plus | Generalist daily driver, images, breadth | GPT-5.6 Sol | 7.7% |
| Claude Pro | Writing, coding, long documents | Claude Sonnet 5 | 24.4% |
Do not buy on that last column. It is the most honest number on this page and it still cannot rank these three, because we ran the identical test seven days earlier and Claude came out at 14.0% while Gemini came out at 11.6%. Two of the three rows moved about ten points with no model release in between.
That movement is worth more to you than the ranking, because it tells you what a twenty dollar decision can and cannot be made on.
Every review picks an accuracy winner
Pull the first page of results for this question and you get comparison posts from a dozen sites, most of them careful and useful on the things you can see. They agree on the shape: ChatGPT for breadth, Claude for writing and code, Gemini for anything that touches Google.
Then each one reaches an accuracy verdict. Claude wins for accuracy. Claude handles ambiguity better and produces outputs that are easier to trust. Claude is the most reliable for multi-step reasoning. Read four of these pages and you will meet the same claim four times, phrased four ways, sourced to nothing.
We fetched the two closest ranking pages to check. One is a 4,000-word piece that runs real prompts through all three models and screenshots the answers, and its accuracy section is a general warning that chatbots hallucinate. The other says up front that the author paid for all three and ran the same work through each, then reaches a buying recommendation without a single accuracy measurement anywhere in it.
The pattern holds off the search page too. A plan review that landed this month, 34,000 views in five days, tells its audience that one model tends to stay closest to the facts on long answers and invents less when a question gets complicated. No test, no number, no source. In the comments underneath, one viewer said they had cancelled a subscription that day on the strength of it.
The reviewers are not being sloppy. Usage limits, connectors, storage and video are all things a careful person can check in a month with three subscriptions open. Fabrication rate is not. It needs a frozen claim set, a registry lookup and a scoring pass that no model is allowed anywhere near, and none of that fits in a review cycle.
So the accuracy leg is the one you cannot outsource to a reviewer. You can borrow somebody else's judgment on whether the video generator is any good. You cannot borrow their judgment on whether the thing you are about to publish is true. TrueStandard runs your draft past four models from different labs at once and surfaces every claim they cannot agree on, in about 60 seconds.
What we measured on all three
Our benchmark asks a model for peer-reviewed sources on 30 frozen claims, half with real literature behind them and half constructed for the test with none. Every DOI that comes back gets resolved against Crossref. A DOI that exists in neither Crossref nor doi.org is scored fabricated. No language model touches the scoring path, because a test about models inventing detail cannot be graded by a model.
Fabricated citations by model, 30 claims, ungrounded, 2026-08-10
| Model | Sources given | Fabricated | Rate | 95% interval |
|---|---|---|---|---|
| Gemini 3.1 Pro | 41 | 1 | 2.4% | 0.4 - 12.6% |
| GPT-5.6 Sol | 39 | 3 | 7.7% | 2.7 - 20.3% |
| GPT-5.5 | 45 | 5 | 11.1% | 4.8 - 23.5% |
| Claude Sonnet 5 | 45 | 11 | 24.4% | 14.2 - 38.7% |
Read the intervals before the rates. Gemini's runs from 0.4% to 12.6%. GPT-5.6's runs from 2.7% to 20.3%. Those two ranges overlap across most of their width, on 30 claims, which is what a small sample buys you. The gap you can see in the middle column is thinner than it looks.
The shape of the failures matters more than the count. In our first run, all 24 fabrications landed on claims that have genuine published literature behind them. Not one landed on a claim we invented. Partial knowledge is the danger zone: the model knows the topic and knows roughly which paper answers it, then invents the identifier. Claude gave us a handwashing citation with the right journal, right year and right author, and a DOI that does not exist. Nothing about it looks wrong until you resolve it.
The full study, all four original models, every claim and every individual verdict, is written up in our four-model citation benchmark. For the published academic numbers instead of ours, see the 2026 hallucination-rate reference.
The same test, one week apart
On 10 August we kept every model from the previous run in the same arm. That turned the run into something we had not done before: an exact repeat. Same 30 claims, same prompt, same settings, same scorer, seven days later, with no model version changing in between.
If the measurement is stable, every row should barely move.
Identical claims, identical settings, one week apart
| Model | 3 Aug | 10 Aug |
|---|---|---|
| GPT-5.5 | 11.1% | 11.1% |
| Claude Sonnet 5 | 14.0% | 24.4% |
| Gemini 3.1 Pro | 11.6% | 2.4% |
GPT-5.5 reproduced to the cell: five fabrications out of forty-five, both times. So did Grok 4.3, at eleven out of forty-five. The pipeline is deterministic and the questions did not change, so the movement in the other two rows belongs to the models.
Claude nearly doubled. Gemini fell to a fifth of its previous rate. On 3 August the three were separated by about three points and you could not tell them apart. On 10 August they were separated by twenty-two points and Gemini looked untouchable.
What that invalidates
Our 10 August run produces two pairwise gaps that clear the usual significance bar, Gemini against Claude at p = 0.004 among them. Run the identical comparison on the 3 August data and neither one exists. Both depend on Claude sitting at a one-week high and Gemini at a one-week low on the day we happened to measure. A p-value computed inside a single run knows nothing about the variance between runs, so it is more confident than the evidence deserves.
So we will not publish those as vendor findings, and nobody quoting this page should either. Our earlier caution was that 30 claims cannot rank vendors. The version this run forced on us is harder: one run cannot rank vendors at any sample size. Ranking needs repeated sampling across days, and two data points is where we are.
The practical reading for a subscriber: the gap between one model's good week and its bad week was larger than the gap between most of the models. Picking the safest vendor is not a strategy when the safest vendor is a coin that lands differently on Monday. Running the same claim past several models at once is what TrueStandard does instead, because a claim several independent labs corroborate is standing on something firmer than one vendor's Tuesday.
ChatGPT Plus
$20 / monthWhat it is bought for
The generalist. Image generation with the strongest text-inside-the-image handling of the three, the deepest third-party ecosystem, a cloud coding agent, and the clearest usage allowance on the market: OpenAI publishes an actual message count and an actual reset window, where the other two weigh how heavy your requests are and leave you guessing.
What we measured
GPT-5.6 Sol fabricated 3 of the 39 sources it offered, a rate of 7.7%, with an interval running to 20.3%. Its predecessor GPT-5.5 sat at 11.1% in the same arm. The point estimate fell by about a third and a Fisher exact test returns p = 0.719, so this test cannot show the drop is real.
One thing did change in kind, not degree. On the 15 claims we invented, the ones with no literature behind them, GPT-5.6 emitted no source at all. Fifteen out of fifteen declined. GPT-5.5 had leaked on two. A model that says nothing when it has nothing is doing the thing that protects you, and that behaviour is worth more than a point or two off a rate.
Best for
Anyone who wants one assistant that does most things well and would rather know exactly when they will hit the ceiling. It is also the least surprising of the three, which counts for more than people admit.
The full generational comparison, including why we published a null result, is in our GPT-5.6 write-up. For ChatGPT accuracy on its own terms, see how accurate ChatGPT actually is.
Claude Pro
$20 / monthWhat it is bought for
Writing and code, mostly. Its prose reads less like a machine than the other two, it can take a sample of your writing and match the tone, and it keeps projects as saved workspaces so the context survives between sessions. On the coding benchmarks it leads, and its desktop agent works on the ordinary files sitting on your own machine, not only on a code repository.
What we measured
Claude Sonnet 5 fabricated 11 of the 45 sources it offered on 10 August, a rate of 24.4%, the worst of the three plans. On 3 August, same test, it was 14.0%. In our first run, under a prompt that told the model declining was acceptable, it was 29.4%. Across every version of this test we have run, Claude's number has never been the lowest and has never been in the same place twice.
The caveat that matters
We tested Sonnet 5. Claude Pro's headline model is Opus 4.8, which we have not put through this benchmark. So this is not a measurement of the model most Claude Pro subscribers open by default, and we are not going to pretend otherwise. What it does show is that the vendor the whole first page of Google calls the most accurate has, on the one axis anybody has actually measured, been the worst of the three every time we looked.
Best for
People whose output is words or code, who will read every line before it goes anywhere. The tighter usage ceiling is the real trade, not the fabrication rate, because you should be checking sources regardless of which of the three you pay for.
Google AI Pro
About $20 / monthWhat it is bought for
You are buying a bundle more than an assistant. It is the only one of the three that generates video, its image editing is level with ChatGPT's on photo work, and it drops into Gmail, Docs and Sheets where the work already happens. On top of the assistant you get cloud storage, higher research-tool limits and a set of perks unrelated to chat, which is why it wins the value argument on every comparison page without much of a fight.
What we measured
Gemini 3.1 Pro fabricated 1 of the 41 sources it offered on 10 August, a rate of 2.4%, the lowest number any model has posted on this benchmark. Read the interval: 0.4% to 12.6%. And read the week before, where the identical test put it at 11.6%, roughly where GPT-5.5 sat.
So the honest sentence about Gemini is that it had an exceptional week on a small test, and we cannot tell you yet whether that is the model or the day. Anyone quoting 2.4% as Gemini's fabrication rate is quoting one Monday.
Best for
Anyone already living inside Google Workspace, and anyone who wants the most tool per dollar. It is also the one least likely to cut you off mid-afternoon, which for heavy users decides more than any benchmark does.
Your plan does not pin your model
There is a structural reason the accuracy column on every comparison page goes stale faster than the rest of the table. You are not buying a model. You are buying a subscription that points at whichever model the vendor has decided to serve this week.
GPT-5.6 Sol became callable on 9 July. Our own benchmark was rebuilt on 3 August and still measured GPT-5.5 as OpenAI's current model, 25 days after its successor shipped. We published a fabrication rate for a model that had already been replaced, and it was the second time we had done it. In the comments under that plan review, somebody asked whether ChatGPT was not already on 5.6. They were right, and the video was out of date on the day it went up.
So even a correctly measured accuracy number is attached to a model your subscription can swap out from under you, without an email, in the middle of a billing cycle. Our GPT rows are a case in point: two generations of the same family, four points apart, and both of them inherited the same two invented identifiers verbatim.
A generation change did not disturb either fabrication. Both are stable artifacts of how that family completes a citation pattern. One of them, a JSTOR identifier, was also invented independently by Gemini for the same claim in an earlier run. Three model instances, two vendors, two generations, one record that has never existed.
Which sets a limit on cross-checking too. Comparing models works because independent models fail independently, and two models completing the same pattern can fail identically. That is why our benchmark resolves identifiers against a registry instead of asking one model to grade another, and why TrueStandard shows you the disagreements instead of one confident verdict. A claim four labs agree on is a stronger bet. It is still a bet, and you should be able to see the seams.
How to choose without an accuracy verdict
Take the accuracy column out of the decision. It is not stable enough to buy on and it will have moved before your second invoice. What is left is workflow, ceilings and what else comes in the box, all of which you can check yourself in a week.
Which twenty dollars, by what you actually do
| If you | Buy | Because |
|---|---|---|
| Live in Gmail, Docs and Drive | Google AI Pro | The integration is native and the bundle is the widest of the three at the same price |
| Want one assistant for everything | ChatGPT Plus | Strongest image work, deepest ecosystem, and the only published usage numbers you can plan around |
| Write or ship code for a living | Claude Pro | The prose reads most naturally, it leads on coding, and its desktop agent works on your own files |
| Need video | Google AI Pro | It is the only one of the three that generates video at all |
| Use AI hard, all day | Google AI Pro | Compute-weighted limits are the most forgiving; Claude runs out fastest and drops you to nothing |
| Publish anything with sources in it | Any of them, then check | The best rate we have measured on any model still means one fake source in roughly forty |
The part that does not change with the plan
Whichever you pick, the number that should govern how you work is the one at the top of the range, not the middle. Gemini's interval runs to 12.6%. GPT-5.6's runs to 20.3%. Claude's starts at 14.2%. Those are the numbers to plan against, and none of them is low enough to publish a sourced claim without resolving it.
For the conceptual version of this argument instead of the buying version, we wrote it up as whether a most accurate AI model exists at all, and as why picking the model that hallucinates least is the wrong move.
Buy the one that fits your week. Then treat every source it hands you as unverified, because on the only axis that has been measured, the leader changed between two Mondays.
FAQ
Which is most accurate, ChatGPT, Claude or Gemini?
On our 10 August run, Gemini 3.1 Pro fabricated the fewest citations at 2.4%, GPT-5.6 Sol came next at 7.7% and Claude Sonnet 5 was worst at 24.4%. We do not think that is a ranking. The identical test seven days earlier put Claude at 14.0% and Gemini at 11.6%, with no model version changing in between. The honest answer is that the three are close enough, and the measurement noisy enough, that no single run separates them.
Is Claude really the most accurate AI model?
That claim is repeated across most comparison pages and we have not found a measurement behind it. On the one axis we test, whether the sources a model gives you exist, Claude Sonnet 5 has been the worst of the three in every run we have done: 29.4% under our first prompt, then 14.0% and 24.4% under the current one. Claude Pro's default model is Opus 4.8, which we have not benchmarked, so this is not a verdict on the model most subscribers use. It does mean the consensus is running on reputation, not data.
Which $20 AI plan is worth it?
Google AI Pro for the widest bundle and the most forgiving usage limits, ChatGPT Plus for a predictable generalist with the best image tools, Claude Pro if your day is writing or code. Do not let a claimed accuracy advantage break the tie, because that number moves ten points week to week on the same test.
How often does ChatGPT make up sources?
In our latest run GPT-5.6 Sol invented 3 of the 39 sources it volunteered, about one in thirteen, with a 95% interval from 2.7% to 20.3%. Its predecessor GPT-5.5 invented 5 of 45. Every fabrication landed on a topic with real published literature behind it, which is what makes them hard to spot: right journal, right author, right year, identifier that resolves to nothing.
Why do accuracy rankings between AI models keep changing?
Partly because vendors ship new models every few weeks, and partly because the tests are small. We re-ran 30 identical claims a week apart with nothing changed and two of five models moved about ten points. At that sample size, run-to-run variance is larger than the gap between most of the models, so a leaderboard built on one run is measuring the day as much as the model.
Does paying for a plan get you a more accurate model?
It gets you a bigger model and a higher usage ceiling, not a factual guarantee. Every model in our benchmark was a current paid-tier frontier model, and the best of them still invented a source. The gains from paid tiers show up in reasoning depth, context length and usage ceilings. None of them decides whether a citation exists.
Can I trust an AI comparison that did not run a test?
On the things you can see, usually yes. Usage limits, connectors, storage, video and the feel of the writing are all checkable by one person with three subscriptions. Accuracy is not. It needs a claim set that never changes and a scoring pass with no model anywhere in the loop. When a comparison asserts an accuracy winner with no method attached, treat that line as an impression, not a finding.
What should I do about fake citations regardless of which plan I buy?
Resolve every identifier before you publish. A DOI that does not open is the fastest tell, and it catches the failure mode that looks most convincing. For anything longer than a paragraph, run the draft past several models from different labs and read what they disagree on, since independent models tend to fail on different claims.
Keep reading
Is There a Most Accurate AI Model?
The honest answer is no. The ranking changes with the task, the benchmark, and the month, and even the leader still hallucinates.
How Accurate Is ChatGPT?
Accurate enough to trust for everyday questions, and wrong often enough to get you sued if you publish it unchecked. Here is what the measurements actually say, and what to do about it.
Why AI Is Confidently Wrong
Models sound certain every time, even when wrong. The confident tone you trust in people is worthless here. Here is the fix.
Should You Stop Using ChatGPT?
Researchers found AI made experts measurably worse on hard tasks. Here is when to trust ChatGPT, and when it is just telling you what you want to hear.
TrueStandard vs Parafact
Both verify claims before you publish. The real difference is what one model can miss — and whether your long-form draft fits inside the check at all.
The leader changed between two Mondays. Check the sources anyway.
The lowest fabrication rate we have ever measured on any model still means about one invented source in forty, and the model that posted it was middle of the pack the week before. No subscription solves this, and switching vendors does not either. Paste your draft into TrueStandard and four models from different labs check every claim in parallel, flagging what none of them can corroborate in about 60 seconds.
Check Your Draft →