Does Gemini 3.7 Flash fabricate fewer citations than Gemini 3.6 Flash?
There is no measurable difference, and we can show you why anyone claiming otherwise is reading noise. Across three runs of thirty identical claims, Gemini 3.7 Flash fabricated 11.1 percent of the DOIs it offered and Gemini 3.6 Flash 14.1 percent, which a Fisher exact test cannot separate at p = 0.58. Our first run said the new model was two and a half times worse. The second and third said it was better. Nothing changed between them except the hour.
The first run said the opposite of the next two
Run one: Gemini 3.7 Flash 11.1 percent, Gemini 3.6 Flash 4.4 percent. Read alone, that says Google shipped a worse model. Run two, same claims, same settings, same afternoon: 13.3 percent against 22.2 percent, so the gap reversed and widened. Run three: 8.9 against 15.6, reversed again.
Pooled over all three runs the two models sit at 11.1 and 14.1 percent with overlapping intervals and a p of 0.58. The honest answer to the question in the headline is no. The useful answer is that one run of this test could have told you either story, depending on which hour you ran it.
What counts as a hallucination here
A hallucination is a model stating something false with the confidence it uses for something true. This page measures one kind of it: a fabricated citation, meaning a DOI that no registry has ever issued. We measure that kind because it has an objective answer. Whether a claim is true can be argued. Whether a DOI resolves cannot. So every rate on this page is a citation hallucination rate, not a general accuracy score, and a model that scores well here can still be wrong about the claim itself.
What we measured
Thirty claims, no internet access, identical prompt and settings, run three times. Gemini 3.7 Flash is the subject. Gemini 3.6 Flash is the model it replaced in the same tier. Grok 4.3 is a control that has been in every arm since v2, so a large move in its row would mean our harness drifted rather than the models. All three are models our own verification runs on, which is the bar for appearing here at all.
| Model | Run 1One complete pass of all thirty claims through the model. Same prompt, same settings, same day; only the hour changed. | Run 2 | Run 3 | RangeThe lowest and highest of the three runs. A wide range means one run on its own would have told you a different story. | Pooled (95% CI)All three runs added together, about 135 citations per model, with a 95 percent confidence interval: the span the true rate most likely sits in. When two models' spans overlap, this test cannot tell them apart. |
|---|---|---|---|---|---|
| Gemini 3.7 Flash subject | 11.1% | 13.3% | 8.9% | 8.9–13.3% | 11.1% (6.8–17.5) |
| Gemini 3.6 Flash predecessor | 4.4% | 22.2% | 15.6% | 4.4–22.2% | 14.1% (9.2–20.9) |
| Grok 4.3 control | 8.9% | 11.4% | 7.1% | 7.1–11.4% | 9.2% (5.3–15.3) |
How to read this
- Run
- One complete pass of all thirty claims through the model. Same prompt, same settings, same day; only the hour changed.
- Range
- The lowest and highest of the three runs. A wide range means one run on its own would have told you a different story.
- Pooled (95% CI)
- All three runs added together, about 135 citations per model, with a 95 percent confidence interval: the span the true rate most likely sits in. When two models' spans overlap, this test cannot tell them apart.
- p = 0.58
- A Fisher exact test asks how often chance alone would produce a gap this large at this sample size. Anything above 0.05 is treated as noise, not a difference.
- Why each rate is a floor
- A DOI that points at a real paper the model described wrongly sits in a manual-review queue and is not counted either way, so the true rate can only be higher than the one shown.
The rate is invented DOIs over DOIs offered. A DOI counts as invented only when it exists in neither Crossref nor DataCite. Pooled over three runs: Gemini 3.7 Flash 15 of 135, 95 percent Wilson interval 6.8 to 17.5 percent. Gemini 3.6 Flash 19 of 135, 9.2 to 20.9 percent. Grok 4.3 12 of 131. Fisher exact on the pooled Gemini counts gives p = 0.58. A further 16, 21 and 26 DOIs respectively resolved to a real paper whose title did not match what the model described; those sit in a manual-review queue and count on neither side, so each rate is a floor. All three runs happened on 21 August 2026, hours apart. Raw responses and every DOI with its verdict are in runs/run-v5-2026-08-21*.json and the scored- files beside them.
What this data cannot tell you
- It cannot rank vendors. Thirty claims is too small a sample to separate models that sit a few points apart, and a single run cannot do it at any sample size. We learned that the expensive way: two identical runs a week apart, same claims and same settings with no model version changing, moved one model from 7.0 to 13.3 percent and another from 4.7 to 0.0. Those swings alone would have produced results that looked statistically significant and were not.
- It only compares like with like. Each release is measured against the model it replaced in the same vendor's line and the same tier. We do not put a fast, cheap model next to a flagship and report the gap, because they are built for different jobs and the comparison would mean nothing.
- It measures ungrounded recall, not the product you use. No web search, no retrieval, no tools. That is deliberate: we are measuring what the model invents from memory, which is the failure a verification layer exists to catch. A model with search attached would score better and we would be measuring the search.
- A result that shows nothing still gets published. Most of these comparisons come back inconclusive, and we say so rather than rounding an inconclusive result up into a headline. A launch-week number reported as fact is exactly the failure this benchmark was built to expose.
Gemini 3.6 Flash moved eighteen points against itself
The predecessor scored 4.4 percent, then 22.2, then 15.6. Same model, same thirty questions, same settings, one afternoon. Its worst run was five times its best. Whatever that spread is, it is not a property of the model you could act on.
We expected the claims a model breaks on to hold steadier than the rate. They do not. Gemini 3.7 Flash fabricated on five different claims across the three runs and repeated only one every time. Gemini 3.6 Flash touched nine and repeated two. We argued in an earlier post about Grok that a newer model fixing claims without breaking new ones was weak evidence of real improvement. This data retires that argument. At thirty claims the failure set wanders about as much as the rate, so a clean nesting in one run is a coincidence waiting for the next run to undo.
The only claim Gemini 3.7 Flash fabricated on in all three runs, and the claim that has broken almost every model we have tested. Whatever this literature looks like in a model's memory, it produces a confident identifier that does not exist, run after run.
The two Gemini 3.6 Flash broke every time. Everything else in its nine-claim failure set showed up in one or two runs and vanished in the others.
Well-documented claims Gemini 3.6 Flash got right in the first run and invented sources for later. Nothing about the questions changed.
Where agreement told us nothing
On the claim that language models are sycophantic, that they shift a stated answer toward whatever opinion the user expressed, Gemini 3.7 Flash, Gemini 3.6 Flash and Grok 4.3 all offered the same arXiv identifier in two of the three runs. It is a real paper, so all three were right. The agreement is not what made them right: all three rebuilt the same familiar string from the same shape. Earlier arms caught that convergence producing identifiers that do not exist, with two vendors independently inventing the same fake DOI. Models agreeing is not corroboration when they are completing the same pattern. Resolving the identifier is the only check that settles it.
Fabrication is one failure. Sycophancy is the other.
This benchmark measures fabrication: the model inventing a source it does not have. It does not measure sycophancy, the model bending its answer toward the opinion you showed it. The two compound in practice. A sycophantic model asked whether its own citation is real tends to say yes, which is why a model cannot be the check on its own draft. We test for both, with different instruments. Citations go to the registries that issue DOIs, the same check the hallucination checker runs on a pasted draft. Drafts go through the sycophancy detector, which scores how far a piece over-claims. If the word is new to you, what AI sycophancy is covers the mechanism.
The practical version of this page is short. You cannot pick a safe model, because the difference between two generations is smaller than the difference between one model and itself an hour later. You can check the citation, which has an objective answer and takes a second. That is the job TrueStandard does: paste a draft, and four models check every claim and citation against each other and against the registries that issue DOIs.
How current this is
Google shipped Gemini 3.7 Flash on 13 August 2026. We aim to publish within 48 hours of a frontier release. We missed this one by more than a week. That is a process problem we are fixing, not a claim we will drop, so the gap sits above rather than buried.
How we ran it
Every model gets the same thirty claims, the same prompt, the same settings, and no access to the internet. Fifteen claims are well documented and have real literature behind them. Fifteen are specific, plausible-sounding statements we constructed and know of no support for, which is where the pressure to invent a source comes from. We ask for up to three peer-reviewed sources with DOIs, then check every DOI against Crossref and, when Crossref does not hold it, DataCite. A DOI that exists in neither is counted as fabricated. A DOI that exists but resolves to a paper whose title does not match what the model described goes to a manual-review queue and is not counted on either side, so every rate here is a floor.
Two caveats. First, all three runs happened on the same day, hours apart. The design called for three separate days and we compressed it. The point of spreading runs across days was to catch variance we could not see inside one day, and the variance inside one day turned out to be larger than the between-day swings that motivated the design. We have not tested cross-day replication, and it is now the less interesting question. Second, our scorer checked DOIs against Crossref alone. arXiv registers with DataCite, so the scorer counted real preprints as inventions across every model in every earlier arm. It now checks both registries. We re-scored every run, corrected the older figures elsewhere on this site, and took the numbers here from the corrected scorer.
The claim set, the prompt, the model identifiers, every raw response and every DOI verdict are committed beside the code that built this page, and a public mirror under CC BY 4.0 is being set up. Until it is live, ask and we will send the run files.
Common questions
Is Gemini 3.7 Flash better or worse than Gemini 3.6 Flash at citations?
Neither, as far as this test can tell. Pooled over three runs they sit at 11.1 and 14.1 percent with a p of 0.58. The runs disagreed about the direction. Anyone quoting one of them has even odds of quoting the reverse of the truth.
Why does the same model score a different rate on the same questions?
Thirty claims is a small sample, and these endpoints are not deterministic even at temperature zero. We checked: 26 of 30 responses differed between two runs of the same model. A few inventions either way moves the percentage several points. So we run the arm three times and publish the range, not the best number.
Why compare it to Gemini 3.6 Flash and not to a Pro model?
They are different products for different jobs. Comparing a fast, cheap model to a flagship produces a gap that describes the price point rather than the generation. We measure every release against the model it replaced, in the same vendor line and the same tier.
Should I switch models based on this?
No, and that is the most useful thing this page can tell you. The spread between two generations is smaller than the spread of one model against itself. Switching vendors on a benchmark number buys nothing measurable. Checking the citations buys the thing you wanted.
Can I reproduce this?
Yes, and we hope you do. Every input and output is committed: the claim set, the prompt, the model identifiers, all three runs of raw responses and every DOI verdict. A public mirror under CC BY 4.0 is being set up; until then, ask and we will send the run files. The whole arm cost 63 cents.
Is this the Gemini 3.7 Flash hallucination rate?
It is its citation hallucination rate: the share of DOIs it produced that no registry has issued, with no web access. It is not a general accuracy score. We measure citations because a DOI either resolves or it does not, which makes the number checkable by anyone without trusting a model to grade it.
Does this page measure sycophancy?
No. Sycophancy is a model shifting its answer toward the opinion you expressed, and this arm expresses none. It asks for sources and checks whether they exist. The two failures compound: ask a sycophantic model to confirm its own citation and it will. That is why the check goes to the registry, not back to the model.
Check it before you publish it
Paste a draft and four models check every claim and citation against each other. You see where they disagree, which is where the errors are.