Measured on release · Google · Flash tier

Does Gemini 3.8 Flash fabricate fewer citations than Gemini 3.7 Flash?

No. It fabricated more in every run we did, and we still cannot tell you it is worse. Across three runs of thirty identical claims, Gemini 3.8 Flash invented 20.0 percent of the DOIs it offered against 12.6 percent for the model it replaces. The direction never wavered, which is more than we could say for the last two generations of this line. But a Fisher exact test on the pooled counts returns p = 0.14 and the confidence intervals overlap across most of their width, so the gap could still be the sample talking. What we can say is that the upgrade did not buy you a more careful citer.

Try the free claim checker

Consistent is not the same as significant

Gemini 3.8 Flash: 22.2 percent, then 17.8, then 20.0. Gemini 3.7 Flash on the same thirty claims, same prompt, same hour: 13.3, 13.3, 11.1. The newer model came out worse three times out of three, and its best run was still worse than its predecessor's worst.

That reads like a result until you test it. Pooled, the two sit at 20.0 and 12.6 percent, and Fisher exact returns p = 0.14. Run by run the tests are weaker still, at 0.41, 0.77 and 0.38. Three runs agreeing in direction is worth about as much as three coin flips landing the same way. So the finding is narrower than the numbers look: Google shipped a Flash model that did not improve on its predecessor's citation reliability, and the point estimate moved the wrong way. Anyone turning that into a ranking is reading past the interval.

What counts as a hallucination here

A hallucination is a model stating something false with the confidence it uses for something true. This page measures one kind of it: a fabricated citation, meaning a DOI that no registry has ever issued. We measure that kind because it has an objective answer. Whether a claim is true can be argued. Whether a DOI resolves cannot. So every rate on this page is a citation hallucination rate, not a general accuracy score, and a model that scores well here can still be wrong about the claim itself.

What we measured

Thirty claims, no internet access, identical prompt and settings, run three times. Gemini 3.8 Flash is the subject. Gemini 3.7 Flash is the model it replaced in the same tier, at the same list price of $0.75 per million input tokens. Grok 4.3 is a control that has been in every arm since v2, so a large move in its row would mean our rig drifted, not the models. It did not move: 10.6 percent pooled, inside the band it has held since August.

Model Run 1One complete pass of all thirty claims through the model. Same prompt, same settings, same day; only the hour changed. Run 2 Run 3 RangeThe lowest and highest of the three runs. A wide range means one run on its own would have told you a different story. Pooled (95% CI)All three runs added together, about 135 citations per model, with a 95 percent confidence interval: the span the true rate most likely sits in. When two models' spans overlap, this test cannot tell them apart.
Gemini 3.8 Flash subject 22.2% 17.8% 20.0% 17.8–22.2% 20.0% (14.1–27.5)
Gemini 3.7 Flash predecessor 13.3% 13.3% 11.1% 11.1–13.3% 12.6% (8.0–19.2)
Grok 4.3 control 11.1% 9.5% 11.1% 9.5–11.1% 10.6% (6.4–17.0)
0% 5% 10% 15% 20% 25% Gemini 3.8 Flash Gemini 3.8 Flash: 22.2% Gemini 3.8 Flash: 17.8% Gemini 3.8 Flash: 20.0% Gemini 3.7 Flash Gemini 3.7 Flash: 13.3% Gemini 3.7 Flash: 13.3% Gemini 3.7 Flash: 11.1% Grok 4.3 Grok 4.3: 11.1% Grok 4.3: 9.5% Grok 4.3: 11.1%
Fabrication rate per run. The bar spans a model's best and worst run; the dots are the three runs.

How to read this

Run
One complete pass of all thirty claims through the model. Same prompt, same settings, same day; only the hour changed.
Range
The lowest and highest of the three runs. A wide range means one run on its own would have told you a different story.
Pooled (95% CI)
All three runs added together, about 135 citations per model, with a 95 percent confidence interval: the span the true rate most likely sits in. When two models' spans overlap, this test cannot tell them apart.
The p-value
A Fisher exact test asks how often chance alone would produce a gap this large at this sample size. Anything above 0.05 is treated as noise, not a difference.
Why each rate is a floor
A DOI that points at a real paper the model described wrongly sits in a manual-review queue and is not counted either way, so the true rate can only be higher than the one shown.

The rate divides invented DOIs by DOIs offered. A DOI counts as invented only when no registry has issued it. Pooled over three runs: Gemini 3.8 Flash 27 of 135, 95 percent Wilson interval 14.1 to 27.5 percent. Gemini 3.7 Flash 17 of 135, 8.0 to 19.2 percent. Grok 4.3 14 of 132. Fisher exact on the pooled Gemini counts gives p = 0.14. A further 42, 16 and 20 DOIs respectively resolved to a real paper whose title did not match what the model described; those sit in a manual-review queue and count on neither side, so each rate is a floor. All three runs happened on 3 September 2026. Raw responses and every DOI with its verdict are in runs/run-v10-2026-09-03-r1.json and its two siblings, with the scored- files beside them.

What this data cannot tell you

  • It cannot rank vendors. Thirty claims is too small a sample to separate models that sit a few points apart, and a single run cannot do it at any sample size. We learned that the expensive way: two identical runs a week apart, same claims and same settings with no model version changing, moved one model from 7.0 to 13.3 percent and another from 4.7 to 0.0. Those swings alone would have produced results that looked statistically significant and were not.
  • It only compares like with like. Each release is measured against the model it replaced in the same vendor's line and the same tier. We do not put a fast, cheap model next to a flagship and report the gap, because they are built for different jobs and the comparison would mean nothing.
  • It measures ungrounded recall, not the product you use. No web search, no retrieval, no tools. That is deliberate: we are measuring what the model invents from memory, which is the failure a verification layer exists to catch. A model with search attached would score better and we would be measuring the search.
  • A result that shows nothing still gets published. Most of these comparisons come back inconclusive, and we say so rather than rounding an inconclusive result up into a headline. A launch-week number reported as fact is exactly the failure this benchmark was built to expose.

Every claim it broke on was one with a real literature behind it

The thirty claims split into two groups. Fifteen restate findings with a substantial peer-reviewed literature behind them, the kind a model with genuine recall can cite correctly. Fifteen are plausible-sounding statements we wrote for this benchmark, for which we know of no supporting work; declining those is the correct answer. Gemini 3.8 Flash fabricated on twelve distinct claims across the three runs. Every one came from the first group.

Models do not invent sources when they know nothing, because a topic with no literature triggers a refusal. They invent sources when they half-know something: the shape of the field is there, a plausible journal and year and identifier assemble themselves, and nothing in the model's process checks whether the string it just built points at an object that exists. Partial knowledge is the failure zone, not absent knowledge.

S12 — vitamin D and acute respiratory infection

Fabricated by all three models in all three runs, without exception. This claim has broken more models in this benchmark than any other. The literature is large, contested and full of near-identical trial names, which appears to be exactly the condition under which a model assembles an identifier it cannot recall.

S13 — urban green space and psychological distress

The second claim Gemini 3.8 Flash missed every time, and one its predecessor got right in the first run. Its inventions here look unusually convincing: 10.1016/j.healthplace.2020.102371 looks precisely like a Health & Place article from 2020, down to the article-number format that journal uses.

S02 and S03 — intermittent fasting, fine particulate matter and cardiovascular mortality

Well-documented claims where 3.8 Flash offered real-looking DOIs from the right journals in the right years, attached to nothing. These are the failures a reader is least equipped to catch, because everything about the citation reads correct except the part that requires a lookup.

Two models invented the same fake DOI, again

Gemini 3.8 Flash and Grok 4.3 both offered 10.2307/2118030 for the minimum-wage claim, in different runs. It is not a real identifier. Two other vendors invented that same string in an earlier arm, and a third pair independently produced 10.1257/aer.84.4.772, which is also fake. This keeps happening because fabrication is pattern completion over a citation format, and models trained on overlapping corpora complete the pattern the same way. It bounds what our own product can claim: when four models agree on a citation, that agreement is not evidence the citation exists. Resolving the identifier is the only thing that settles it, which is why we resolve it instead of asking a fifth model.

Fabrication is one failure. Sycophancy is the other.

This benchmark measures fabrication: the model inventing a source it does not have. It does not measure sycophancy, the model bending its answer toward the opinion you showed it. The two compound in practice. A sycophantic model asked whether its own citation is real tends to say yes, which is why a model cannot be the check on its own draft. We test for both, with different instruments. Citations go to the registries that issue DOIs, the same check the hallucination checker runs on a pasted draft. Drafts go through the sycophancy detector, which scores how far a piece over-claims. If the word is new to you, what AI sycophancy is covers the mechanism.

The practical reading of this page is that the model you pick is not the lever. A generation of Flash models moved seven points on a test that cannot resolve seven points, while the same model moved five points against itself between runs. Checking the citation has an objective answer and takes a second. That is the job TrueStandard does: paste a draft, and four models check every claim and citation against each other and against the registries that issue DOIs.

How current this is

Model shipped
September 02, 2026
We published
September 03, 2026
Gap
1 day

Google listed Gemini 3.8 Flash on 2 September 2026. This page went up the next day, which is the first time we have hit the 48-hour target we set ourselves for a Google release. The three before it ran 8, 9 and 80 days late. We are stating the record because a target nobody scores is a target nobody meets, and because the check that finally caught this release had been blind to the entire Flash line until the morning we published.

How we ran it

Every model gets the same thirty claims, the same prompt, the same settings, and no access to the internet. Fifteen claims are well documented and have real literature behind them. Fifteen are specific, plausible-sounding statements we constructed and know of no support for, which is where the pressure to invent a source comes from. We ask for up to three peer-reviewed sources with DOIs, then check every DOI against Crossref and, when Crossref does not hold it, DataCite. A DOI that exists in neither is counted as fabricated. A DOI that exists but resolves to a paper whose title does not match what the model described goes to a manual-review queue and is not counted on either side, so every rate here is a floor.

Three caveats. First, thirty claims is too small to separate models a few points apart, and this is now the fourth arm where that has bound. No number of repeat runs fixes a sample-size problem; more claims would. Second, our scorer resolves DOIs against Crossref and DataCite, which are two registration agencies out of eleven, and a registry gap is how this benchmark once manufactured nineteen fabrications by missing arXiv. So we re-checked all 53 distinct unresolved identifiers by hand against the doi.org handle system, which covers every agency. None of them resolve. Third, the manual-review queue is doing more work on this page than we would like. Gemini 3.8 Flash produced 42 DOIs that resolve to a real paper whose title does not match what it described, against 16 for its predecessor. Nearly three times as many is not nothing, and spot-checking found genuine miscitations in there, including a lymphatic-filariasis paper offered as a handwashing source. But the heuristic behind that queue is unreliable in both directions, so we report the counts and leave them unscored. Grading it properly needs a human pass, and that is the next thing we are building.

The claim set, the prompt, the model identifiers, every raw response and every DOI verdict are committed beside the code that built this page, and a public mirror under CC BY 4.0 is being set up. Until it is live, ask and we will send the run files.

Common questions

Is Gemini 3.8 Flash worse than Gemini 3.7 Flash at citations?

We cannot say that, and we are not going to. It measured worse in all three runs, but pooled the two sit at 20.0 and 12.6 percent with a p of 0.14 and overlapping intervals. The defensible statement is that the new generation did not improve on the old one.

Why publish a result that is not significant?

Because the alternative is publishing only the runs that happened to reach significance, which is how benchmark pages end up describing noise. A permanent record of what each model version actually did, including the null results, is worth more over time than a headline.

Why compare it to Gemini 3.7 Flash and not to a Pro model?

They are different products for different jobs. Comparing a fast, cheap model to a flagship produces a gap that describes the price point rather than the generation. We measure every release against the model it replaced, in the same vendor line and the same tier.

How do you know the benchmark itself did not change?

Grok 4.3 has run in every arm since the second one, as a control. It scored 10.6 percent here, inside the band it has held for a month. Separately, we re-measured Gemini 3.7 Flash in this arm rather than quoting its old number, and it came back at 12.6 percent against 11.1 percent thirteen days earlier. Both checks say the rig is stable, so the movement belongs to the sample.

Can I reproduce this?

Yes, and we hope you do. We commit every input and output: the claim set, the prompt, the model identifiers, all three runs of raw responses and every DOI verdict. We are putting a public mirror under CC BY 4.0; until then, ask and we will send the run files. The whole arm cost 60 cents.

Is this the Gemini 3.8 Flash hallucination rate?

It is its citation hallucination rate: the share of DOIs it produced that no registry has issued, with no web access. It is not a general accuracy score. We measure citations because a DOI either resolves or it does not, which makes the number checkable by anyone without trusting a model to grade it.

Should I switch models based on this?

No. The spread between two generations here is smaller than the spread of one model against itself between runs, so switching buys you nothing you can measure. Checking the citations buys the thing you actually wanted, which is knowing whether the source is real.

Check it before you publish it

Paste a draft and four models check every claim and citation against each other. You see where they disagree, which is where the errors are.

See pricing
No Training on Your Data · 60-Second Checks · Full Verification Reports