AI Reliability

Does Gemini Hallucinate Less With More Thinking?

We asked Gemini 3.7 Flash for sources on the same 30 claims at all four of its thinking levels. Then we checked every DOI it gave us against the real registries.

Does Gemini Hallucinate Less With More Thinking?

Gemini 3.7 Flash made up 13.3% of the citations it gave at its minimal thinking level and 4.5% at high. The low and medium levels barely moved the rate. Almost all of the gain came from the last step, from medium to high.

Google lets you set how much Gemini 3 models think, from minimal up to high. Anthropic's post on effort made the case that more thinking pays off on hard coding work. We ran the same test on three vendors to see if it pays off on a simpler job: naming real sources for a claim, from memory, without search.

Only the high thinking level cut Gemini's fake citations

We scored each DOI, the permanent ID printed on a research paper. Each row is 90 answers: 30 claims, asked three times. Half the claims have real research behind them. We made up the other half, so no real paper supports them.

Gemini 3.7 Flash at four thinking levels, 30 claims, three runs, September 2026

Thinking level DOIs given Made up Rate 95% range Cost per 100 answers
Minimal 135 18 13.3% 8.6 to 20.1% $0.26
Low 135 16 11.9% 7.4 to 18.4% $0.24
Medium 137 15 10.9% 6.7 to 17.3% $0.40
High 133 6 4.5% 2.1 to 9.5% $0.57

Made up means the DOI exists in neither Crossref nor DataCite, the two registries that issue DOIs. The 95% range is where the true rate likely sits, given how many DOIs we saw. Cost is what we paid per 100 answers at list price. The claims and the scoring are the same ones behind our first citation test.

The drop came in one step

From minimal to medium, the rate slid from 13.3% to 10.9%. That slide is small enough to be chance. Then high cut it to 4.5%, less than half of medium.

The gap between minimal and high has about a 1 in 60 chance of showing up between two identical settings. Against low it is about 1 in 23. We call a result real at 1 in 20 or better, so high beats both. Against medium the gap falls just short, at about 1 in 15.

High also held up in every run. Its three runs came in at 4.4%, 7.0% and 2.2%. No minimal run was lower than 8.9%.

One bad run at minimal

At minimal, two runs came in at 8.9% and one at 22.2%, with the same claims and the same setting. The low level showed the same shape: 8.9%, 17.8%, 8.9%.

That is why we run every test three times. A single run at minimal could have told you Gemini makes up one citation in eleven, or nearly one in four. Three runs put it at 13.3%, with a range from 8.6% to 20.1%.

A model that swings that much between identical runs cannot be trusted on any one answer. Its average says little about the draft in front of you. That is why we check each claim against four models from different labs instead of trusting one.

High costs 2.2 times as much as minimal

More thinking costs money because the model writes more before it answers. At minimal, Gemini 3.7 Flash thought for about 551 tokens, or word pieces, per answer. At high it thought for about 1,148.

High cost $0.57 per 100 answers, about 2.2 times minimal. For that you get about a third of the made-up DOIs. On this task, that is the cheapest real gain in our whole effort test.

Which thinking level to use

This is for one job: asking Gemini to name sources for a claim, from memory. Coding and agent work are different tasks, and thinking may pay off there in other ways.

Thinking level for listing sources with Gemini 3.7 Flash

If you care most about Use Why
Fewer made-up DOIs High 4.5%, against 10.9% to 13.3% at the other three levels.
Cost Low $0.24 per 100 answers. Its rate, 11.9%, is within noise of minimal and medium.

Skip medium for this job. It cost more than low and gave no gain our test could see.

How we ran it

We used the same 30 claims as our earlier citation tests, word for word. We asked for up to three peer-reviewed sources per claim, with a DOI for each. The model had no web search, so every source came from memory.

The only thing we changed was the thinking level. We ran all four levels, three times each. That is 360 answers, run on 26 September 2026. We gave every answer room for 32,000 tokens, so the model was never cut off by us. None was.

A script then looked up every DOI in Crossref and DataCite. No AI model graded anything. A DOI that exists in neither registry counts as made up. Gemini turned down every invented claim except one, at medium, where it offered three sources. Two of them did not exist.

Thirty claims is a small test. Treat these rates as a direction, not a ranking. Thinking levels mean different things on different models, so compare levels within one model, never across two.

We ran the same test on two other vendors:

FAQ

Does more thinking make Gemini hallucinate less?

At the high level, yes. On our 30-claim citation test, Gemini 3.7 Flash made up 13.3% of its DOIs at minimal thinking, 11.9% at low, 10.9% at medium and 4.5% at high, across three runs. Only high separated clearly from the lower levels.

What thinking level should I use for Gemini 3.7 Flash?

For listing sources from memory, use high. It cut made-up DOIs to 4.5%. It cost $0.57 per 100 answers, about 2.2 times the cost of minimal. For coding or agent work, test your own tasks.

Is Gemini's medium thinking level worth it?

Not for citations. Medium made up 10.9% of its DOIs against low's 11.9%, a gap within noise, and it cost $0.40 per 100 answers against low's $0.24.

Why did Gemini's results change between runs?

Models do not give the same answer every time. At the minimal level, two of our runs came in at 8.9% and one at 22.2%, with nothing changed. That is why we ran every level three times and report the range.

How many Gemini citations are fake?

In our test, between 4.5% and 13.3% of the DOIs Gemini 3.7 Flash gave did not exist, depending on the thinking level. That is about one in twenty-two at high and one in seven at minimal. Check every DOI before you publish it.

Keep reading

Model Selection | 11 min read

Is There a Most Accurate AI Model?

The honest answer is no. The ranking changes with the task, the benchmark, and the month, and even the leader still hallucinates.

Model Selection | 12 min read

ChatGPT vs Claude vs Gemini: Which Is Most Accurate?

The three big plans cost the same twenty dollars, and every comparison of them reaches a verdict on accuracy without running a test. We ran one on all three, twice, a week apart. The order changed.

Model Selection | 11 min read

We Benchmarked the Models Behind Our Own Product

Six jobs in our product each pick an AI model. Exactly one of those choices had ever been tested. We built a deterministic benchmark. It ran 227 graded calls against a slate of four to five candidates. We changed five of the six. Two of the tables decided nothing. One of our own test labels turned out to be wrong. And the exercise found three bugs it was not looking for.

AI Reliability | 12 min read

Every AI Tier List Measures the Same Six Things. Truth Isn't One.

A developer who built a $200k business on AI ranked 11 companies and 17 models across six categories. It is a good ranking. But in 27 minutes the words hallucination, accurate and sycophancy never come up. And one model gets marked down, in so many words, for checking its work. So we ran the missing column ourselves: nine models, both failure modes, three times each. It separates, and it does not agree with any capability ranking.

Comparisons | 9 min read

TrueStandard vs FactCheckTool

These two tools look alike, but they solve opposite problems. One tells you if the media you read is fake. The other tells you if the draft you are about to publish is true.

Even at high, one in twenty-two was fake

The top thinking level cut Gemini's made-up DOIs by about two thirds. It did not cut them to zero. Paste your draft into TrueStandard. Four models from different labs check each claim and flag what none of them can back up, in about 60 seconds.

Check Your Draft →