AI Reliability

Does Claude Hallucinate Less at Higher Effort?

We asked Sonnet 5 and Opus 5.5 for sources on the same 30 claims at three effort levels each. Then we checked every DOI they gave us against the real registries.

Does Claude Hallucinate Less at Higher Effort?

Claude Sonnet 5 made up 14.7% of the citations it gave at low effort and 3.4% at max. Claude Opus 5.5 made up none, at low, medium or max. More effort helped the model that had room to improve. It did nothing for the one that was already clean.

Anthropic released Opus 5.5 on 22 September with five effort levels and a new default of medium. Anthropic's own post on effort shows it paying off on coding tasks. We wanted to know if it pays off on a different job: naming real sources for a claim, from memory, without search.

Effort helped Sonnet 5 and did nothing for Opus 5.5

We scored each DOI, the permanent ID printed on a research paper. Each row is 90 answers: 30 claims, asked three times. Half the claims have real research behind them. We made up the other half, so no real paper supports them.

Claude at three effort levels, 30 claims, three runs, September 2026

Model and effort DOIs given Made up Rate 95% range Cost per 100 answers
Sonnet 5, low 129 19 14.7% 9.6 to 21.9% $0.20
Sonnet 5, high 121 11 9.1% 5.2 to 15.5% $0.52
Sonnet 5, max 119 4 3.4% 1.3 to 8.3% $4.64
Opus 5.5, low 153 0 0.0% 0.0 to 2.4% $1.01
Opus 5.5, medium 160 0 0.0% 0.0 to 2.3% $1.24
Opus 5.5, max 153 0 0.0% 0.0 to 2.4% $8.80

Made up means the DOI exists in neither Crossref nor DataCite, the two registries that issue DOIs. The 95% range is where the true rate likely sits, given how many DOIs we saw. Cost is what we paid per 100 answers at list price. The claims and the scoring are the same ones behind our first citation test.

Effort cut Sonnet 5's fake citations by three quarters

Sonnet 5 went from 19 made-up DOIs at low effort to 4 at max. The chance of seeing a gap this big between two identical models is about 1 in 500.

It also held in every run. At low effort the three runs ranged from 9.5% to 18.6%. At max they ranged from 2.4% to 5.1%. The worst max run beat the best low run.

High effort sat in between, at 9.1%. So each step up helped, and the last step helped most.

Opus 5.5 made up nothing at any level

Opus 5.5 gave us 466 DOIs across all three levels. Every one of them is real, so more effort had nothing to fix.

It also turned down every claim we made up, at every level. On 43 of those 135 answers it went one step further. It said no source supports the claim, then listed real papers on a nearby topic and said they do not support it either.

I can't provide sources for this claim. I'm not aware of any 2022 meta-analysis, or any peer-reviewed study, comparing oxytocin levels when reading fiction on paper versus a tablet.

Opus 5.5 at medium effort, on one of the claims we invented.

For a claim with no research behind it, that is the correct answer. A simple count of DOIs would score it as a wrong citation, so we read each of these answers by hand before scoring them.

A 0% rate on 30 claims is not a promise about your 31st. Opus 5.5 was clean on this test; the next claim may be the one it gets wrong. That is why we check each claim against four models from different labs instead of trusting one, however good it looks this week.

The cheapest clean setting

Opus 5.5 at low effort cost $1.01 per 100 answers and made up nothing. Sonnet 5 at max cost $4.64 per 100 answers and still made up 3.4%.

So on this task the bigger model at its lowest setting was about 4.6 times cheaper than the smaller model at its highest. It was also cleaner. Max effort on Opus 5.5 cost $8.80 per 100 answers and bought nothing we could measure.

Opus costs more per token than Sonnet, so that result sounds backwards. It happens because at max effort Sonnet 5 thinks for about 1,900 tokens, or word pieces, before a short list. At low effort Opus 5.5 thinks for about 260.

What went wrong at max effort

Max effort failed in a way the lower levels never did. Twice, Sonnet 5 spent its whole 32,000-token budget thinking and never wrote an answer. Both times it was asked about a claim with plenty of real research behind it.

Once, Opus 5.5 at max had its answer blocked by a filter after about 1,900 tokens of thinking. Low and medium never did either.

We left all three answers out of the table and did not re-run them. Say both empty Sonnet answers had held three made-up DOIs each. Sonnet 5 at max would then sit at 8.0%, still well under its 14.7% at low.

Which effort level to use

This is for one job: asking Claude to name sources for a claim, from memory. Coding and agent work are different tasks, and effort may pay off there in other ways.

Effort level for listing sources

Model Use Why
Sonnet 5 High or max Max cut made-up DOIs from 14.7% to 3.4%. High is a cheaper middle step at 9.1%.
Opus 5.5 Low Zero made-up DOIs at every level. Medium and max cost more and changed nothing here.

Neither setting means you can skip the check. Claude Fable 5 tops a well-known hallucination benchmark, and that score still says more about accuracy than about how often it makes things up.

How we ran it

We used the same 30 claims as our earlier citation tests, word for word. We asked for up to three peer-reviewed sources per claim, with a DOI for each. The models had no web search, so every source came from memory.

The only thing we changed was the effort setting. We ran each model at three of its five levels, three times each. That is 540 answers, run on 26 September 2026. We gave every answer room for 32,000 tokens, so a model that thought hard was never cut off by us.

A script then looked up every DOI in Crossref and DataCite. No AI model graded anything. A DOI that exists in neither registry counts as made up.

Thirty claims is a small test. Treat these rates as a direction, not a ranking. Effort levels also mean different things on different models, so compare levels within one model, never across two. We have tested GPT and Gemini the same way, and they get their own posts.

FAQ

Does higher effort make Claude hallucinate less?

For Sonnet 5, yes. On our 30-claim citation test it made up 14.7% of its DOIs at low effort, 9.1% at high and 3.4% at max, across three runs. For Opus 5.5, effort made no difference: it made up none of 466 DOIs at low, medium or max.

What effort level should I use for Claude Opus 5.5?

For listing sources from memory, low was enough in our test. Opus 5.5 made up zero DOIs at low, medium and max. Low cost $1.01 per 100 answers and max cost $8.80. Coding and agent work may reward higher effort, so test those tasks on their own.

What is the default effort for Claude Opus 5.5?

Medium. On Opus 5, a request with no effort set ran at high. On Opus 5.5 it runs at medium, one level lower. In our test medium was as clean as the other two levels we ran, with no made-up DOIs.

Is Opus 5.5 cheaper than Sonnet 5 for citations?

At the settings that made up the fewest DOIs, yes. Opus 5.5 at low cost $1.01 per 100 answers and made up none. Sonnet 5 at max cost $4.64 and made up 3.4%. Sonnet at max thinks for about 1,900 tokens per answer, and that outweighs its lower price per token.

Can max effort make Claude worse?

It can fail in new ways. Twice in 90 answers, Sonnet 5 at max thought for its whole 32,000-token budget and returned nothing. Opus 5.5 at max had one answer blocked by a filter. Neither happened at lower levels in our test.

Does Opus 5.5 still make up citations?

Not in this test. It gave 466 DOIs across three effort levels and all of them were real. It also refused every claim we made up. Thirty claims is a small sample, so a clean result here does not mean it never makes one up.

Keep reading

Model Selection | 11 min read

Is There a Most Accurate AI Model?

The honest answer is no. The ranking changes with the task, the benchmark, and the month, and even the leader still hallucinates.

Model Selection | 12 min read

ChatGPT vs Claude vs Gemini: Which Is Most Accurate?

The three big plans cost the same twenty dollars, and every comparison of them reaches a verdict on accuracy without running a test. We ran one on all three, twice, a week apart. The order changed.

Comparisons | 9 min read

TrueStandard vs Parafact

Both verify claims before you publish. The real difference is what one model can miss, and whether your long-form draft fits inside the check at all.

Model Selection | 11 min read

We Benchmarked the Models Behind Our Own Product

Six jobs in our product each pick an AI model. Exactly one of those choices had ever been tested. We built a deterministic benchmark. It ran 227 graded calls against a slate of four to five candidates. We changed five of the six. Two of the tables decided nothing. One of our own test labels turned out to be wrong. And the exercise found three bugs it was not looking for.

AI Reliability | 12 min read

Every AI Tier List Measures the Same Six Things. Truth Isn't One.

A developer who built a $200k business on AI ranked 11 companies and 17 models across six categories. It is a good ranking. But in 27 minutes the words hallucination, accurate and sycophancy never come up. And one model gets marked down, in so many words, for checking its work. So we ran the missing column ourselves: nine models, both failure modes, three times each. It separates, and it does not agree with any capability ranking.

Even the clean models get checked

Opus 5.5 was clean on 30 claims. Sonnet 5 still made up one DOI in thirty at its best setting. Paste your draft into TrueStandard. Four models from different labs check each claim and flag what none of them can back up, in about 60 seconds.

Check Your Draft →