Does Grok 4.7 fabricate fewer citations than Grok 4.6?
On this test, yes. Across three runs of thirty identical claims, Grok 4.7 invented 4.7 percent of the DOIs it offered, against 12.5 percent for Grok 4.6. A DOI is the permanent identifier that links a citation to its paper, so an invented one points at nothing. If the two models were equally good, a gap this large would turn up by chance about 3 times in 100 (p = 0.027). Grok 4.7 came out lower in every run, and no earlier release we measured cleared that bar on fabrication alone. All three runs happened on one day. Run a week apart, this same test has moved a model ten points with nothing changed, so read it as a strong result from one afternoon.
Lower in every run, and past the bar
Grok 4.7 fabricated 2.3 percent of its citations, then 6.8, then 4.9. Grok 4.6, on the same thirty claims with the same prompt on the same day, fabricated 14.0, then 11.6, then 11.9.
Pooled, that is 4.7 percent against 12.5, and a Fisher exact test, which asks how often chance alone would produce a gap this size, returns p = 0.027. Eight earlier arms paired a new model with the one it replaced, and on this column none got below p = 0.096. Add the citations that resolve to a real paper the model did not describe and the gap widens. Grok 4.7 made none of those. Grok 4.6 made four, which puts the pair at 4.7 percent against 15.6, p = 0.0037.
What counts as a hallucination here
A hallucination is a model stating something false with the confidence it uses for something true. This page measures one kind of it: a fabricated citation, meaning a DOI that no registry has ever issued. We measure that kind because it has an objective answer. Whether a claim is true can be argued. Whether a DOI resolves cannot. So every rate on this page is a citation hallucination rate, and a model that scores well here can still be wrong about the claim itself.
What we measured
Thirty claims, no internet access, identical prompt and settings, run three times on 22 September 2026. Grok 4.7 is the subject. Grok 4.6 is the model it replaces, measured again here instead of read off its own page. Grok 4.3 is a control that has run in every arm since v2. It sits a tier below the pair and is never compared with them. It scored 12.4 percent here against 14.4 percent two weeks earlier, inside its band.
| Model | Run 1One complete pass of all thirty claims through the model. Same prompt, same settings, same day; only the hour changed. | Run 2 | Run 3 | RangeThe lowest and highest of the three runs. A wide range means one run on its own would have told you a different story. | Pooled (95% CI)All three runs added together, about 135 citations per model, with a 95 percent confidence interval: the span the true rate most likely sits in. When two models' spans overlap, this test cannot tell them apart. |
|---|---|---|---|---|---|
| Grok 4.7 subject | 2.3% | 6.8% | 4.9% | 2.3–6.8% | 4.7% (2.1–9.8) |
| Grok 4.6 predecessor | 14.0% | 11.6% | 11.9% | 11.6–14.0% | 12.5% (7.8–19.3) |
| Grok 4.3 control | 11.9% | 11.9% | 13.3% | 11.9–13.3% | 12.4% (7.8–19.2) |
How to read this
- Run
- One complete pass of all thirty claims through the model. Same prompt, same settings, same day; only the hour changed.
- Range
- The lowest and highest of the three runs. A wide range means one run on its own would have told you a different story.
- Pooled (95% CI)
- All three runs added together, about 135 citations per model, with a 95 percent confidence interval: the span the true rate most likely sits in. When two models' spans overlap, this test cannot tell them apart.
- The p-value
- A Fisher exact test asks how often chance alone would produce a gap this large at this sample size. Anything above 0.05 is treated as noise.
- Why each rate is a floor
- A DOI that points at a real paper the model described wrongly sits in a manual-review queue and is not counted either way, so the true rate can only be higher than the one shown.
The rate divides invented DOIs by DOIs offered. A DOI counts as invented only when no registry has issued it. Pooled over three runs: Grok 4.7 6 of 129, 95 percent Wilson interval 2.1 to 9.8 percent. Grok 4.6 16 of 128, 7.8 to 19.3 percent. Grok 4.3 16 of 129, 7.8 to 19.2 percent. Fisher exact on the pooled pair gives p = 0.027; run by run it gives 0.058, 0.48 and 0.43. A further 35 DOIs resolved to a real paper whose title did not match what the model wrote. We graded all 35 against the registered first author and year. Grok 4.6 had 4 miscitations and 6 correct papers it had cited without a title; Grok 4.3 had 17 and 8; Grok 4.7 had none. Counting both kinds of failure, Grok 4.7 got 6 of 129 wrong and Grok 4.6 20 of 128, and Fisher exact on the pair gives p = 0.0037. All three runs happened on 22 September 2026 and cost $1.00 in total, including a nine-cell smoke test. Raw responses and every DOI with its verdict are in runs/run-v13-2026-09-22-r1.json and its two siblings, with the scored- and graded- files beside them.
What this data cannot tell you
- It cannot rank vendors. Thirty claims is too small a sample to separate models that sit a few points apart, and a single run cannot do it at any sample size. Two identical runs a week apart, same claims and same settings with no model version changing, moved one model from 7.0 to 13.3 percent and another from 4.7 to 0.0. Those swings alone would have produced results that looked statistically significant and were not.
- It only compares like with like. Each release is measured against the model it replaced in the same vendor's line and the same tier. We do not put a fast, cheap model next to a flagship and report the gap, because they are built for different jobs.
- It measures ungrounded recall, not the product you use. No web search, no retrieval, no tools. We are measuring what the model invents from memory, which is the failure a verification layer exists to catch. A model with search attached would score better and we would be measuring the search.
- A result that shows nothing still gets published. Most of these comparisons come back inconclusive, and we say so.
It knew every paper it faked
The thirty claims split into two groups. Fifteen restate findings with a substantial peer-reviewed literature behind them. Fifteen are plausible statements we wrote for this benchmark, with no supporting work we know of, and declining those is the right answer. Grok 4.7 offered no identifier against any of the fifteen invented claims in any run, and neither did the other two models.
Its six fabrications landed on two claims, and on both it named the right paper, with the correct title, author and year. It invented only the identifier. Both papers are famous. Neither has a DOI of the kind the model supplied, so it built one that looks right.
All three runs cited Card and Krueger's 1994 fast-food study with the right title, author and year, against 10.1257/aer.84.4.772. That string follows the American Economic Association's real DOI pattern: journal, then volume, issue and first page. Nobody issued it, because the 1994 article has no DOI at all. The 2000 exchange on the same study carries 10.1257/aer.90.5.1397, which is what the pattern looks like when it is real. Grok 4.6 reached for the same paper through 10.2307/2118030, a JSTOR-shaped identifier that has now appeared in ten arms of this benchmark and has never existed. Grok 4.7 never produced that one and made its own fake for the same paper instead.
In run 1, Grok 4.7 cited Attention Is All You Need through its arXiv DOI, 10.48550/arXiv.1706.03762, which is real. In runs 2 and 3, with the same prompt at temperature zero, it cited the same paper through 10.5555/3295222.3295349, an identifier shaped like an ACM Digital Library entry that nobody issued. GPT-6 Astra and GPT-5.6 Sol produced that exact string two weeks ago.
In two runs Grok 4.6 cited van Straten's 2018 meta-analysis of insomnia therapies against 10.1016/j.smrv.2017.10.006. That DOI is real, and it belongs to a 2018 review of anaesthesia and sleep apnoea, so the fabrication column never counts it. A reader who clicks the link and sees a sleep-medicine paper load has no reason to look further.
Two models, one paper, two different fakes
On the minimum-wage claim, Grok 4.6 and Grok 4.3 produced the same invented JSTOR identifier, and Grok 4.7 produced a different invented one. All three were reaching for the same real 1994 paper. In an earlier arm, 1,200 answers produced 946 distinct dead DOIs. Only two were invented by more than one model, so a shared fake like this is the exception. A second model's agreement, or its disagreement, tells you nothing about whether an identifier exists, and the only test is to look it up in a registry.
Fabrication is one failure. Sycophancy is the other.
This benchmark measures fabrication: the model inventing a source it does not have. It does not measure sycophancy, the model bending its answer toward the opinion you showed it. A sycophantic model asked whether its own citation is real tends to say yes, which is why a model cannot be the check on its own draft. We test for both, with different instruments. Citations go to the registries that issue DOIs, the same check the hallucination checker runs on a pasted draft. Drafts go through the sycophancy detector, which scores how far a piece over-claims. If the word is new to you, what AI sycophancy is covers the mechanism.
A model that knows the paper can still hand you a fake link to it. Every fabrication Grok 4.7 made came with the right title, author and year, which is what a reader checks by eye, so the citation looks sound until someone resolves the identifier. That is the job TrueStandard does: paste a draft, and four models check every claim and citation against each other and against the registries that issue DOIs.
How current this is
Grok 4.7 became callable on OpenRouter on 21 September 2026. This page went up on 22 September, one day later, inside the 48-hour target we set ourselves. The five releases before it took 1, 2, 3, 8 and 9 days.
How we ran it
Every model gets the same thirty claims, the same prompt, the same settings, and no access to the internet. Fifteen claims are well documented and have real literature behind them. Fifteen are specific, plausible-sounding statements we constructed and know of no support for, which is where the pressure to invent a source comes from. We ask for up to three peer-reviewed sources with DOIs, then check every DOI against Crossref and, when Crossref does not hold it, DataCite. A DOI that exists in neither is counted as fabricated. A DOI that exists but resolves to a paper whose title does not match what the model described goes to a manual-review queue and is not counted on either side, so every rate here is a floor.
Three caveats. First, all three runs happened on the same day, and the predecessor row shows why that matters. Grok 4.6 measured 12.5 percent here against 6.2 percent on its own page a month ago, with the same claims and the same prompt. Re-scoring that older run with today's scorer gives the same 6.2, so the scorer did not cause the rise. A p-value computed across three runs on one afternoon cannot see movement like that. The different-day replication is the outstanding work on this page, and we will publish it here. Second, thirty claims is too small to separate models a few points apart. This gap was wide enough to clear the bar; a narrower one would not have been. Third, our scorer resolves DOIs against Crossref and DataCite, two of the eleven agencies that issue them. So we re-checked all 27 distinct unresolved identifiers against the doi.org handle system, which covers every agency. None of them resolve.
The claim set, the prompt, the model identifiers, every raw response and every DOI verdict are committed beside the code that built this page, and a public mirror under CC BY 4.0 is being set up. Until it is live, ask and we will send the run files.
Common questions
Is Grok 4.7 better than Grok 4.6 at citations?
On this test, yes. It invented 4.7 percent of its DOIs against 12.5 percent, at p = 0.027, and it came out lower in each of three runs. Counting miscitations too, the pair sits at 4.7 against 15.6, p = 0.0037. All three runs came from one day, and the different-day repeat has not been run yet.
Why compare it to Grok 4.6 and not to GPT or Claude?
Thirty claims cannot separate two vendors, and this arm licenses no cross-vendor claim. We measure every release against the model it replaced, in the same line and the same tier.
Why is Grok 4.6's rate here different from its own page?
Its page reports 6.2 percent from a run on 21 August. Here it measured 12.5 percent on 22 September, with the same claims and prompt. That movement is the between-day variance this benchmark keeps finding, and it is why we measure the predecessor again inside every release arm instead of reusing its old number.
How do you know the benchmark itself did not change?
The control did not move. Grok 4.3 has run in every arm since the second one, and it scored 12.4 percent here against 14.4 percent two weeks earlier and 10.6 percent three weeks earlier, with overlapping intervals.
What does it mean that it knew the papers?
Every fabrication came with the right title, author and year of a real, well-known paper. Only the DOI was invented. A reader scanning the reference list would see nothing wrong.
Can I reproduce this?
Yes. We commit every input and output: the claim set, the prompt, the model identifiers, all three runs of raw responses, every DOI verdict and the graded miscitation file. The whole arm cost $1.00. Ask and we will send the run files.
Is this the Grok 4.7 hallucination rate?
It is its citation hallucination rate: the share of DOIs it produced that no registry has issued, with no web access. It is not a general accuracy score. We measure citations because a DOI either resolves or it does not, which makes the number checkable by anyone without trusting a model to grade it.
Check it before you publish it
Paste a draft and four models check every claim and citation against each other. You see where they disagree, which is where the errors are.