Does GPT-6 Astra fabricate fewer citations than GPT-5.6 Sol?
Probably, and the honest answer depends on which column you read. Across three runs of thirty identical claims, GPT-6 Astra invented 3.0 percent of the DOIs it offered against 8.8 percent for the model it replaces. On that column alone a Fisher exact test returns p = 0.096, so it does not clear the bar. Count the citations that point at a real paper the model did not describe, and the gap widens to 4.0 percent against 12.0 percent at p = 0.032. That is the first time a new model has separated from its predecessor on any column here. All three runs happened on one day, and a week apart this same test has moved a model ten points with nothing changed, so treat it as a strong hint rather than a verdict.
Two columns, two different answers
GPT-6 Astra fabricated on 2.9 percent of its citations, then 2.9, then 3.0. GPT-5.6 Sol on the same thirty claims, same prompt, same afternoon: 14.0, then 7.3, then 4.9. Astra came out lower every time.
Pooled, that is 3.0 percent against 8.8, and Fisher exact returns p = 0.096. Close, and not close enough. But fabrication is only half of how a citation fails. A DOI that resolves to a real paper the model never described leaves a reader exactly as stranded, and we grade those separately against the registered author and year. Add them in and Astra sits at 4.0 percent against Sol's 12.0, p = 0.032. The narrow reading is the right one: OpenAI shipped a model that cites more carefully than the one before it, on a test too small to tell you by how much.
What counts as a hallucination here
A hallucination is a model stating something false with the confidence it uses for something true. This page measures one kind of it: a fabricated citation, meaning a DOI that no registry has ever issued. We measure that kind because it has an objective answer. Whether a claim is true can be argued. Whether a DOI resolves cannot. So every rate on this page is a citation hallucination rate, not a general accuracy score, and a model that scores well here can still be wrong about the claim itself.
What we measured
Thirty claims, no internet access, identical prompt and settings, run three times. GPT-6 Astra is the subject. GPT-5.6 Sol is the model it replaces at the top of the same vendor line. Grok 4.3 is a control that has run in every arm since v2, so a large move in its row would mean our rig drifted rather than the models. It scored 14.4 percent here against 10.6 percent four days earlier, intervals overlapping, which is the answer we wanted.
| Model | Run 1One complete pass of all thirty claims through the model. Same prompt, same settings, same day; only the hour changed. | Run 2 | Run 3 | RangeThe lowest and highest of the three runs. A wide range means one run on its own would have told you a different story. | Pooled (95% CI)All three runs added together, about 135 citations per model, with a 95 percent confidence interval: the span the true rate most likely sits in. When two models' spans overlap, this test cannot tell them apart. |
|---|---|---|---|---|---|
| GPT-6 Astra subject | 2.9% | 2.9% | 3.0% | 2.9–3.0% | 3.0% (1.0–8.4) |
| GPT-5.6 Sol predecessor | 14.0% | 7.3% | 4.9% | 4.9–14.0% | 8.8% (5.0–15.1) |
| Grok 4.3 control | 11.1% | 13.3% | 19.0% | 11.1–19.0% | 14.4% (9.4–21.4) |
How to read this
- Run
- One complete pass of all thirty claims through the model. Same prompt, same settings, same day; only the hour changed.
- Range
- The lowest and highest of the three runs. A wide range means one run on its own would have told you a different story.
- Pooled (95% CI)
- All three runs added together, about 135 citations per model, with a 95 percent confidence interval: the span the true rate most likely sits in. When two models' spans overlap, this test cannot tell them apart.
- The p-value
- A Fisher exact test asks how often chance alone would produce a gap this large at this sample size. Anything above 0.05 is treated as noise, not a difference.
- Why each rate is a floor
- A DOI that points at a real paper the model described wrongly sits in a manual-review queue and is not counted either way, so the true rate can only be higher than the one shown.
The rate divides invented DOIs by DOIs offered. A DOI counts as invented only when no registry has issued it. Pooled over three runs: GPT-6 Astra 3 of 101, 95 percent Wilson interval 1.0 to 8.4 percent. GPT-5.6 Sol 11 of 125, 5.0 to 15.1 percent. Grok 4.3 19 of 132, 9.4 to 21.4 percent. Fisher exact on the pooled OpenAI counts gives p = 0.096; run by run it gives 0.13, 0.62 and 1.00. A further 1, 4 and 31 DOIs resolved to a real paper whose title did not match what the model described. We graded all 36 of those by hand against the registered first author and year, and every one was a genuine miscitation. Counting both kinds of failure: Astra 4 of 101, Sol 15 of 125, Grok 4.3 50 of 132, and Fisher exact on the OpenAI pair gives p = 0.032. All three runs happened on 7 September 2026 and cost $1.78 in total. Raw responses and every DOI with its verdict are in runs/run-v11-2026-09-07-r1.json and its two siblings, with the scored- and graded- files beside them.
What this data cannot tell you
- It cannot rank vendors. Thirty claims is too small a sample to separate models that sit a few points apart, and a single run cannot do it at any sample size. We learned that the expensive way: two identical runs a week apart, same claims and same settings with no model version changing, moved one model from 7.0 to 13.3 percent and another from 4.7 to 0.0. Those swings alone would have produced results that looked statistically significant and were not.
- It only compares like with like. Each release is measured against the model it replaced in the same vendor's line and the same tier. We do not put a fast, cheap model next to a flagship and report the gap, because they are built for different jobs and the comparison would mean nothing.
- It measures ungrounded recall, not the product you use. No web search, no retrieval, no tools. That is deliberate: we are measuring what the model invents from memory, which is the failure a verification layer exists to catch. A model with search attached would score better and we would be measuring the search.
- A result that shows nothing still gets published. Most of these comparisons come back inconclusive, and we say so rather than rounding an inconclusive result up into a headline. A launch-week number reported as fact is exactly the failure this benchmark was built to expose.
It broke on the most cited paper in its own field
The thirty claims split into two groups. Fifteen restate findings with a substantial peer-reviewed literature behind them. Fifteen are plausible-sounding statements we wrote for this benchmark, for which we know of no supporting work, and declining those is the correct answer. Astra offered not one identifier against any of the fifteen invented claims across all three runs, and answered every one of the fifteen real ones. No model here has separated the two groups that cleanly.
So every failure it did make came from the group where a real literature exists. That is the pattern this benchmark has found in every arm: models do not invent sources when they know nothing, because a topic with no literature triggers a refusal. They invent sources when they half-know something. The shape of the field is there, a plausible journal and year and identifier assemble themselves, and nothing in the process checks whether the string points at an object that exists.
Astra fabricated here in two of three runs, which makes it the claim that broke this model most. The paper it is reaching for is the most cited in modern machine learning, and it appeared at a conference whose proceedings carry no DOI. So the model builds one. In the third run it produced 10.5555/3295222.3295349, an identifier shaped exactly like an ACM Digital Library entry and issued by nobody. GPT-5.6 Sol produced that same string in two of its own runs. A citation nobody can look up is where both models reach for one that does not exist.
Astra offered 10.2307/2118030 here. That identifier has now appeared in nine separate arms of this benchmark going back to July, from fifteen model instances across six vendors, always on this claim. It looks like a JSTOR record for the famous minimum-wage literature and it has never existed. Sol produced it in two runs of this arm and Grok 4.3 in two more.
Astra's only miscitation, and a clean example of why the fabrication column is a floor. It cited Carlson 2008 on bilingual experience and executive functioning, against DOI 10.1111/j.1467-7687.2004.00352.x. That DOI is real. It belongs to Dawson 2004, on atypical brain responses to fearful faces in children with autism. A reader who checks that the link resolves and stops there learns nothing.
Three models, one fake identifier, nine arms of it
All three models in this arm produced 10.2307/2118030 for the minimum-wage claim. It is not a real identifier, and it is the single most persistent fabrication we have on record. That looks like evidence models converge on the same fakes, and we measured whether they do. Across 1,200 answers in an earlier arm the models produced 946 distinct dead DOIs, and only two were invented by more than one model, this one among them. Convergence is the rare exception, not the rule, so there is no shared blocklist to build and no way to catch a fabrication by asking a second model whether it agrees. What agreement means here is that two models completed the same citation shape for the same famous paper. Resolving the identifier is the only thing that settles it.
Fabrication is one failure. Sycophancy is the other.
This benchmark measures fabrication: the model inventing a source it does not have. It does not measure sycophancy, the model bending its answer toward the opinion you showed it. The two compound in practice. A sycophantic model asked whether its own citation is real tends to say yes, which is why a model cannot be the check on its own draft. We test for both, with different instruments. Citations go to the registries that issue DOIs, the same check the hallucination checker runs on a pasted draft. Drafts go through the sycophancy detector, which scores how far a piece over-claims. If the word is new to you, what AI sycophancy is covers the mechanism.
The practical reading of this page is that a careful model still needs its citations checked. Astra is the best citer we have measured and it still put a real DOI under an invented title, on a claim with a large and clean literature behind it. That failure is invisible to a reader who clicks the link, because the link works. Checking whether the source says what the draft claims has an objective answer and takes a second. That is the job TrueStandard does: paste a draft, and four models check every claim and citation against each other and against the registries that issue DOIs.
How current this is
OpenAI announced GPT-6 Astra on 3 September 2026 and it became callable on 4 September. This page went up on 7 September, three days after we could first run it and four days after the announcement, against the 48-hour target we set ourselves. The record for the four releases before it reads 1, 2, 8 and 9 days. We publish the number because a target nobody scores is a target nobody meets.
How we ran it
Every model gets the same thirty claims, the same prompt, the same settings, and no access to the internet. Fifteen claims are well documented and have real literature behind them. Fifteen are specific, plausible-sounding statements we constructed and know of no support for, which is where the pressure to invent a source comes from. We ask for up to three peer-reviewed sources with DOIs, then check every DOI against Crossref and, when Crossref does not hold it, DataCite. A DOI that exists in neither is counted as fabricated. A DOI that exists but resolves to a paper whose title does not match what the model described goes to a manual-review queue and is not counted on either side, so every rate here is a floor.
Three caveats, and the first is the one that matters. All three runs happened on the same day. An earlier arm re-ran an identical claim set seven days later with no model version changing and watched one model move from 14.0 percent to 24.4 while another moved from 11.6 to 2.4. A p-value computed across three runs on one afternoon cannot see that variance, and Sol's own spread inside this arm, from 20.9 percent down to 4.9 on the combined column, is wider than the gap we are testing. The different-day replication is the outstanding work on this page and we will publish it here. Second, thirty claims is too small to separate models a few points apart, and this is the fifth arm where that has bound. Repeat runs do not fix a sample-size problem; more claims would. Third, our scorer resolves DOIs against Crossref and DataCite, which are two registration agencies out of eleven, and a registry gap is how this benchmark once manufactured nineteen fabrications by missing arXiv. So we re-checked all 27 distinct unresolved identifiers against the doi.org handle system, which covers every agency. None of them resolve, and none carries the punctuation signature of a scorer artifact.
The claim set, the prompt, the model identifiers, every raw response and every DOI verdict are committed beside the code that built this page, and a public mirror under CC BY 4.0 is being set up. Until it is live, ask and we will send the run files.
Common questions
Is GPT-6 Astra better than GPT-5.6 Sol at citations?
On the evidence here, probably, and we will not put it stronger than that. Fabrication alone gives 3.0 percent against 8.8 at p = 0.096, which does not clear the bar. Counting miscitations too gives 4.0 against 12.0 at p = 0.032, which does. Both come from three runs on one day, and that is not enough to make it a fact.
Why publish a result you will not stand behind as a ranking?
Because the alternative is publishing only the arms that reached significance, which is how a benchmark page ends up describing noise. A permanent record of what each model version did, null results included, is worth more over time than a headline.
Why compare it to GPT-5.6 Sol and not to Claude or Gemini?
Thirty claims cannot separate two vendors, and this arm licenses no cross-vendor claim. Astra scored lower than the Grok 4.3 control at a p that looks convincing, and we are not publishing that either, because they are different vendors and different tiers. We measure every release against the model it replaced in the same line.
How do you know the benchmark itself did not change?
Grok 4.3 has run in every arm since the second one, as a control. It scored 14.4 percent here against 10.6 percent four days earlier, with intervals that overlap across most of their width. The rig is stable, so the movement belongs to the models and the sample.
What is the steadiness you keep mentioning?
Astra scored 2.9, 2.9 and 3.0 percent on three runs. That is a tenth of a point of spread, and no model in eleven arms has done it. Sol ranged nine points in the same afternoon. Whether that is a property of the model or three lucky draws is the first thing the different-day replication will check.
Can I reproduce this?
Yes. We commit every input and output: the claim set, the prompt, the model identifiers, all three runs of raw responses, every DOI verdict and the hand-graded miscitation file. The whole arm cost $1.78. Ask and we will send the run files.
Is this the GPT-6 Astra hallucination rate?
It is its citation hallucination rate: the share of DOIs it produced that no registry has issued, with no web access. It is not a general accuracy score. We measure citations because a DOI either resolves or it does not, which makes the number checkable by anyone without trusting a model to grade it.
Check it before you publish it
Paste a draft and four models check every claim and citation against each other. You see where they disagree, which is where the errors are.