Which AI Uses the Most Em Dashes? Six Models Counted
We gave six models from Anthropic, OpenAI and Google the same 35 writing jobs, three times each. Then we counted every em dash, and 22 other tells, in all 629 drafts.
Read articleGuides on AI verification, model selection, and working with artificial intelligence.
We gave six models from Anthropic, OpenAI and Google the same 35 writing jobs, three times each. Then we counted every em dash, and 22 other tells, in all 629 drafts.
Read articleWe asked four models from three labs for sources on the same 30 claims, at each effort level. Then we checked every DOI against the real registries.
Read articleWe asked Gemini 3.7 Flash for sources on the same 30 claims at all four of its thinking levels. Then we checked every DOI it gave us against the real registries.
Read articleWe asked GPT-5.6 Luna for sources on the same 30 claims at all five of its effort levels. Then we checked every DOI it gave us against the real registries.
Read articleWe asked Sonnet 5 and Opus 5.5 for sources on the same 30 claims at three effort levels each. Then we checked every DOI they gave us against the real registries.
Read articleThe per-judgment price said we would cut the bill by a quarter. Measured against the real pipeline it saved nothing at all, and the reason is more useful than the saving would have been.
Read articleWe handed TypeSafe's decision model thirteen factual claims with nothing to check them against. It refused to guess on twelve of them, which is the best thing we can say about it.
Read articleWe measured the same model two ways and got 1.7x and 100x. Both numbers are honest. What separates them is the thing you put on the other side.
Read articleEvery write-up of TypeSafe's new model repeats the same line about calibration. We checked it against two cheap chat models on 108 claims, and the advantage sits somewhere else.
Read articleFour AI judges scored an answer with dead sources about as highly as the same answer with real ones.
Read articleSame question, raw API calls, no app in between. The first four split 2 to 2, the next four split 2 to 2, then the two newest models agreed.
Read articleThe cheapest model on the price list stopped being the cheapest model on the invoice.
Read articleEvery impressive result carries a condition, and the conditions are where the story sits.
Read articleSix prompting practices, and two numbers that mean close to the opposite of how they read.
Read articleRecall is easy to advertise, because a tool that flags everything scores 100 percent on it. The number that decides whether a checker is usable is how often it flags something true. We measured ours on 30 labelled claims, and the more useful result was what the test could not tell us.
Read articleTake one generation of models: published hallucination rates for it run from under 2 percent to over 60. The standard explanation is that different labs test different things, and that is true. Almost nobody has tested it directly, so we did. One model, one prompt, one scorer, three runs a side, and only the questions changed.
Read articleTwo models posted the lowest sycophancy scores in our field, and both had stopped answering. The scorer could not tell the difference.
Read articleIt leads AA-Omniscience with a score of 40 — the highest recorded. Artificial Analysis says that score comes from accuracy, not from low hallucination. And 9% of its answers came from a different model.
Read articleClaude marks the words it chooses, and if you wrote them, there is almost nothing to mark.
Read articleA developer who built a $200k business on AI ranked 11 companies and 17 models across six categories. It is a good ranking. But in 27 minutes the words hallucination, accurate and sycophancy never come up. And one model gets marked down, in so many words, for checking its work. So we ran the missing column ourselves: nine models, both failure modes, three times each. It separates, and it does not agree with any capability ranking.
Read articleSix jobs in our product each pick an AI model. Exactly one of those choices had ever been tested. We built a deterministic benchmark. It ran 227 graded calls against a slate of four to five candidates. We changed five of the six. Two of the tables decided nothing. One of our own test labels turned out to be wrong. And the exercise found three bugs it was not looking for.
Read articleThe three big plans cost the same twenty dollars, and every comparison of them reaches a verdict on accuracy without running a test. We ran one on all three, twice, a week apart. The order changed.
Read articleWe asked both generations for peer-reviewed sources on the same 30 claims, on the same day. We resolved each DOI they produced against a real registry. GPT-5.6 came out lower. Then we noticed something in the data that matters more than which model won.
Read articleGraph engineering says the checker should be a separate job, and never a separate model. In practice the skeptic node runs on the one that wrote the answer.
Read articleGrok 4.5 shipped on 8 July, and our benchmark ran three weeks later. It tested Grok 4.3, so the number we published was already a generation stale. We fixed that: both models went through the same 30 claims on the same day, and the rate dropped from 11.1% to 4.4%. Here is why we are not calling that a win.
Read articleWe asked four frontier models for peer-reviewed sources on 30 claims. Then we resolved every DOI they gave us against Crossref and DataCite, and 21 of 146 did not exist. The fabrication rate varies three and a half times between vendors. And every single fake landed on a claim that has real literature behind it.
Read articleA sourced reference of the 2026 hallucination-rate numbers. What each benchmark measured. Which models did best and worst. And where the widely quoted figures get misread. Built to be linked and kept current as new data lands.
Read articleIn 2026, the strongest AI systems stopped trusting one model. They send each task to a different one. That choice has a plain lesson for anyone who publishes AI-assisted work.
Read articleDeep research tools are fast, fluent, and wrong about their sources more often than not. Here is what the audits found, and the two-part fix.
Read articleA model carries the same blind spots into review that it had while writing. Dressing it up as a critic is a costume, not a second mind.
Read articleBoth run several models, and both show where those models agree. The difference is the moment each one is built for. Omniscient checks what you are reading. TrueStandard checks what you are about to publish.
Read articleA running list of documented cases where AI made up citations in law, academia, and media. Every entry is sourced, and it is here to be linked, cited, and updated as new cases surface.
Read articleThey flag honest writers. They mark famous human documents as AI, and OpenAI quietly killed its own detector. But the deeper problem is not that detection is unreliable. It is that 'was this written by AI?' was never the question that protects you.
Read articleA six-step workflow that separates drafting from verification. And a clear line between editing, which you already do, and fact-checking, which you probably skip.
Read articleIt catches citations that don't exist. By GPTZero's own admission, it does not check whether what you wrote is true. Here is exactly what its hallucination and source tools do. Here is where the gap is, and what closes it.
Read articleAI does not just get facts wrong. It invents whole sources: cases, studies, DOIs. Then it cites them with the same confidence it uses for real ones. Here is why it happens, and the disasters it has already caused. And here is how to catch a fake citation before your name is on it.
Read articleAccurate enough to trust for everyday questions. Wrong often enough to get you sued if you publish it unchecked. Here is what the measurements actually say, and what to do about it.
Read articleDELEGATE 52, GPT-5.5, and a Purdue impossibility proof. Three April 2026 results that move 'hallucinations are structural' from take to documented fact.
Read articleFour checks catch a fabricated reference before your readers do, and one of them is new. Here is the change: in 2026, a DOI that resolves no longer means the citation is real.
Read articleResearchers found AI made experts measurably worse on hard tasks. Here is when to trust ChatGPT, and when it is just telling you what you want to hear.
Read articleOne asks who wrote this. The other asks is this true. Before you publish, only one of those questions protects your name, and most teams are watching the wrong one.
Read articleThree different problems hurt real creators when AI is involved. Identity attestation, AI detection, and claim verification each need a different tool.
Read articleIf ChatGPT wrote the draft, can Claude safely verify it? Sometimes helpful, not sufficient by default. The reason is what these models share, not what they don't.
Read articleSlop and AI-assisted work can look identical on the page. The line between them is whether you verified the output, and whether you can prove it.
Read articleThe honest answer is no. The ranking changes with the task, the benchmark, and the month, and even the leader still hallucinates.
Read articleSix patterns cover almost every agent you'll build. Five are routine. The sixth, verification, breaks when you wire it with a single model, and most teams wire it that way.
Read articleModels sound certain every time, even when wrong, and the confident tone you trust in people is worthless here. Here is the fix.
Read articleOriginality.AI is a strong AI-detection suite. But say the job you care about is verifying claims before you publish. Fact-checking is only one of its five bundled checks, and it runs on a single model.
Read articleBoth verify claims before you publish. The real difference is what one model can miss, and whether your long-form draft fits inside the check at all.
Read articleDrafting got faster and verification did not. The work didn't disappear: it moved to the step right before your name goes on it.
Read articleThese two tools look alike, but they solve opposite problems. One tells you if the media you read is fake. The other tells you if the draft you are about to publish is true.
Read articleSolo operators ship AI-assisted content under deadline with no editor. The math only works if subscribers trust you. Here is what newsletter operators need to verify before send.
Read articleThree states proposed or enforced 'independent verification' for AI work in 30 days. Here is what 'independent' actually requires.
Read articleAI builders use both terms as if they meant the same thing. They are different architectures with different strengths. The difference matters most for the one job neither term sells: catching AI errors before you publish.
Read articleYour AI sometimes makes things up and sounds completely confident doing it. Anthropic explains why hallucinations happen and what you can do about them.
Read articleIn April 2026, top AI builders ran real experiments instead of demos. The results were more interesting than the demos. Here is what each test reveals, and why none of them fully answers the question writers care about.
Read articleThree things just changed about how AI handles your documents. Here is what works for content teams, and why better retrieval still does not mean better truth.
Read articleYour AI agrees with you too much. Anthropic's safeguards team explains why models tell you what you want to hear, and what you can do about it.
Read articleIn six weeks, Andrej Karpathy and AI builders shipped three viral reliability methods. Each is real and useful. None of them solves the checking problem for writers.
Read articleAnthropic just shipped a feature that quietly admits there is no single best Claude model. Here is how writers and content teams should actually pick.
Read articleFrom large language models to coding agents: what each type of AI does, which tools lead each group, and how to pick the right one for your work.
Read article