Blog

Guides on AI verification, model selection, and working with artificial intelligence.

AI Reliability | | 9 min read

We Put Jev in Our Pipeline and Saved Nothing

The per-judgment price said we would cut the bill by a quarter. Measured against the real pipeline it saved nothing at all, and the reason is more useful than the saving would have been.

Read article
AI Reliability | | 9 min read

Is Jev Really 193x Faster?

We measured the same model two ways and got 1.7x and 100x. Both numbers are honest. What separates them is the thing you put on the other side.

Read article
AI Reliability | | 10 min read

Jev's Accuracy, Measured on 108 Claims

Every write-up of TypeSafe's new model repeats the same line about calibration. We checked it against two cheap chat models on 108 claims, and the advantage sits somewhere else.

Read article
AI Verification | | 10 min read

The Number AI Fact-Checkers Do Not Publish

Recall is easy to advertise, because a tool that flags everything scores 100 percent on it. The number that decides whether a checker is usable is how often it flags something true. We measured ours on 30 labelled claims, and the more useful result was what the test could not tell us.

Read article
AI Reliability | | 11 min read

The Same Model Fabricated at 10.9% and 30.9%

Take one generation of models: published hallucination rates for it run from under 2 percent to over 60. The standard explanation is that different labs test different things, and that is true. Almost nobody has tested it directly, so we did. One model, one prompt, one scorer, three runs a side, and only the questions changed.

Read article
AI Reliability | | 9 min read

A Quiet Model Wins Your Benchmark

Two models posted the lowest sycophancy scores in our field, and both had stopped answering. The scorer could not tell the difference.

Read article
AI Reliability | | 12 min read

Every AI Tier List Measures the Same Six Things. Truth Isn't One.

A developer who built a $200k business on AI ranked 11 companies and 17 models across six categories. It is a good ranking. But in 27 minutes the words hallucination, accurate and sycophancy never come up. And one model gets marked down, in so many words, for checking its work. So we ran the missing column ourselves: nine models, both failure modes, three times each. It separates, and it does not agree with any capability ranking.

Read article
Model Selection | | 11 min read

We Benchmarked the Models Behind Our Own Product

Six jobs in our product each pick an AI model. Exactly one of those choices had ever been tested. We built a deterministic benchmark. It ran 227 graded calls against a slate of four to five candidates. We changed five of the six. Two of the tables decided nothing. One of our own test labels turned out to be wrong. And the exercise found three bugs it was not looking for.

Read article
AI Reliability | | 10 min read

Does GPT-5.6 Fabricate Fewer Citations Than GPT-5.5?

We asked both generations for peer-reviewed sources on the same 30 claims, on the same day. We resolved each DOI they produced against a real registry. GPT-5.6 came out lower. Then we noticed something in the data that matters more than which model won.

Read article
AI Reliability | | 9 min read

Does Grok 4.5 Fabricate Fewer Citations Than 4.3?

Grok 4.5 shipped on 8 July, and our benchmark ran three weeks later. It tested Grok 4.3, so the number we published was already a generation stale. We fixed that: both models went through the same 30 claims on the same day, and the rate dropped from 11.1% to 4.4%. Here is why we are not calling that a win.

Read article
AI Reliability | | 12 min read

Which AI Fabricates Citations? We Tested Four Models

We asked four frontier models for peer-reviewed sources on 30 claims. Then we resolved every DOI they gave us against Crossref and DataCite, and 21 of 146 did not exist. The fabrication rate varies three and a half times between vendors. And every single fake landed on a claim that has real literature behind it.

Read article
AI Architecture | | 12 min read

The Best AI Uses Many Models

In 2026, the strongest AI systems stopped trusting one model. They send each task to a different one. That choice has a plain lesson for anyone who publishes AI-assisted work.

Read article
AI Verification | | 12 min read

Is AI Deep Research Reliable?

Deep research tools are fast, fluent, and wrong about their sources more often than not. Here is what the audits found, and the two-part fix.

Read article
AI Reliability | | 11 min read

Why AI Can't Check Its Own Work

A model carries the same blind spots into review that it had while writing. Dressing it up as a critic is a costume, not a second mind.

Read article
Comparisons | | 9 min read

TrueStandard vs Omniscient AI

Both run several models, and both show where those models agree. The difference is the moment each one is built for. Omniscient checks what you are reading. TrueStandard checks what you are about to publish.

Read article
AI Reliability | | 12 min read

AI Fake-Citation Disasters: A 2026 Reference

A running list of documented cases where AI made up citations in law, academia, and media. Every entry is sourced, and it is here to be linked, cited, and updated as new cases surface.

Read article
AI Verification | | 10 min read

AI Detectors Ask the Wrong Question

They flag honest writers. They mark famous human documents as AI, and OpenAI quietly killed its own detector. But the deeper problem is not that detection is unreliable. It is that 'was this written by AI?' was never the question that protects you.

Read article
Comparisons | | 9 min read

GPTZero's Hallucination Detector, Explained

It catches citations that don't exist. By GPTZero's own admission, it does not check whether what you wrote is true. Here is exactly what its hallucination and source tools do. Here is where the gap is, and what closes it.

Read article
AI Reliability | | 15 min read

When AI Cites Studies That Don't Exist

AI does not just get facts wrong. It invents whole sources: cases, studies, DOIs. Then it cites them with the same confidence it uses for real ones. Here is why it happens, and the disasters it has already caused. And here is how to catch a fake citation before your name is on it.

Read article
AI Reliability | | 12 min read

How Accurate Is ChatGPT?

Accurate enough to trust for everyday questions. Wrong often enough to get you sued if you publish it unchecked. Here is what the measurements actually say, and what to do about it.

Read article
AI Architecture | | 12 min read

Why AI Hallucinations Are Structural

DELEGATE 52, GPT-5.5, and a Purdue impossibility proof. Three April 2026 results that move 'hallucinations are structural' from take to documented fact.

Read article
AI Reliability | | 12 min read

How to Check If AI Citations Are Fake

Four checks catch a fabricated reference before your readers do, and one of them is new. Here is the change: in 2026, a DOI that resolves no longer means the citation is real.

Read article
AI Reliability | | 11 min read

Should You Stop Using ChatGPT?

Researchers found AI made experts measurably worse on hard tasks. Here is when to trust ChatGPT, and when it is just telling you what you want to hear.

Read article
AI Verification | | 9 min read

AI Detector vs Fact Checker

One asks who wrote this. The other asks is this true. Before you publish, only one of those questions protects your name, and most teams are watching the wrong one.

Read article
Creator Economy | | 12 min read

AI Cloned Your Podcast. Now What?

Three different problems hurt real creators when AI is involved. Identity attestation, AI detection, and claim verification each need a different tool.

Read article
Creator Economy | | 11 min read

What Is AI Slop, and How to Avoid It

Slop and AI-assisted work can look identical on the page. The line between them is whether you verified the output, and whether you can prove it.

Read article
AI Reliability | | 11 min read

Why AI Is Confidently Wrong

Models sound certain every time, even when wrong, and the confident tone you trust in people is worthless here. Here is the fix.

Read article
Comparisons | | 10 min read

An Originality.AI Alternative

Originality.AI is a strong AI-detection suite. But say the job you care about is verifying claims before you publish. Fact-checking is only one of its five bundled checks, and it runs on a single model.

Read article
Comparisons | | 9 min read

TrueStandard vs Parafact

Both verify claims before you publish. The real difference is what one model can miss, and whether your long-form draft fits inside the check at all.

Read article
Comparisons | | 9 min read

TrueStandard vs FactCheckTool

These two tools look alike, but they solve opposite problems. One tells you if the media you read is fake. The other tells you if the draft you are about to publish is true.

Read article
AI Architecture | | 11 min read

Multi-Agent vs Multi-Model AI in 2026

AI builders use both terms as if they meant the same thing. They are different architectures with different strengths. The difference matters most for the one job neither term sells: catching AI errors before you publish.

Read article
AI Reliability | | 12 min read

What Are AI Hallucinations?

Your AI sometimes makes things up and sounds completely confident doing it. Anthropic explains why hallucinations happen and what you can do about them.

Read article
Case Studies | | 11 min read

3 AI Stress Tests from Q2 2026

In April 2026, top AI builders ran real experiments instead of demos. The results were more interesting than the demos. Here is what each test reveals, and why none of them fully answers the question writers care about.

Read article
AI Architecture | | 13 min read

Long Context vs RAG in 2026

Three things just changed about how AI handles your documents. Here is what works for content teams, and why better retrieval still does not mean better truth.

Read article
AI Reliability | | 10 min read

What Is AI Sycophancy?

Your AI agrees with you too much. Anthropic's safeguards team explains why models tell you what you want to hear, and what you can do about it.

Read article
AI Reliability | | 12 min read

What Karpathy's AI Methods Don't Fix

In six weeks, Andrej Karpathy and AI builders shipped three viral reliability methods. Each is real and useful. None of them solves the checking problem for writers.

Read article
AI Fundamentals | | 12 min read

Every Type of AI, Explained

From large language models to coding agents: what each type of AI does, which tools lead each group, and how to pick the right one for your work.

Read article