Blog

Guides on AI verification, model selection, and working with artificial intelligence.

Model Selection | | 11 min read

We Benchmarked the Models Behind Our Own Product

Six jobs in our product each pick an AI model. Exactly one of those choices had ever been tested. We built a deterministic benchmark, ran 227 graded calls against a slate of four to five candidates, and changed five of the six. Two of the tables decided nothing, one of our own test labels turned out to be wrong, and the exercise found three bugs it was not looking for.

Read article
AI Reliability | | 10 min read

Does GPT-5.6 Fabricate Fewer Citations Than GPT-5.5?

We asked both generations for peer-reviewed sources on the same 30 claims, on the same day, and resolved every DOI they produced against a real registry. GPT-5.6 came out lower. Then we noticed something in the data that matters more than which model won.

Read article
AI Reliability | | 9 min read

Does Grok 4.5 Fabricate Fewer Citations Than 4.3?

Grok 4.5 shipped on 8 July. Our benchmark, run three weeks later, tested Grok 4.3, so the number we published was already a generation stale. We fixed that by running both models through the same 30 claims on the same day. The rate dropped from 11.1% to 4.4%. Here is why we are not calling that a win.

Read article
AI Reliability | | 12 min read

Which AI Fabricates Citations? We Tested Four Models

We asked four frontier models for peer-reviewed sources on 30 claims, then resolved every DOI they gave us against Crossref. 24 of 146 did not exist. The fabrication rate varies four-fold between vendors, and every single fake landed on a claim that has real literature behind it.

Read article
AI Architecture | | 12 min read

The Best AI Uses Many Models

In 2026, the strongest AI systems stopped trusting any single model. They route each task to a different one. That design choice has a direct lesson for anyone who publishes AI-assisted work.

Read article
AI Verification | | 12 min read

Is AI Deep Research Reliable?

Deep research tools are fast, fluent, and wrong about their sources more often than not. Here is what the studies show, and the two-part fix.

Read article
AI Reliability | | 11 min read

Why AI Can't Check Its Own Work

A model carries the same blind spots into review that it had while writing. Dressing it up as a critic is a costume, not a second mind.

Read article
Comparisons | | 9 min read

TrueStandard vs Omniscient AI

Both run several models and show where they agree. The difference is the moment they are built for: Omniscient checks what you are reading; TrueStandard checks what you are about to publish.

Read article
AI Reliability | | 12 min read

AI Fake-Citation Disasters: A 2026 Reference

A running catalogue of documented cases where AI fabricated citations, across law, academia, and media. Every entry is sourced. It is here to be linked, cited, and updated as new cases surface.

Read article
AI Verification | | 10 min read

AI Detectors Ask the Wrong Question

They flag honest writers, clear famous human documents as AI, and OpenAI quietly killed its own detector. But the deeper problem is not that detection is unreliable. It is that 'was this written by AI?' was never the question that protects you.

Read article
Comparisons | | 9 min read

GPTZero's Hallucination Detector, Explained

It catches citations that don't exist. It does not, by GPTZero's own admission, check whether what you wrote is true. Here is exactly what its hallucination and source tools do, where the gap is, and what closes it.

Read article
AI Reliability | | 13 min read

When AI Cites Studies That Don't Exist

AI does not just get facts wrong. It invents whole sources, cases, studies, DOIs, and cites them with the same confidence it uses for real ones. Here is why it happens, the disasters it has already caused, and how to catch a fabricated citation before your name is on it.

Read article
AI Reliability | | 12 min read

How Accurate Is ChatGPT?

Accurate enough to trust for everyday questions, and wrong often enough to get you sued if you publish it unchecked. Here is what the measurements actually say, and what to do about it.

Read article
AI Architecture | | 12 min read

Why AI Hallucinations Are Structural

DELEGATE 52, GPT-5.5, and a Purdue impossibility proof. Three April 2026 results that move 'hallucinations are structural' from take to documented fact.

Read article
AI Reliability | | 12 min read

How to Check If AI Citations Are Fake

Four checks catch a fabricated reference before your readers do. One of them is new: in 2026, a DOI that resolves no longer means the citation is real.

Read article
AI Reliability | | 11 min read

Should You Stop Using ChatGPT?

Researchers found AI made experts measurably worse on hard tasks. Here is when to trust ChatGPT, and when it is just telling you what you want to hear.

Read article
AI Verification | | 9 min read

AI Detector vs Fact Checker

One asks who wrote this. The other asks is this true. Before you publish, only one of those questions protects your reputation — and most teams are watching the wrong one.

Read article
Creator Economy | | 12 min read

AI Cloned Your Podcast. Now What?

Three different problems hurt real creators when AI is involved. Identity attestation, AI detection, and claim verification each need a different tool.

Read article
AI Verification | | 10 min read

Can One AI Reliably Fact-Check Another AI?

If ChatGPT wrote the draft, can Claude safely verify it? Sometimes helpful, not sufficient by default — and the reason is what these models share, not what they don't.

Read article
AI Reliability | | 11 min read

Why AI Is Confidently Wrong

Models sound certain every time, even when wrong. The confident tone you trust in people is worthless here. Here is the fix.

Read article
Comparisons | | 10 min read

An Originality.AI Alternative

Originality.AI is a strong AI-detection suite. But if the job you care about is verifying claims before you publish, fact-checking is only one of its five bundled checks — and it runs on a single model.

Read article
Comparisons | | 9 min read

TrueStandard vs Parafact

Both verify claims before you publish. The real difference is what one model can miss — and whether your long-form draft fits inside the check at all.

Read article
Comparisons | | 9 min read

TrueStandard vs FactCheckTool

These tools look similar and solve opposite problems. One tells you if the media you're consuming is fake. The other tells you if the draft you're about to publish is true.

Read article
AI Reliability | | 11 min read

Why AI Citations Keep Showing Up Wrong

A 12-fold rise in fake biomedical references, four legal sanctions in 30 days, public defenders flooded with ChatGPT case theories. The same failure shape, across professions.

Read article
AI Architecture | | 11 min read

Multi-Agent vs Multi-Model AI in 2026

AI builders use both terms interchangeably. They are different architectures with different strengths, and the difference matters most for the one job neither term usually advertises: catching AI errors before you publish.

Read article
AI Reliability | | 12 min read

What Are AI Hallucinations?

Your AI sometimes makes things up and sounds completely confident doing it. Anthropic explains why hallucinations happen and what you can do about them.

Read article
Case Studies | | 11 min read

3 AI Stress Tests from Q2 2026

When top AI builders ran real experiments instead of demos in April 2026, the results were more interesting than the demos. Here is what each test reveals, and why none of them fully answers the question writers care about.

Read article
AI Architecture | | 13 min read

Long Context vs RAG in 2026

Three things just changed about how AI handles your documents. Here is what actually works for content teams, and why better retrieval still does not mean better truth.

Read article
AI Reliability | | 10 min read

What Is AI Sycophancy?

Your AI agrees with you too much. Anthropic's safeguards team explains why models tell you what you want to hear, and what you can do about it.

Read article
AI Reliability | | 12 min read

What Karpathy's AI Methods Don't Fix

In six weeks, Andrej Karpathy and the AI builder community shipped three viral reliability methods. Each is real and useful. None of them solves the verification problem for writers.

Read article
AI Fundamentals | | 12 min read

Every Type of AI, Explained

From large language models to coding agents: what each type of AI does, which tools lead each category, and how to choose the right one for your work.

Read article