AI Detector
Paste text. See whether it reads as human or AI, passage by passage.
The number to ask any AI detector for is how often it flags a person as a machine. Ours is on this page. In a test on 1,081 human texts the detector had never seen, it flagged none, including 91 TOEFL essays by non-native English writers. It also says where it is weak, and when it is unsure it says unclear.
About two seconds.
AI Detector runs the moment you sign in, on the text you just pasted. Nothing to enter twice.
A saved check is kept for 24 hours.
You've used up today's free runs on this check.
Come back tomorrow for another free run.
How the AI detector works
It is a trained classifier, which is what the commercial detectors use. Ours was built on a smaller budget and tested in the open.
Your text is cut into 300-word passages
Each passage is scored alone. A human essay with two pasted AI paragraphs shows up as mixed, and the average does not hide them.
Two signals are read from every passage
An embedding captures how the prose is built. A fast classifier adds three judgments: how likely a model wrote it, whether it looks edited by one, and what genre it is.
A model trained on paired examples combines them
We took 2,300 human texts written before 2022, before ChatGPT existed, and had current GPT, Claude, Gemini, Grok and GLM models write their own version of each. The classifier learned the difference from those pairs.
The threshold is set to protect the writer
A passage is called AI only above a score that flagged 1% of human texts during training. Between that line and even odds the result says unclear.
The method follows the one Pangram Labs published for its detector. It needs human text that predates the models, an AI rewrite of every document, and a threshold chosen on human writing held back from training. Ours is the small version of it.
The record, measured on texts the detector never saw
- Human texts wrongly flagged Reddit posts, news, reviews, abstracts, Wikipedia, Paul Graham essays, two blogs, eight classic books and 324 student essays.
- 0 of 1,081
- Non-native English writers wrongly flagged TOEFL essays from the Stanford study in which seven detectors flagged more than half of them.
- 0 of 91
- AI essays and blog posts missed Including 130 written in the voice of an author the detector was never trained on.
- 3 of 216
- Short-form AI text missed Reddit posts, reviews, news items and abstracts under 300 words. This is the weak spot.
- 103 of 386
Where it fails
- Polished first-person narrative, the kind written for college applications. It wrongly flagged 21 of 70 real admission essays. We are retraining on that genre and will update this number.
- Human writing that a model has edited. It missed 84 of 91 essays polished by GPT-4, because most of the writing is still the person's.
- Short text. Under 150 words there is not much to measure.
- Text rewritten to evade detectors by tools we have not tested. We tested prompts that ask the model to sound human, not commercial humanizers.
Knowing who wrote it does not tell you whether it is true. Why detection and verification are different questions
If the text makes claims, check them: run the AI fact checker
Example: Thoreau, and a model writing as Thoreau
Two real runs from 19 September 2026. The first is a 300-word passage from Walden, a book the detector was never trained on. The second is Gemini 3.7 Flash, a model it was never trained on, asked to write the same passage. Each took about 1.4 seconds.
Each stick was carefully mortised or tenoned by its stump, for I had borrowed other tools by this time. My days in the woods were not very long...
Score −9.24. The flag line is 3.10.
A man never knows the true virtue of his hands until he undertakes to fashion his own roof from the standing timber of the wilderness...
Score 4.59, above the flag line of 3.10.
Frequently Asked Questions
On 1,081 human texts it had never seen, it flagged none. On essays and blog posts written by AI it missed 3 of 216. On short AI text such as Reddit posts and reviews it missed 103 of 386, about one in four. Those are our own measurements on a test set we built, and the full breakdown is on this page.
It flagged none of 91 TOEFL essays written by non-native speakers. Those essays come from a Stanford study in which seven detectors flagged more than half of them. One test set is not a guarantee, so if it ever flags your own writing, treat that as our error.
No. A score is evidence about style and it can be wrong. It wrongly flags some polished personal essays, and it misses most human writing that a model has edited. Use it to decide what to look at more closely.
It was trained on text from GPT, Claude, Gemini, Grok and GLM models current in September 2026, and tested on Gemini 3.7 Flash, which it had never seen. New models shift the signal, so we retrain when one ships.
Eighty words is the shortest passage it was trained on. Below that there is too little to measure. Under 150 words the result carries a caution, because that is where it misses the most.
It shows its false-positive count and its weak spots next to the tool, and it has a third answer, unclear, for text it cannot call either way.
One text a day without an account, up to 12,000 characters. A free account gets three a day.
Powered by TrueStandard
Four independent models check every claim in your draft.