How AI Detectors Work: And Why Honest Writing Sometimes Gets Flagged
A plain-English guide to perplexity, burstiness, and the stylometry behind tools like GPTZero: plus how to check your own text.
If a teacher, editor, or recruiter has ever run your writing through an "AI detector," you have probably wondered the same thing everyone does: how does it actually know? The honest answer is that it does not know anything. It estimates. And understanding what it measures is the difference between trusting a score blindly and reading it for what it really is: a probability, not a verdict.
This is a practical, no-hype guide to how AI text detection works, why it sometimes flags genuinely human writing, and how to sanity-check any passage yourself.
Detectors measure predictability, not authorship
Large language models are trained to pick the most likely next word. Do that millions of times and the output becomes statistically smooth: fluent, evenly paced, and a little too tidy. Detectors look for that smoothness. Two measurements do most of the work:
- Perplexity: how surprised a language model is by your word choices. AI text tends to be low-perplexity (predictable). Human writing zig-zags into less expected words.
- Burstiness: how much your sentence length and rhythm vary. Humans write in bursts: a long, winding sentence followed by a short one. AI tends to hold a steady, uniform cadence.
Rule of thumb: low perplexity + low burstiness = "reads like AI." High variation in both = "reads like a person." That is the entire core idea behind GPTZero and most statistical detectors.
What else detectors look at
Beyond those two signals, stronger detectors stack on a handful of stylometric features, the same fingerprints linguists use to identify authorship:
- Vocabulary diversity: how often you reuse the same words.
- Punctuation patterns: AI overuses the em dash and the semicolon; humans are messier.
- Phrasing fingerprints: "delve into," "it is important to note," "in today's fast-paced world," "a testament to," and the tidy three-item list are instruction-tuned tells.
- Repetition: models reuse templated transitions more than people do.
Why honest writing gets false-flagged
This is the part most detectors quietly downplay. Because the signal is "predictability," any writing that happens to be clean and even-paced can trip the alarm, even when no AI was involved.
- Non-native English writers often use simpler vocabulary and more uniform sentence structures. Studies have found detectors flag non-native essays at alarming rates; one widely-cited result hit a 61% false-positive rate on TOEFL essays.
- Concise, formulaic formats: lab reports, legal summaries, technical docs are naturally low-perplexity.
- Heavily edited human writing can be polished into exactly the smoothness detectors punish.
This is why no responsible tool should ever output a flat "this is AI" verdict. A score is evidence, not proof. Never use one to accuse someone.
The cat-and-mouse: humanizers and deep detection
Because statistical detectors only measure surface predictability, "humanizer" tools can defeat them: they raise perplexity and burstiness, swap synonyms, and vary sentence length until the text reads human, even if a model wrote it. Pure statistical detectors largely miss humanized text. That is a known limitation of the whole category.
The next generation of detection adds a model-based layer on top of the statistics: instead of only measuring rhythm, it reads the underlying structure and meaning, which survives surface rewriting. That is how reworded AI still gets caught. Our own Deep analysis works this way, and it is also why we built strong protection for non-native writers directly into the scoring.
How to make AI-assisted writing read as your own
If you use AI to draft and then make the work genuinely yours, the goal is not to "trick" anything; it is to put your voice, judgment, and specifics back in. The same edits that lower an AI score also make writing better:
- 1Vary your sentence length on purpose. Follow a long sentence with a short, blunt one.
- 2Cut the generic connectors such as "moreover," "furthermore" and "in conclusion," and the tidy three-item lists.
- 3Add concrete specifics only you would know: a real number, a name, a date, a small story.
- 4Take a position. AI hedges; people commit. Say what you actually think.
- 5Read it aloud. If it sounds like a brochure, rewrite the parts that do.
Check any text in seconds
Our free AI Detector shows you the same signals described here: a perplexity-and-burstiness breakdown plus sentence-by-sentence highlighting, so you can see why a passage reads the way it does, not just a number. Signed-in users can run Deep analysis, which catches humanized and paraphrased AI that statistical-only tools miss, and which is specifically tuned to protect non-native English writing from false flags.
Whatever tool you use, treat every result as a conversation starter, not a confession. The technology is genuinely useful, and genuinely fallible. Knowing the difference is the whole game.
Related Features
Notes
Auto-structured notes from any audio, video, or document, organised with headings, highlights, and key takeaways.
Learn more about NotesFlashcards
Generate study-ready flashcards from your transcripts and notes using proven spaced-repetition principles.
Learn more about FlashcardsQuiz Mode
Test your understanding with auto-generated quizzes pulled directly from your content.
Learn more about Quiz ModeReading Mode
Transform static PDFs and documents into interactive, searchable study material.
Learn more about Reading Mode