Why "Polished" AI Writing Gets Flagged More Than Raw AI Text
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI
Quick answer
A new Carnegie Mellon study found something that sounds backwards: when an AI model finishes a sentence without any chat-style polish, GPTZero and Pangram call it human almost every time. Run the identical model through instruction-tuning first, the version behind every consumer chatbot, and the same detectors flag it as AI most of the time. The detectors aren't reading intent or authorship. They're reading a statistical fingerprint that instruction-tuning and RLHF leave behind, and that fingerprint, not "AI-ness" itself, is mostly what a commercial detector actually measures.
The study, in plain terms
Researchers at Carnegie Mellon (Yixuan Even Xu, Ziqian Zhong, Aditi Raghunathan, Fei Fang, and J. Zico Kolter) tested Llama-3-8B and Qwen3-8B in two forms each: the raw "base" model, which just continues text with no chat training, and the standard instruction-tuned version that powers assistant-style products. They fed both versions the same human-written and AI-written prefixes and asked GPTZero and Pangram to score the continuations.
The gap was large. On Llama-3-8B given a human-written prefix, GPTZero rated the base model's continuation 96.7% human and the instruction-tuned model's continuation only 30.3% human. Pangram's numbers moved the same direction: 98.8% human for the base model, 17.1% for the instruction-tuned model. Qwen3-8B showed a similar split. Same underlying model family, same prefix, same task. The only thing that changed was whether the model had gone through instruction-tuning, and that alone swung the detector's verdict by 60 to 80 percentage points.
The paper (posted May 2026, cs.CL, arXiv:2605.19516) explains why this happens. Instruction-tuning and reinforcement learning from human feedback reshape a model's outputs in specific, measurable ways: more predictable sentence rhythm, characteristic phrasing patterns, a tendency toward the "helpful assistant" register. Prior research already knew RLHF changes response length and introduces sycophancy. This study shows the same post-training reshaping is a big part of what current commercial detectors are actually keying on, more than any stable signature of "text a machine produced" as opposed to "text a person wrote."
Why this matters more than it sounds
This is not a story about detectors being broken. GPTZero and Pangram both report vanishingly low false-positive rates on ordinary human writing, and independent testing backs that up: real reporting from Nature on the state of AI detection found Pangram scored zero false positives across hundreds of human texts in an Epoch AI benchmark, with GPTZero close behind. Nobody is disputing that these tools catch chatbot output when a student pastes a ChatGPT answer straight into an essay.
What the CMU paper shows is narrower and, for anyone using detection results to make a decision, more important: a detector's score is a measurement of how closely text matches the statistical shape of instruction-tuned assistant output. That's usually a good proxy for "did a chatbot write this," because almost every AI text a person encounters in the wild, from a ChatGPT reply to a Claude draft, came out of an instruction-tuned, RLHF-shaped model. But it means the detector is measuring a specific kind of machine writing, not a universal, unfakeable signature of non-human origin. Text that avoids that particular statistical shape, whether because it came from an unusual model, went through several rounds of editing, or happens to share surface features with base-model continuation text, sits in a genuine blind spot. The researchers used that blind spot to build and test a rewriting method that pushed detector scores down while trying to preserve meaning. We're not walking through how it works or how to reproduce it here. The point that matters for our readers isn't the exploit. It's what the existence of a working exploit proves: detection is reading a fingerprint left by how text was produced, and any fingerprint that can be identified this precisely can, at least at the margins, be altered.
That's exactly why we've said all along that a detection score is a data point, not a verdict. A number close to 100% AI on a document that came from a chatbot is strong evidence. A number close to 100% human on a document engineered specifically to avoid the instruction-tuning fingerprint is weaker evidence than it looks, because it's answering "does this match the shape we usually catch," not "did a human definitely write every word." Neither failure mode means detection is worthless. Both mean a single score, read in isolation, is answering a narrower question than most people assume when they see it.
How this connects to what detectors actually look at
We've written before about the base mechanics of how AI text detection works: perplexity, burstiness, and increasingly, learned statistical patterns from training on paired human and machine text, the same supervised-classifier approach Pangram and GPTZero both use. The CMU findings sharpen that picture. A learned classifier doesn't need to understand meaning or authorship to work well most of the time; it needs the training data to represent the kinds of AI text it will actually see. Since nearly all AI text people encounter comes from instruction-tuned, RLHF-shaped assistants, training on that distribution produces a detector that's extremely good at catching exactly that kind of writing, and comparatively weaker on writing that falls outside it. This is a variant of the same theme in why detectors still miss AI text even when nobody is trying to fool them: the gap isn't always adversarial. Sometimes it's just a mismatch between what a detector was trained to recognize and what's actually in front of it.
It's also a useful companion to the honesty we've already built into our coverage of false positives and what they do and don't prove. A false positive on human writing and a false negative on machine writing that dodges the fingerprint are two sides of the same underlying fact: the detector is scoring similarity to a pattern, not certainty about origin.
What this means if you're the one being scored
If a detector flags your writing, or clears someone else's, the CMU paper is a reminder to ask what's actually being measured before treating a single number as the end of the conversation. A high AI-probability score on writing that reads like typical chatbot output is meaningful; that's the exact pattern these tools are built to catch, and they catch it well. A low score doesn't automatically mean a human wrote every sentence, especially as more people run AI text through some form of rewriting before submitting it. Neither situation calls for panic. It calls for treating the score as one input, the same way we've argued in every honest evaluation on this blog: a sentence-level read of where the signal actually concentrates in a document tells you more than a single whole-document percentage, and it's a better starting point for a real conversation than an unexamined number either way. Hard to say how long any specific evasion gap stays open, for what it's worth. The CMU team built and tested a rewriting method that pushed detector scores down in the lab, and neither the paper's write-up nor this post walks through the mechanics. History gives a hint anyway: earlier paraphrase-attack research got folded into vendor retraining within months of publication, so a gap this well-documented probably doesn't survive long.
FAQ
Does this mean AI detectors don't actually work? No. It means they work by recognizing the statistical fingerprint of instruction-tuned, RLHF-shaped writing, which covers the overwhelming majority of AI text people actually encounter. The finding narrows what a score proves, it doesn't erase what detectors catch correctly every day.
See where a score actually comes from, not just what it says
A single percentage can't tell you whether you're looking at typical chatbot output, a base-model edge case, or an ordinary sentence a detector happens to be uncertain about. Try TheChecker.AI's free demo to see a sentence-level read instead of a single number, or check our published accuracy data to see exactly what our detector is built to catch and where honest limits still apply.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.
Related posts
What Actually Decides an AI-Cheating Dispute? Five 2026 Cases Say the Same Thing
Five AI-cheating disputes reached real verdicts in 2026. None hinged on detector accuracy — all came down to whether the appeal process held up.
Read more
Russian State Media Used AI to Write News Scripts. No AI Detector Would Have Caught the Finished Broadcast.
Anthropic's Sept 2026 report shows AI-polished state media content that detectors can't reliably catch after editing.
Read more
Can You Check If a Research Paper Is AI-Written Before You Cite It?
AlphaXiv now flags AI-written sections in arXiv papers. NeurIPS's own test shows window size alone swings a paper's score from 43% to 13%.
Read more