Back to Blog
Detection Guide 8 min read

Can Turnitin Detect ChatGPT? What Turnitin's Own Data Actually Says

By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Ink-wash illustration of a torn exam paper passing through a stamping press, half emerging clean and half emerging smudged, symbolizing inconsistent AI-detection outcomes

Quick answer

With a document-level false-positive rate under 1%, Turnitin's AI-writing detector reliably distinguishes machine-generated work from human writing, and its own Chief Product Officer has said the tool is tuned to catch roughly 85% of AI writing, trading some misses for fewer false accusations (BestColleges, "Testing Turnitin's New AI Detector"). That's a genuinely strong number — but it isn't the same as "Turnitin is always right," and two major universities have published their own reasons for pulling back on it anyway. Here's what Turnitin's own data says, what independent testing found, and what a single score should and shouldn't decide.

What Turnitin actually claims, in Turnitin's own words

Turnitin publishes the numbers directly. The company's stated design goal is a document-level false-positive rate under 1%. Fewer than one in a hundred fully human-written documents should be incorrectly flagged as AI-generated. This requirement applies specifically to documents where 20% or more of the text is AI-written (Turnitin, "Understanding false positives within our AI writing detection capabilities").

That's the headline. The more interesting number sits one level down: Turnitin's own follow-up post states its sentence-level false-positive rate is around 4% — meaning that within a document, individual sentences highlighted as "AI-written" have roughly a 1-in-25 chance of actually being human-written (Turnitin, "Understanding AI writing detection: False positive rates"). Turnitin explains that this happens most often right at the seam where human and AI writing meet in the same document — 54% of these sentence-level false positives sit directly next to genuine AI text, per the company's own analysis.

Annie Chechitelli, Turnitin's Chief Product Officer, put the trade-off in plain terms to BestColleges: "We would rather miss some AI writing than have a higher false positive rate. So we are estimating that we find about 85% of it. We let probably 15% go by in order to reduce our false positives to less than 1 percent." In other words, Turnitin itself confirms the detector is built on purpose to under-flag rather than over-accuse — so "can it detect ChatGPT" gets a real yes, right alongside an equally real "not every time, by design."

Turnitin also keeps updating the underlying model rather than shipping it once. Its public release log shows detection updates in both October 2025 and February 2026 described in Turnitin's own words as improving recall "while maintaining a low false positive rate," plus a May 2026 update specifically for Spanish-language detection (Turnitin, AI writing detection model release notes) — evidence the 85%-catch-rate figure from 2023 isn't a permanently fixed number, in either direction.

What happens when institutions test it themselves

Two research universities ran their own numbers in the past year. Their findings diverged from what Turnitin's marketing copy alone would suggest.

The University of Waterloo discontinued Turnitin's AI-detection feature in September 2025 after its own Instructional Technologies and Media Services group tested it internally. Waterloo's public rationale cites three independent research papers questioning detector reliability generally, and states plainly that "in more than one instance the product flagged human written text as 100% generated by AI" during its own testing (University of Waterloo, "Discontinuing use of AI detection functionality in Turnitin").

Washington State University went further in February 2026, cancelling its Turnitin AI-detection contract outright. The provost's memo does the arithmetic explicitly: in Fall 2024 alone, Turnitin analyzed 148,547 assessments at WSU — and even accepting Turnitin's own claimed 1% false-positive rate at face value, that implies roughly 1,485 human-written assessments were likely flagged as AI-generated in a single semester. The same memo reports that 33% of academic-integrity hearing board cases involving AI allegations between 2023 and 2025 ended in a finding of "not responsible," specifically because AI-detection output had been submitted as the only evidence (WSU Office of the Provost, "Cancellation of Turnitin AI Detection software").

The claim that Turnitin never catches AI writing is absent from either institution's statement — both memos acknowledge the software catches a real share of it. The argument is narrower and more useful: a headline accuracy number reflects performance across a large batch of documents, not certainty about any one specific submission, and that's where both schools drew the line before treating a single score as sufficient evidence for a misconduct finding.

So does it detect ChatGPT specifically?

Independent academic testing backs up that detectors generally can distinguish ChatGPT output from human writing at a meaningfully high rate, though this specific study doesn't test Turnitin itself. A 2025 study published in Acta Neurochirurgica ran 1,000 texts — a mix of pre-ChatGPT human writing and abstracts generated by GPT-3.5, GPT-4, and GPT-4o — through three other AI-output detectors (GPTZero, ZeroGPT, and Corrector App) and found the detectors' scores reliably distinguished AI-generated content from human-written texts, with area-under-curve scores between 0.75 and 1.00 across models — the paper's own scale labels 0.75 "fair" and reserves "excellent" for 0.90 and above, so the fit ranged from solid to very strong depending on the tool and model. The study's actual quoted conclusion is the part that matters for a claim like "detects ChatGPT": "none of the detectors achieved 100% reliability in distinguishing AI-generated content" (Erol et al., "Can we trust academic AI detective? Accuracy and limitations of AI-output detectors," PMC).

Every credible dataset on this question lands in roughly the same shape: strong detection of raw, unedited ChatGPT output, a real but small false-positive rate on genuine human writing, and a far murkier picture once the text gets lightly edited, blended with human writing, or run through a paraphrasing pass. None of that makes Turnitin, or any detector, useless. A line like "Turnitin detected 82% AI" is a fact worth discussing, not a verdict to act on by itself — exactly the position Turnitin's own chief product officer, Waterloo, and WSU each reach on their own.

What to actually do with a score like this

  • Treat a mid-range score as a prompt, not proof. The 20–50% range is exactly where the sources above say confidence is lowest, so it's a band to ask about, not act on.
  • Ask for process evidence before assuming intent. Draft history, version timestamps, and research notes settle unclear cases faster and more fairly than a percentage alone.
  • Don't rely on one tool's read of one submission. Detectors disagree with each other in documented testing, so a second independent read is a real check, not overkill.

That last point is exactly what a free check is for. TheChecker.AI's demo gives you a second, sentence-level read on any text — up to 1,000 characters, no signup needed — so you can see where a detection score is confident versus where it's sitting in that murky middle ground before treating any single number as an answer. If you want the underlying mechanics of how any detector, including ours, generates that score in the first place, how AI detectors actually work breaks down perplexity and burstiness in plain terms. And if you're weighing detection tools generally, our own accuracy page publishes the same kind of transparent, tested figures Turnitin does — including where a single score should stop being treated as a verdict, as we've written about before. If you're specifically comparing TheChecker.AI against Turnitin for classroom or team use, our side-by-side breakdown covers that directly.

FAQ

Is Turnitin's AI detector accurate? Turnitin's own published data claims a document-level false-positive rate under 1% and an estimated 85% detection rate for AI-generated content, by design trading some missed detections for fewer false accusations. Independent academic testing generally supports strong — not perfect — accuracy on unedited AI text.

Why did some universities stop using Turnitin's AI detector? Not because it never works. The University of Waterloo and Washington State University both cite the real, non-zero false-positive rate — and the harm a false accusation causes a student — as reason enough to stop treating a detection score as sufficient evidence on its own, even at a false-positive rate Turnitin itself calls low.

Can Turnitin be fooled by lightly edited or paraphrased AI text? Light editing or paraphrasing can deceive Turnitin, according to its own documentation. The vendor's release notes outline continuous updates specifically aimed at countering "AI bypasser" and paraphrasing tools, which implies the detector's confidence drops on edited text compared with raw AI output — consistent with what independent research finds across detectors generally, not a Turnitin-specific weakness.

Dusan Boljevic

AI/ML Engineer at TheChecker.AI

Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.