A Tool Just Flagged 260,000 Cancer Papers for Paper-Mill Fraud. It Has Nothing to Do With AI Text Detection.
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI
Quick answer
A BERT classifier trained on retracted papers just screened 2.6 million cancer-research titles and abstracts and flagged 261,245 of them, 9.87%, as resembling known paper-mill output. That number is getting reported alongside AI-writing statistics, but it measures something completely different from an AI-text detector. One looks for phrasing that matches known fraud templates. The other looks for the statistical fingerprints of machine-generated prose. Confusing the two leads editors to draw the wrong conclusion from either tool.
What the study actually did
Researchers Baptiste Scancar, Jennifer A. Byrne, David Causeur, and Adrian G. Barnett published "Machine learning based screening of potential paper mill publications in cancer research: methodological and cross sectional study" in The BMJ (2026;392:e087581). They trained a BERT-based text classifier on 2,202 retracted cancer papers tagged "paper mill" in the Retraction Watch database, plus 3,094 more papers flagged by independent research-integrity sleuths. On internal validation the model hit 0.91 accuracy, 0.87 sensitivity, and 0.96 specificity. Tested against the separate integrity-experts dataset it never trained on, it held up at 0.93 accuracy.
They then ran the trained model across the entire indexed cancer-research literature from 1999 to 2024, roughly 2.6 million papers. It flagged 261,245 (9.87%, 95% CI 9.83 to 9.90) as resembling the retracted paper-mill set. Of those flagged papers, 177,907 carried a Chinese institutional affiliation, about 35% of all flagged output. The flagged share rose almost exponentially from under 1% of annual output in the early 2000s to over 15% by the early 2020s, before flattening after 2022. Flagging wasn't confined to obscure journals either: the percentage of flagged papers in the top 10% of journals by impact factor also climbed steadily over the same period.
Why this isn't an AI-text detector, even though the numbers get lumped together
Here's the mechanical difference that matters. The BMJ model learned to recognize phrasing patterns statistically associated with a training set of already-confirmed paper-mill papers. It works purely on the text of titles and abstracts, and it makes no claim about whether any specific fact, number, or method in a paper is genuine. It's pattern-matching against known fraud templates, the same family of technique publishers already use to search for "tortured phrases," like "cruel temperature" standing in for "mean temperature," a telltale sign of paraphrasing software laundering plagiarized text.
An AI-text detector like TheChecker.AI works on an entirely different signal. It measures statistical properties of the writing itself, perplexity (how predictable each next word is) and burstiness (how much sentence length and structure vary), to estimate whether a large language model most likely generated the prose, independent of whether the underlying claims are true or false. A paper mill's fabricated data could, in principle, be written by a human ghostwriter using zero AI assistance, and it would still get flagged by the BMJ classifier for matching known fraud templates. Conversely, a legitimate researcher who used ChatGPT to help draft a section of an honest, factually sound paper would show up on an AI-text detector but never trip a paper-mill classifier, because nothing about the claims resembles fraud.
The authors of the BMJ study are explicit about this boundary too. They describe their tool as "a probabilistic model, not a definitive arbiter of misconduct," intended to flag papers for human follow-up rather than confirm fraud. A flag from their model means "this abstract's phrasing statistically resembles phrasing in papers we already know were paper-mill products." It is not a claim that any specific sentence was written by AI.
Where the two problems are starting to blur
The overlap is real, just not where the headline number suggests. As generative AI tools get better at producing fluent, "academic-sounding" text on demand, paper mills have an obvious incentive to adopt them, both to write fabricated manuscripts faster and to evade the very phrasing-pattern detectors publishers built against older, template-based fraud. The BMJ authors flag this directly in their discussion: the flattening of flagged papers after 2022 could reflect publishers fighting back, or it could reflect paper mills shifting to new templates "following the rise of AI," their words, not ours. Either way, the two detection problems are converging even though they started from opposite directions: one asks "does this resemble known fraud," the other asks "does this resemble machine-generated prose."
That convergence is also why journals are increasingly running both kinds of checks side by side rather than picking one. A cancer-research publisher's own review-report screening already found reviewers were quietly using AI to write their review comments far more often than they disclosed. Separately, a corrected detection method found signs of AI writing in 89% of biomedical manuscripts submitted in one recent month, a number that measures writing assistance, not fraud. None of that tells you whether a paper's underlying data is real. Fraud detection and AI-writing detection answer two different questions, and an editor who only runs one check is leaving the other question completely unanswered.
What the model gets wrong, in the authors' own words
The BMJ paper is unusually candid about its own limitations, and they're worth naming because they cut against treating "9.87% flagged" as a clean fraud rate. Training data skewed toward retractions from 2013 to 2023, and within that toward Chinese-affiliated papers, creates a residual risk the model partly learned to recognize country-of-origin writing patterns rather than fraud signals specifically. The authors checked for this directly: false positives were rare (39 out of 3,375 control papers), and the overrepresentation of Chinese papers showed up among false negatives, papers the model missed, not papers it wrongly accused. That's evidence against systematic over-flagging, but it doesn't erase the underlying bias risk in a model trained mostly on one class of publicly-tagged fraud.
The model is also not explainable. As a deep-learning classifier, it can't point to which specific words or phrases drove any individual flag, which means a "flagged" label carries no built-in justification a human reviewer could check against. And the ground truth itself is fuzzy: the "paper mill" tag in Retraction Watch reflects each publisher's own retraction notice, with no uniform standard across publishers for how much evidence justified the label in the first place.
The takeaway for editors and researchers
If you're a journal editor evaluating a submission, a fraud-pattern classifier and an AI-writing detector answer different questions, and you need both to get a complete picture. A paper can pass a fraud screen cleanly (its phrasing doesn't resemble known paper-mill templates) while still containing large sections of undisclosed AI-generated prose, which most journals now require authors to declare. Running a manuscript through TheChecker.AI's free AI detector tells you specifically whether the writing itself carries the statistical signature of an LLM, a separate and complementary check to any fraud-pattern screening a publisher's integrity team already runs.
FAQ
Does a paper-mill classifier flag the same papers an AI-text detector would flag? Not necessarily. The two tools measure different things: template-matching against known fraud versus statistical prose patterns associated with LLM generation. A paper can trip one and not the other.
Is a 9.87% flag rate the same as a 9.87% fraud rate? No. The authors describe flags as a probabilistic screening signal requiring human follow-up, not a confirmed fraud rate. False positives in their validation sets were rare, but the flagged share includes papers that plausibly resemble fraud templates for reasons other than actual fraud.
Why does this matter for AI-detection specifically? Because as paper mills adopt generative AI to write fabricated manuscripts, and as legitimate researchers increasingly use AI to help draft honest papers, publishers need to run fraud detection and AI-writing detection as separate, complementary checks rather than assuming one covers the other.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.
Related posts
AI Chatbots Keep Inventing the Same Fake Scientists, and Some Now Have Real DOIs
An AI preprint found Claude, Gemini, and GPT invent the same fake experts, some now with 1,655 real, citable DOIs.
Read more
Why AI-Generated Wikipedia Edits Get Caught by Verification, Not Detection
Wiki Education flagged 178 AI-suspect Wikipedia articles. Most failed citation verification, not fake sources. Detection flags where to look.
Read more
Harvard's "Two-Way Suspicion": What AI Detection Is Doing to Classroom Trust
A new Harvard report names "two-way suspicion": teachers' AI-dar hunches and student fear, eroding trust without proof.
Read more