Back to Blog
Research & Academic Integrity 7 min read

AI Chatbots Keep Inventing the Same Fake Scientists, and Some Now Have Real DOIs

By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Torn paper diorama of a diploma dissolving into ghostly silhouettes with a blank official seal, ink-on-paper style

Quick answer

Elena Vasquez has camped on an active volcano, survived 11 eruptions alongside her research partner Marcus Chen, and co-authored papers on Zenodo and ResearchGate. None of it is true. She does not exist. A June 2026 preprint from researchers at the Samsung AI Center Warsaw found that Claude, Gemini, and GPT each default to the same small set of fictional "expert" names when prompted to invent one, and that these invented identities have already picked up 1,655 real, citable DOIs on a CERN-operated academic repository. This is not a text-quality problem an AI detector solves by itself. It is an identity-verification problem, and the same fabricated bios that fool a search engine are themselves AI-generated text worth screening before you trust them.

The names a chatbot reaches for when it invents a person

Researchers Michał Brzozowski and Neo Christopher Chung ran a simple experiment. They prompted nine versions of Claude, ten versions of GPT, and one version of Gemini to generate stories about fictional experts, research partners, and collaborators, then logged which names came up. The results, published as The Ghost Couple, were not random.

Claude models consistently reached for the pairing of Elena Vasquez and Marcus Chen, sometimes joined by a third name, Amara Okafor. Gemini defaulted to Aris Thorne and Lena Petrova together in 37% of prompts. GPT models favored a single recurring name, Elara Voss, without a fixed partner. These are not coincidental overlaps. The paper describes them as "correlated character ensembles" whose co-occurrence rates "far exceed chance" and stay consistent across independent runs of the same model.

The pattern also moves with model versions. An earlier Claude model frequently produced a different fictional name, Elena Rodriguez. Claude Sonnet 4, released in May 2025, shifted almost entirely to the Vasquez-and-Chen pairing, which Nature reported showed up together in about 23% of the researchers' prompts. By Claude Sonnet 4.6, released in 2026, that pairing had largely vanished from sampled outputs. That means the specific fake name a model generates can function as a rough timestamp for which model and version produced a given piece of text, the same way a car's design details can date the decade it was built.

Where this stops being a curiosity

A chatbot inventing a name for a story is harmless on its own. The problem is what happens when that invented name gets attached to something that looks like a real academic record. The researchers searched Zenodo, a CERN-operated repository that issues real DataCite digital object identifiers, and identified 1,655 ghost-authored records under names like Elena Vasquez claiming to belong to journals that do not exist. Server-side DataCite timestamps show the publication dates on many of these records were deliberately backdated. The researchers found 991 of these records were registered in a single month alone.

That matters because a DOI is supposed to be a durable, trustworthy pointer to a real piece of scholarship. Once one is minted, it can be harvested by any citation index, scholarly search engine, or aggregator that ingests DOI metadata without checking whether the underlying work or its authors are real. The same ghost names turned up on ResearchGate, where they formed what the researchers call synthetic research groups, with fabricated collaborators drawn from multiple different AI models appearing together on the same fake profile.

Sidney Wong, a computational linguist at the University of Otago, told Nature the scale of the problem is "scary," even though some signs suggest it has partly improved since the preprint's publication. The concern is not that any single fake paper fools a careful reader. It is the volume: hundreds of fabricated identities accumulating real, trackable, citable metadata inside the infrastructure researchers and journals rely on every day.

This is a different failure than the one text detectors usually catch

Most of what we cover on this blog is about detecting whether a passage of prose was generated or substantially written by AI, which is a text-pattern problem. Ghost co-authors are a related but separate failure mode: the text of a fake bio might read perfectly plausibly, because an AI model wrote a plausible-sounding paragraph about a plausible-sounding volcanologist. The actual lie is in the identity, not necessarily in any single sentence's phrasing.

That distinction matters for anyone deciding whether to trust an unfamiliar name on a paper, a preprint, or a conference program. The practical defense looks less like fact-checking one claim and more like the kind of verification habit we described in how to check whether a research paper is AI-written before you cite it: look for an ORCID with a real publication history, check whether the claimed institution lists the person, and treat a bio you cannot independently confirm as a reason to look further, not a reason to assume good faith.

There is also a direct overlap with the AI-text-detection side of this problem. The bios, abstracts, and author descriptions attached to these ghost identities were themselves generated by the same models that invented the names. A suspicious author bio, an unusually polished "about the researcher" paragraph, or a submitted short biography that reads like promotional copy for a person nobody can verify is exactly the kind of text an AI detector is built to flag. We've written before about why paper mills already use detection tools at scale to screen submissions for fabricated or AI-written content, and the same logic applies one level up: screening the author's own words, not just the paper's.

What journals and editors can actually do with this

None of the standard fixes here are exotic. Editors and conference organizers already have tools for verifying identity, they just are not always applied consistently to every new or unfamiliar name on a submission:

  • Check whether a claimed ORCID iD resolves to a real, populated profile with a prior publication history, not a blank or recently created one.
  • Search the claimed institution's own faculty or staff directory for the name, rather than trusting a bio paragraph alone.
  • Treat a DOI as a starting point for verification, not proof of legitimacy. A real DOI only confirms a record was registered, not that the person or journal behind it exists.
  • Run an unfamiliar author's submitted bio or cover letter through TheChecker.AI's free detector as one more signal alongside identity checks, the same way we've argued a single AI-detector score should never be the only thing deciding a case involving a student or a writer.

FAQ

Is Elena Vasquez a real person who had her identity misused? No. Elena Vasquez does not correspond to any real individual. She is a name that large language models, particularly Claude, generate by default when prompted to invent a fictional expert or research partner, and the name has since been attached to fabricated content across the web and academic repositories by people using these models.

Can an AI detector catch a fake co-author? Not directly. An AI detector analyzes whether a block of text shows patterns consistent with AI generation. It cannot confirm whether a named person exists. What it can do is flag whether a submitted bio, abstract, or cover letter was itself AI-written, which is a useful supporting signal alongside identity checks like ORCID verification and institutional lookups.

Why does the fake name change between model versions? The researchers found that specific fictional name pairings are tied to how individual model versions were trained. The Vasquez-and-Chen pairing that showed up in about 23% of prompts under Claude Sonnet 4 had largely vanished from sampled outputs by Claude Sonnet 4.6. That makes the presence of a particular ghost name a rough, informal indicator of which model and version likely produced a piece of content.

Does this affect real, legitimate preprint servers like arXiv? The researchers' findings focused on Zenodo and ResearchGate, platforms that allow broader self-publishing with less editorial gatekeeping than arXiv's moderated submission process. That does not mean no risk exists elsewhere, but the documented 1,655 ghost-authored records were concentrated on platforms with lighter verification, which is a useful distinction when deciding how much scrutiny a given source deserves.

Dusan Boljevic

AI/ML Engineer at TheChecker.AI

Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.

Interested in using TheChecker.AI?

Try it free