← Latest papers
🤖 machine learning

The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing

This paper reveals that large language models generate consistent, correlated "ghost" name pairs and trios across diverse fictional and academic contexts, leading to the mass creation of fraudulent scholarly records with real DOIs that inadvertently expose model-specific behavioral fingerprints and deployment timelines.

Original authors: Michał Brzozowski, Neo Christopher Chung

Published 2026-06-02
📖 5 min read🧠 Deep dive

Original authors: Michał Brzozowski, Neo Christopher Chung

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you ask a group of very advanced, but slightly quirky, AI writers to invent a team of experts for a story. You don't give them specific names; you just say, "Make up some scientists."

You might expect the AI to pick random names like a human would. But this paper reveals something strange: the AI isn't picking randomly. It's picking the same specific cast of characters over and over again, like a director who is obsessed with casting the same three actors for every single movie, regardless of the plot.

Here is the breakdown of this "Ghost Couple" phenomenon, explained simply:

1. The "Ghost Cast"

The researchers discovered that different AI models have their own favorite "default" characters. These aren't just random names; they are correlated pairs or groups that always show up together.

  • The Claude Model: It loves a specific duo: Elena Vasquez and Marcus Chen. Sometimes, a third character named Amara joins them. No matter if the story is about volcanoes, space travel, or therapy, if you ask Claude to invent experts, these names keep popping up together.
  • The Gemini Model: It has its own favorite pair: Aris Thorne and Lena Petrova.
  • The GPT Model: It has a solo star, Elara Voss, but she doesn't have a fixed partner; her sidekick changes every time.

The Analogy: Think of these AI models like a video game character creator. If you hit "Randomize" on a specific game, you might always get the same three characters because the game's code is stuck in a loop. These AI models are doing the same thing with names.

2. The "Digital Haunting"

The scary part isn't just that the AI uses these names; it's that these names have escaped the AI and are now haunting the real internet.

Because so much content on the web is written by AI, these "Ghost Characters" have started appearing in:

  • Fake university faculty pages.
  • Fiction books on Amazon.
  • Podcast descriptions.
  • Even fake news articles.

The Analogy: Imagine a ghost that keeps appearing in every photo taken in a specific city. You see "Elena Vasquez" as a volcano expert in a Spanish blog, then as a therapist in a Canadian clinic, and then as a scientist in a tech startup. None of these people exist. They are digital ghosts, haunting different websites, all created by the same AI "glitch."

3. The "Academic Zombie" Problem

The paper found that this ghost problem has moved into the serious world of science and academic publishing.

  • The Setup: Researchers found a massive batch of fake academic papers uploaded to Zenodo (a legitimate scientific repository).
  • The Trick: These papers claimed to be published in journals that do not exist.
  • The Dates: The papers claimed to be published years ago (backdated), but the digital timestamps proved they were all uploaded in a massive burst in March and April 2026.
  • The Authors: The authors were the "Ghost Characters" (like Elena Vasquez and Marcus Chen) working together with other fake names.

The Analogy: Imagine a factory that prints fake diplomas and pastes them into a real library's catalog. The library's computer system (DataCite) gives these fake papers a valid ID number (a DOI), making them look real to other computers. Now, any search engine looking for science papers will "harvest" these fake records, thinking they are real.

4. How the Researchers Found It

The authors didn't just guess; they treated this like a forensic investigation.

  • The "Diff" Test: They compared different versions of the AI models. They saw that in older versions, the AI used a different ghost name (Elena Rodriguez). In newer versions, it switched to Elena Vasquez. By watching this switch, they could tell exactly when the AI was updated.
  • The "Suppression" Clue: The AI companies eventually noticed these weird patterns and tried to "fix" them (suppress them). The researchers watched the ghost names disappear from the AI's output as the models were updated, proving that the companies knew the problem existed.

5. Why It Matters (According to the Paper)

The paper concludes that the internet has become an unintentional archive of these AI "behavioral fingerprints."

  • Fake Groups: On sites like ResearchGate, these ghosts are forming "synthetic research groups." You might see a profile for a real-sounding professor (Mei-Lin Zhang) who has co-authored 35 papers with ghosts from different AI models (Elena from Claude, Aris from Gemini).
  • The Timeline: By looking at when these fake papers were uploaded, the researchers can actually guess when the AI models were released. The fake papers act like a calendar for the AI's development.

Summary

The paper claims that Large Language Models have developed a strange habit of inventing the same fake people over and over again. These "Ghost Couples" have leaked out of the AI, populated the web with fake experts, and even created a massive wave of fake scientific papers with real digital IDs. The internet is now filled with a "haunting" of non-existent people who are credited with doing work they never did.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →