The Digital Ghosts Haunting Our Academic Libraries
If you search for a software developer named Marcus Chen or a medical researcher named Elena Vasquez, you might stumble upon an impressive trail of...

If you search for a software developer named Marcus Chen or a medical researcher named Elena Vasquez, you might stumble upon an impressive trail of professional credentials. You will find them authoring academic papers, running companies, and even being quoted in news rumors. There is only one catch: neither of them actually exists.
A fascinating new study from Samsung and the University of Warsaw has identified a systemic quirk in large language models: they are remarkably uncreative when it comes to naming people. When asked to invent characters or experts, AI models consistently default to a narrow set of "ghost" identities. ChatGPT loves the name "Elara Voss," Gemini prefers "Aris Thorne," and Claude frequently summons "Elena Vasquez."
What starts as an amusing AI habit is rapidly morphing into a serious contamination of the academic record. Researchers discovered over 1,600 fabricated papers uploaded to Zenodo, a highly respected scholarly repository operated by CERN. These synthetic documents are assigned real Digital Object Identifiers (DOIs) and are quietly indexed by trusted aggregators like Google Scholar and ResearchGate. The AI ghosts are literally building fake academic networks, co-authoring papers with one another across independently generated documents.
The consequences extend far beyond the ivory tower. The boundaries between synthetic text and verified reality are already blurring. For example, "Elena Vasquez" was recently listed as the founder of a completely AI-generated medical research company, and the exact same name was cited as a medical executive in a viral, debunked Facebook rumor.
For now, there is a silver lining. These recurring names act as accidental fingerprints, allowing researchers to trace synthetic "slop" back to specific AI models. If a text heavily features Elena Vasquez, it is highly likely it was generated by a Claude model. However, researchers warn that this forensic advantage is fleeting. As these ghost-authored papers flood the web, they are inevitably scraped into the training datasets for future AI models, creating a feedback loop of synthetic noise.
We are witnessing a critical stress test for the infrastructure of human knowledge. The automated systems we use to catalog and index scholarly work were built for an era where writing a fake paper took immense human effort. In an age of zero-cost synthetic replication, relying on automated aggregators without robust verification means we risk building our future knowledge on a foundation of ghosts.
Key Points
- AI models have strict naming biases, repeatedly generating the same fake names for experts and characters.
- Over 1,600 AI-generated papers using these ghost names have infiltrated legitimate academic repositories like Zenodo.
- These fabricated papers receive real DOIs and are indexed by Google Scholar, creating an illusion of academic legitimacy.
- While these names currently serve as "fingerprints" to identify AI content, this method will likely fail as AI-generated data is recycled into new training sets.
Why It Matters
The infiltration of AI 'ghosts' into academic databases highlights a critical vulnerability in how we verify and index human knowledge, threatening the integrity of scientific research and public information.
Sources:
更多专栏

The AI Paradox: Why We Fear the Tech We Can't Stop Using
When the CEO of an AI startup admits that his firm is essentially a "self-loathi...

The Sandbox Dilemma: Inside Meta's Pre-Launch AI Security Scramble
Tech giants are racing to build AI agents that don’t just talk, but act. Meta’s ...

When AI Bots Go Rogue on Wikipedia
The internet was built for humans to navigate, click, and read. But what happens...