← 深度专栏/原创观点
原创观点

The Digital Ghosts Haunting Our Academic Libraries

If you search for a software developer named Marcus Chen or a medical researcher named Elena Vasquez, you might stumble upon an impressive trail of...

潜
作者
潜龙编辑部
关注 AI 与社会议题
发布于
2026/10/6
READ
长读
The Digital Ghosts Haunting Our Academic Libraries
illustration · QianLong editorial

If you search for a software developer named Marcus Chen or a medical researcher named Elena Vasquez, you might stumble upon an impressive trail of professional credentials. You will find them authoring academic papers, running companies, and even being quoted in news rumors. There is only one catch: neither of them actually exists.

A fascinating new study from Samsung and the University of Warsaw has identified a systemic quirk in large language models: they are remarkably uncreative when it comes to naming people. When asked to invent characters or experts, AI models consistently default to a narrow set of "ghost" identities. ChatGPT loves the name "Elara Voss," Gemini prefers "Aris Thorne," and Claude frequently summons "Elena Vasquez."

What starts as an amusing AI habit is rapidly morphing into a serious contamination of the academic record. Researchers discovered over 1,600 fabricated papers uploaded to Zenodo, a highly respected scholarly repository operated by CERN. These synthetic documents are assigned real Digital Object Identifiers (DOIs) and are quietly indexed by trusted aggregators like Google Scholar and ResearchGate. The AI ghosts are literally building fake academic networks, co-authoring papers with one another across independently generated documents.

The consequences extend far beyond the ivory tower. The boundaries between synthetic text and verified reality are already blurring. For example, "Elena Vasquez" was recently listed as the founder of a completely AI-generated medical research company, and the exact same name was cited as a medical executive in a viral, debunked Facebook rumor.

For now, there is a silver lining. These recurring names act as accidental fingerprints, allowing researchers to trace synthetic "slop" back to specific AI models. If a text heavily features Elena Vasquez, it is highly likely it was generated by a Claude model. However, researchers warn that this forensic advantage is fleeting. As these ghost-authored papers flood the web, they are inevitably scraped into the training datasets for future AI models, creating a feedback loop of synthetic noise.

We are witnessing a critical stress test for the infrastructure of human knowledge. The automated systems we use to catalog and index scholarly work were built for an era where writing a fake paper took immense human effort. In an age of zero-cost synthetic replication, relying on automated aggregators without robust verification means we risk building our future knowledge on a foundation of ghosts.

Key Points

  • AI models have strict naming biases, repeatedly generating the same fake names for experts and characters.
  • Over 1,600 AI-generated papers using these ghost names have infiltrated legitimate academic repositories like Zenodo.
  • These fabricated papers receive real DOIs and are indexed by Google Scholar, creating an illusion of academic legitimacy.
  • While these names currently serve as "fingerprints" to identify AI content, this method will likely fail as AI-generated data is recycled into new training sets.

Why It Matters

The infiltration of AI 'ghosts' into academic databases highlights a critical vulnerability in how we verify and index human knowledge, threatening the integrity of scientific research and public information.


Sources:

潛
本文完
潜龙编辑部 · 2026/10/6