Researchers discovered a 'ghost couple' phenomenon where LLMs repeatedly generate specific fictional character pairs. These ghost names appear in hundreds of AI-generated documents and are exploited to create fake academic records with real DOIs on Zenodo and ResearchGate.
When LLMs generate fictional expert names, specific name pairs (e.g., Claude: Elena Vasquez + Marcus Chen; Gemini: Aris Thorne + Lena Petrova) appear with high correlation across independent generations. The researchers term this the 'ghost couple' phenomenon and found that these names were used to create 1,655 fake academic records on Zenodo (a CERN-operated repository) with real DataCite DOIs, some with deliberately backdated timestamps.
LLMs do not merely default to high-probability individual names but produce correlated character ensembles that are model-family-specific and version-specific. These 'ghost authors' contaminate the academic publishing ecosystem, undermining trust in scholarly records and injecting garbage information into databases. The Zenodo case demonstrates how AI-generated content can threaten the integrity of academic metadata.
This study highlights the negative impact of LLM-generated content on the academic ecosystem. Publishers and repositories need to develop methods to detect and filter AI-generated author names. Model providers should also mitigate the 'ghost couple' phenomenon by improving name generation prompts. This raises important issues at the intersection of AI safety and academic ethics.