LLMs exhibit correlated name priors, consistently generating specific fictional expert pairs across independent queries. These ghost names have infiltrated academic platforms like Zenodo and ResearchGate with real DOIs, polluting scholarly records.
Researchers discovered that LLMs exhibit correlated name priors: when generating fictional experts, specific name pairs (e.g., Elena Vasquez and Marcus Chen for Claude) consistently co-occur across independent queries, varying by model family and version. These ghost names have been used to create 1,655 fake academic records on Zenodo with real DOIs, and synthetic research groups on ResearchGate.
The proliferation of LLM-generated content on the web and in academic publishing has been a growing concern. This study goes beyond individual name frequency to reveal that name correlations are model-family-specific and deliberately suppressed at version boundaries, leaving temporal fingerprints. It also documents large-scale infiltration of ghost authors into real scholarly databases.
This demonstrates a concrete risk of LLM-generated content contaminating academic records. Records with real DataCite DOIs can be harvested by scholarly aggregators and cited, undermining research integrity. It highlights the urgent need for verification systems and policies to protect scholarly publishing from synthetic content.