LLMs exhibit correlated name priors, repeatedly generating specific fictional expert names like Elena Vasquez and Marcus Chen. This has led to thousands of ghost-authored records on Zenodo and ResearchGate, contaminating academic publishing.
LLMs exhibit correlated name priors, consistently generating specific fictional expert pairs (e.g., Elena Vasquez and Marcus Chen) across model families. This led to 1,655 ghost-authored records on Zenodo and synthetic research groups on ResearchGate.
LLMs reflect statistical patterns in training data when generating names, causing certain name combinations to appear excessively. Researchers found these priors are model-family-specific and version-specific, and are actively suppressed at model release boundaries, providing identifiable fingerprints.
This phenomenon threatens academic publishing integrity. Fake papers can receive real DOIs and be indexed in scholarly databases, potentially distorting research evaluation and meta-analyses. It highlights the need for detecting and regulating AI-generated content proliferation.