Marcel Binz, Elif Akata, Abdullah Almaatouq, Mohammed Alsobay, Oleksii Ariasov, Franziska Brandle, David Broska, Jason W. Burton et al.
This empirical study finds that post-training, the process of making LLMs useful assistants, degrades their performance in modeling human behavior.
LLMs are increasingly used as human surrogates, but it is unclear which models best capture human behavior and why.
The researchers introduce a novel dataset, Psych-201, to measure behavioral alignment with humans at scale across various model families, sizes, and objectives.
Post-training consistently reduces alignment with human behavior, and this misalignment widens in newer model generations. Furthermore, persona-induction techniques do not improve individual-level predictions. The results suggest that the very processes used to make LLMs useful assistants may also make them less accurate models of human behavior.