Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky, Thomas L. Griffiths
This study empirically demonstrates that value-aligned LLMs, which appear unbiased on standard bias benchmarks, strongly learn social stereotypes when measured using methods similar to the psychological Implicit Association Test (IAT).
LLM bias evaluation mainly relies on explicit questions or standard datasets, which can allow models to hide biases by learning socially desirable answers. Therefore, methods to measure implicit biases revealed in actual model behavior are needed.
Word Association and Sentence Completion tasks, analogous to the psychological IAT, were designed to measure how strongly LLMs associate specific social groups (race, gender, religion, health) with negative/positive attributes. Eight recent value-aligned LLMs (GPT-4, Claude, etc.) were evaluated on 21 stereotypes (e.g., Black-criminal, female-science).
Implicit biases consistent with social stereotypes were found in all models, independent of their scores on standard bias benchmarks. Furthermore, these biases significantly affected the models' discriminatory decisions (e.g., simulated hiring, sentencing predictions). This study introduces psychological methodology to LLM bias evaluation, overcoming limitations of existing assessments, and emphasizes the need for behavior-based bias measurement.