Jiawei Chen, Ruoxi Xu, Boxi Cao, Ruotong Pan, Yunfei Zhang, Yifei Hu, Yong Du, Tingting Gao et al.
We construct the first user simulation benchmark based on real human behavior data, integrating long-horizon, cross-scenario, and heterogeneous behaviors, and empirically demonstrate that current LLMs fail to simulate realistic behaviors.
Existing user simulation benchmarks are limited to isolated scenarios, narrow action spaces, or synthetic data, failing to capture the holistic nature of real human behavior (e.g., long-term causal chains, cross-scenario decision-making, individual differences).
We collect real-world data (actual user behavior logs) to build the OmniBehavior benchmark, which includes long-horizon (up to 30+ days), cross-scenario (behaviors across multiple apps/services), and heterogeneous (diverse behavior types) patterns. We evaluate state-of-the-art LLMs (GPT-4, Claude, etc.) on simulation accuracy and systematically compare simulated vs. authentic behaviors.
LLMs exhibit performance plateaus in long-horizon, cross-scenario behavior simulation and show a positive bias (hyper-activity, persona homogenization, utopian bias), losing individual differences and long-tail behaviors. This highlights directions for high-fidelity simulation research.