Suhyun Lee, P. Achananuparp, Neemesh Yadav, Ee-Peng Lim, Yang Deng
A new evaluation framework is proposed to diagnose the safety of LLMs in mental health counseling at the interaction level, incorporating the AI counselor's role.
Existing evaluation methods focus on single responses or static datasets, limiting their ability to diagnose the accumulation and context of harm that emerges gradually over multi-turn interactions typical of counseling.
Large-scale evaluation of state-of-the-art LLMs revealed substantial role-dependent and cumulative safety failures systematically missed by existing static benchmarks. The framework significantly improves failure-mode coverage and diagnostic granularity.