Rahul K. Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, J. Q. Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah et al.
HealthBench is an open-source benchmark that evaluates the performance and safety of LLMs in healthcare using multi-turn conversations and physician-validated criteria.
Existing medical LLM benchmarks rely on multiple-choice or short-answer formats, failing to capture the complexity of real medical conversations and lacking safety evaluation.
262 physicians developed 48,562 unique evaluation criteria for 5,000 multi-turn conversations to assess responses. The benchmark covers various contexts (e.g., emergencies, clinical data transformation, global health) and behavioral dimensions (e.g., accuracy, instruction following, communication).
Performance improved from 16% (GPT-3.5 Turbo) to 32% (GPT-4o) and 60% (o3). GPT-4.1 nano outperforms GPT-4o while being 25 times cheaper. Two variations, HealthBench Consensus and HealthBench Hard, are also released to provide benchmarks for medical LLM development.