Hongjian Zhou, Xinyu Zou, Jinge Wu, Sean Wu, Junchi Yu, Bradley Max Segal, Tobias Erich Niebuhr, Sara Amro et al.
Introduces MedMisBench, a benchmark measuring the phenomenon where LLMs' accuracy plummets when misleading context is added to medical questions.
Existing medical LLM evaluations only measure how well models answer static questions, failing to assess whether they maintain correct judgment in situations mixed with misinformation, such as real patient consultations.
MedMisBench consists of 10,932 medical question items and 48,889 misleading context-option pairs. Misleading context types include authority-based falsehoods, exception poisoning, and alternative facts, covering areas such as medical reasoning, agentic capability, and patient-journey evaluation. Accuracy on original questions versus questions with misleading context was compared across 11 model configurations.
Average accuracy dropped from 71.1% on original questions to 38.0% under misleading context, with an attack success rate of 51.5%. Authority-based falsehoods achieved a 69.5% attack success rate. A clinical panel of 14 members from 7 countries identified serious potential harm in 38.2% of reviewed cases. This benchmark exposes a structural blind spot in evaluating LLM safety in medical settings.