Xiaohang Ren, Chenxiao Fan, Wenyin Ma, Hongliang He, Chongming Gao, Xiaoyan Zhao, Fuli Feng
A comprehensive survey that systematizes LLMs' medical reasoning abilities based on cognitive theory and empirically demonstrates the gap between exam performance and actual clinical capability using a real-world clinical data benchmark.
Although LLMs achieve high performance on medical exam questions, real clinical decision-making is safety-critical, context-dependent, and involves evolving evidence. Thus, robust medical reasoning beyond factual recall is necessary, but existing evaluations focus on exams and fail to reflect actual clinical abilities.
Grounded in cognitive theories of clinical reasoning (iterative process of abduction, deduction, and induction), the paper conceptualizes medical reasoning and organizes existing methods into seven major technical routes spanning training-based and training-free approaches. It also conducts a unified cross-benchmark evaluation of representative medical reasoning models under consistent experimental settings and introduces MR-Bench, derived from real hospital data, to assess clinical decision-making.
The paper provides a systematic taxonomy and unified evaluation framework for existing methods. Evaluations on MR-Bench demonstrate a critical gap: LLMs achieve high scores on exam-level tasks but significantly lower accuracy on real clinical tasks. This suggests that future medical AI research should focus on improving actual clinical reasoning abilities.