Krithik Vishwanath, A. Alyakin, Mrigayu Ghosh, Ali Hage, S. Neifert, C. Orillac, Nataniel J. Mandelberg, Hammad A Khan et al.
This study demonstrates that state-of-the-art general-purpose large language models (LLMs) outperform specialized clinical AI tools on medical benchmarks, highlighting the need for independent evaluation before clinical deployment.
There is a lack of independent, quantitative evaluation of specialized clinical AI tools entering medical practice, particularly in comparison to frontier general-purpose LLMs in real-world clinical contexts.
A three-stage evaluation was conducted using three benchmarks: (1) 500 MedQA questions for medical knowledge, (2) 500 HealthBench items for alignment with clinicians, and (3) the Real Clinical Queries (RCQ) benchmark built from 100 de-identified physician queries. For the RCQ benchmark, 12 US clinicians performed randomized, blinded reviews of model outputs, producing 1,800 annotations.
Frontier LLMs (GPT-5.2, Gemini 3.1 Pro, Claude Opus 4.6) outperformed the specialized clinical AI tools (OpenEvidence, UpToDate Expert AI) across all three evaluations. The clinical AI tools performed comparably to an auto-enabled Google Search AI Overview on the RCQ benchmark. The findings underscore the critical need for independent, real-world evaluation of AI tools before they enter clinical settings, contributing a rigorous methodology for medical AI validation.