General-purpose LLMs (GPT-5.2, Gemini 3.1 Pro, Claude Opus 4.6) outperformed specialized clinical AI tools (OpenEvidence, UpToDate Expert AI) on MedQA, HealthBench, and real clinical query benchmarks. The study involved 12 US clinicians conducting blinded evaluations across 1,800 annotations in a three-stage assessment.
A study published in Nature Medicine found that general-purpose LLMs (GPT-5.2, Gemini 3.1 Pro, Claude Opus 4.6) outperformed specialized clinical AI tools (OpenEvidence, UpToDate Expert AI) on MedQA, HealthBench, and a real clinical query benchmark.
Specialized clinical AI tools claim superior performance due to domain-specific training or RAG, but their architectures and training pipelines are not public. This study conducted a three-stage evaluation with 1,800 blinded annotations from 12 US clinicians to test this claim.
The findings suggest that specialized clinical AI tools may not necessarily outperform general-purpose LLMs in real clinical settings, highlighting the need for independent, real-world evaluation before clinical deployment.