S. Sandmann, Stefan Hegselmann, Michael Fujarski, Lucas Bickmann, Benjamin Wild, Roland Eils, J. Varghese
Open-source DeepSeek LLMs demonstrated clinical decision-making performance on par with proprietary models, proving the feasibility of privacy-compliant medical AI.
Proprietary LLMs like GPT-4o, despite excellent performance, are cloud-based and cannot be deployed on-site in hospitals, making it difficult to comply with stringent medical privacy regulations (e.g., HIPAA, GDPR). Therefore, it was necessary to verify whether open-source LLMs can serve as an alternative and how effective they are in real clinical decision support compared to proprietary models.
The research team compared the diagnostic and treatment recommendation accuracy of DeepSeek-V3, DeepSeek-R1, GPT-4o, and Gemini-2.0 Flash Thinking Experimental using 125 standardized patient cases (including rare diseases). With a sufficient sample size for statistical power, each model's responses were scored by specialist evaluation.
DeepSeek models performed equivalently to proprietary models with no statistically significant difference, and in some tasks showed better results. This empirically demonstrates that open-source LLMs can ensure data privacy through local hospital deployment while providing high-performance clinical support. This is a significant contribution to regulatory compliance and scalability of medical AI.