Karan Singhal, Tao Tu, Juraj Gottweis, R. Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark et al.
Med-PaLM 2 is a model that achieves physician-level medical question answering through LLM improvements, medical domain fine-tuning, ensemble refinement, and chain of retrieval for enhanced reasoning and grounding.
Existing LLMs had limitations in long-form medical question answering and handling real-world clinical workflows. While Med-PaLM surpassed the passing score on USMLE-style questions, reliability and safety in actual medical settings remained insufficient.
Med-PaLM 2 is based on PaLM 2, fine-tuned on medical domain data, and introduces ensemble refinement and chain of retrieval to improve reasoning and evidence provision. Additionally, a human evaluation framework was established to systematically assess physician preference and safety.
Med-PaLM 2 achieved 86.5% on MedQA (a 19% improvement over Med-PaLM) and showed significant performance gains on MedMCQA, PubMedQA, and MMLU clinical topics. In human evaluations, physicians preferred Med-PaLM 2 answers over those from other physicians on eight of nine clinical axes. In a real-world medical question pilot study, specialists preferred Med-PaLM 2 over generalist physician answers 65% of the time, and safety assessments rated it comparable to physician answers.