Guli Zhu, Chenwei Wu, Liyue Shen
This paper introduces M3Bench, a clinically grounded benchmark to systematically evaluate the reliability and generalizability of model editing techniques for medical vision-language models post-deployment.
Existing multimodal model editing benchmarks focus on general tasks and do not adequately consider the specific challenges of the medical domain, such as image and text variation, modality shifts, and clinical knowledge composition, making it difficult to validate model stability in real-world clinical settings.
The authors developed M3Bench, comprising 16,276 questions across diverse anatomy, modalities, and specialties, supporting both single and sequential edits. They systematically evaluated four representative editing methods (gradient-based and memory-based) across six medical and general vision-language models.
The evaluation found that no single method excelled across all criteria. Gradient-based editors showed strong transfer but suffered from catastrophic locality violations, while memory-based methods preserved locality but lacked compositional generality. The study attributes these failures to the latent space geometry of VLMs and how editing methods alter it, providing actionable guidance for safer post-deployment adaptation.