Bohan Liu, Wenqian Ye, Guangzhi Xiong, Zhenghao He, Sanchit Sinha, Aidong Zhang
This paper proposes a training-free mechanistic interpretability method to analyze and mitigate the Typographic Attack vulnerability in CLIP models, achieving significant robustness improvements by intervening on specific attention heads.
Vision-language models like CLIP are vulnerable to Typographic Attacks, where irrelevant text in images biases visual representations towards lexical meaning, posing safety risks in critical applications. Existing defenses often lack interpretability or effectiveness.
The authors introduce a training-free mechanistic interpretability approach. They use sampling-based interpretations to analyze hidden states and quantitatively attribute semantic vs. lexical focus to individual attention heads. Through probabilistic analysis and circuit mining, they isolate specific Vision Transformer (ViT) components that disproportionately encode lexical information. Simple interventions, such as selectively adjusting attention weights on these identified circuits, are then applied.
The interventions substantially improve robustness against Typographic Attacks in object classification without additional training, outperforming both supervised and training-free defense methods. Applying the intervention to the vision encoders of state-of-the-art LVLMs yields substantial gains in Visual Question Answering accuracy under attack on the RIO-Bench, confirming the efficacy and generalizability of the mechanistic approach.