Ramaneswaran Selvakumar, Kaousheik Jayakumar, Sakshi Singh, Sreyan Ghosh, Ruohan Gao, Dinesh Manocha
Analyzing the internal mechanisms of Audio-Visual Large Language Models (AVLLMs) reveals a bias where visual information suppresses audio information, stemming from training data imbalance.
It was unclear whether AVLLMs, which integrate audio and visual information to generate text, actually utilize both modalities in a balanced manner, especially when audio conflicts with vision.
As the first mechanistic interpretability study, we analyzed how audio and visual features evolve and fuse through different layers of an AVLLM. Probing techniques were used to evaluate audio information in intermediate layers and compare it with the final output.
Useful audio information exists in intermediate layers, but deeper fusion layers exhibit a bias where visual representations suppress audio. Additionally, AVLLM audio behavior closely matches its vision-language base model, suggesting limited additional alignment to audio supervision. This study provides new mechanistic insights into modality bias in multimodal LLMs.