Timothy Agboada, Shikha Chandel, Yadav Raj Ghimire, Leila Hashemi-Beni
A PEFT framework injecting lightweight adapters into three VLM architectures for remote sensing VQA, showing hybrid FLAVA achieves the best balance.
Direct application of general-domain VQA models to remote sensing is hindered by domain shift and high computational cost of full fine-tuning.
Inject lightweight bottleneck adapters into attention and MLP layers of frozen CLIP, BLIP, and FLAVA backbones, training less than 5% of parameters on RSVQA-x dataset.
All adapted models converge; hybrid FLAVA offers superior multimodal reasoning and retrieval. Establishes a new baseline for resource-efficient VQA in disaster assessment and urban monitoring.