Jiaju Han, Ma Yaqi, Yahui Chai, Xuemeng Sun, Xin Li, Qike Zhang, Yingying Zhao, Xiang Chen et al.
This paper presents MonoIR-RS, a dataset and benchmark for vision-language learning tailored to infrared remote sensing imagery.
Existing remote sensing vision-language models and datasets primarily focus on visible-band imagery, limiting the understanding of infrared imagery's unique intensity structure and contrast.
The authors constructed a dataset of 600,000 synthesized infrared images and 59,032 IR-aware captions. They fine-tuned five CLIP and six VLM backbones on this data and evaluated their performance against zero-shot baselines.
The study shows synthesized infrared imagery is closer to real thermal imagery than grayscale conversion. IR-aware adaptation improved CLIP mean recall by up to 12.8 points and achieved 100% IR-cue coverage in VLM captioning with near-zero RGB color leakage, providing a controlled testbed for IR vision-language alignment.