Solène Debuysère, Nicolas Trouvé, Nathan Letheule, Elise Colin, Georgia Channing
SARLO-80 is a large-scale multimodal dataset of high-resolution SAR SLC data aligned with optical imagery and natural language descriptions from 257 locations across 72 countries.
Existing SAR-optical datasets rely on low-resolution GRD products and do not preserve complex-valued SAR measurements or native acquisition geometry, limiting physically grounded multimodal learning. Large-scale public datasets combining VHR SAR SLC, aligned optical imagery, and text are lacking.
Collected ~2,500 Umbra spotlight SAR scenes (SICD format), resampled to an 80cm slant-range grid, and tiled into 1024×1024 patches. For each SAR patch, retrieved high-resolution optical tiles and warped them into the SAR grid using local coordinate correspondences for pixel-level alignment. Generated three caption variants (SHORT/MID/LONG) per sample.
Constructed 119,566 triplets (complex and amplitude SAR patches, aligned optical patches, natural language descriptions) covering 257 locations in 72 countries. Released fixed train/validation/test splits and full preprocessing and baseline code for reproducible benchmarks.