Jihan Yang, Shusheng Yang, Anjali Gupta, Rilyn Han, Fei-Fei Li, Saining Xie
We evaluate MLLMs' spatial cognition via the video-based VSI-Bench benchmark and find that generating cognitive maps improves spatial reasoning performance.
Humans possess visuospatial intelligence to remember and reason about spaces from sequential visual observations, but can MLLMs trained on video data 'think in space'? How can we measure and improve this ability?
We construct VSI-Bench with over 5,000 QA pairs and evaluate various MLLMs. We probe models to express spatial understanding linguistically and visually, analyzing bottlenecks. We compare existing linguistic reasoning techniques (CoT, self-consistency, ToT) with cognitive map generation.
MLLMs exhibit competitive but subhuman spatial intelligence, with spatial reasoning as the main bottleneck. Explicitly generating cognitive maps significantly improves spatial distance estimation, offering a new direction for evaluating and enhancing MLLMs' spatial intelligence.