Yalun Dai, Hao Li, Shulin Tian, Runmao Yao, Yuhao Dong, Fangzhou Hong, Zhaoxi Chen, Fangfu Liu et al.
S-Agent is a spatial tool-use agentic paradigm that uses VLM as a semantic planner and leverages 2D detection and 3D lifting tools to accumulate spatio-temporal evidence for spatial intelligence.
Existing VLMs and tool-augmented agents rely on static, stateless inference from isolated visual observations, failing to handle continuous and evolving 3D worlds.
S-Agent casts VLM as a semantic planner to decide what evidence is needed, using a hierarchy of spatial tools (2D detection, 3D lifting, spatial knowledge aggregation) to accumulate evidence. Temporal memory mechanisms (Scene Memory and Agent Memory) integrate evidence across frames and reasoning steps.
On multi-view and video spatial reasoning benchmarks, S-Agent consistently improves both open-source and closed-source VLMs without training. Fine-tuned on S-Agent-generated trajectories (300K), S-Agent-8B significantly surpasses similar-scale baselines (e.g., Qwen3-VL-8B) and performs comparably to advanced closed-source models (e.g., GPT-5.4 and Gemini 3).