Jin Yao, Dhruva Dixith Kurra, Tom Lampo, Zezhou Cheng, Danhua Guo, Burhan Yaman
VLGA is the first method to add geometry as a fourth modality to autonomous driving VLA models, learning dense 3D spatial understanding through LiDAR pointmap reconstruction.
Existing VLA models can reason in language but lack grounding in 3D space. Approaches that inject 3D foundation features or use sparse box/map losses fail to ensure the policy effectively utilizes geometric information.
VLGA introduces geometry as a fourth modality alongside vision, language, and action. It designs a dedicated geometry expert supervised by a per-pixel pointmap regression loss against LiDAR, providing dense spatial signals and forcing the agent to reconstruct the 3D world.
On open-loop nuScenes, VLGA achieves state-of-the-art among VLA methods without ego status (L2 error 0.50m, 3-second collision rate 0.18%). On closed-loop Bench2Drive, it attains a driving score of 79.08, +0.71 over prior VLA, with comparable efficiency and comfort.