Jisang Han, Seonghu Jeon, Jaewoo Jung, René Zurbrügg, Honggyu An, Tifanny Portela, Marco Hutter, Marc Pollefeys et al.
We propose the Geometric Action Model (GAM), which leverages a geometric foundation model (GFM) as a shared substrate for language-conditioned robot manipulation policy learning.
Existing VLAs and video world-action models operate on 2D image frames or 2D-derived latent spaces, failing to explicitly handle the 3D geometry required for contact-rich manipulation.
GAM splits a pretrained GFM at an intermediate layer: shallow layers serve as an observation encoder, and a causal future predictor inserted at the split layer forecasts future latent tokens conditioned on language, proprioception, and action history. The predicted tokens are then routed through the remaining GFM blocks for feature propagation and decoding, enabling a single backbone to produce both future geometry and actions.
Across simulation and real-robot manipulation benchmarks, GAM is more accurate, robust, faster, and lighter than current foundation-model-scale baselines.