Jiaming Liu, Qingpo Wuwu, Nuowei Han, Hao Chen, Zhuoyang Liu, Fan Fei, Yueru Jia, Chenyang Gu et al.
A new VLA framework integrating 3D geometric and dynamic physical understanding for robotic manipulation.
Existing VLA models are limited in utilizing 3D information, suffering from geometric information loss and insufficient temporal action modeling in dynamic environments, hindering effective robotic manipulation.
Extending the Lift3D model, it directly encodes 3D point clouds into the VLA vision encoder. Geometry-Centric Masked Autoencoding (GC-MAE) learns both 3D structure and physical dynamics simultaneously. Layer-wise temporal action modeling using multiple LLM layers predicts consistent action chunks.
Achieves 10.8% and 11.1% higher mean success rates on MetaWorld and RLBench respectively over prior best VLA methods, and outperforms the strongest real-world baseline by 4 percentage points, establishing a new benchmark for 3D-aware robotic manipulation.