Haoyu Zhao, Xingyue Zhao, Siteng Huang, Xin Li, Deli Zhao, Zhongyu Li
This work proposes a 4D generative world model and a derived policy framework for robotic manipulation that jointly predicts the visual, geometric, and temporal dynamics of a scene.
Open-world robotic manipulation requires anticipating how a scene's 3D structure moves under interaction, not just recognizing its appearance. Existing 2D video-based approaches struggle to bridge the gap with low-level robot end-effector actions.
The method leverages synchronized RGB, depth, and optical flow (RGB-DF) as a physically grounded 4D representation. It introduces RynnWorld-4D, a unified diffusion process that generates future frames from a single RGB-D image and language instruction, trained on a large-scale curated dataset (Rynn4DDataset 1.0). An inverse dynamics policy, RynnWorld-4D-Policy, is proposed to directly output robot actions from the model's internal 4D representations.
The model produces temporally and spatially coherent 4D predictions. The derived policy achieves state-of-the-art performance on real-world dexterous bimanual manipulation tasks, particularly excelling in those requiring spatial precision and temporal coordination, thereby narrowing the gap between world prediction and policy learning.