Juncheng Ma, Jianxin Bi, Yufan Deng, Xuanran Zhai, Kewei Zhang, Ye Huang, Bo Liang, Shukai Gong et al.
Egocentric human video, when properly processed, can outperform real-robot teleoperation data for pretraining embodied foundation models.
Embodied foundation models require large-scale data, but real-robot teleoperation data is costly and lacks diversity. Egocentric human video is scalable and cheap but lacks action labels, raising questions about its effectiveness.
We systematically compare egocentric human video and teleoperated real-robot trajectories as pretraining sources under fixed post-training and validation protocols. Human video is processed with a filtering and labeling pipeline to generate action labels.
Models pretrained on egocentric data achieve 24% lower validation loss on real-robot action prediction, and 52.5% and 90% higher success rates on in-distribution and out-of-distribution tasks, respectively. This validates a scalable paradigm: pretrain on human video, then adapt with a small amount of robot data.