Arnav Kumar Jain, Yilin Wu, Jesse Farebrother, Gokul Swamy, Andrea Bajcsy
WEAVER is a multi-view world model that predicts future latents and rewards using a flow-matching loss, improving fidelity, consistency, and efficiency for robotic manipulation tasks.
Existing world models suffer from inconsistency over long horizons or slow speed, limiting their application in real robotics. A world model that jointly satisfies fidelity, consistency, and efficiency is needed.
WEAVER takes multi-view inputs and predicts future latents and rewards via a flow-matching loss. It distills key design decisions in model architecture, memory, and prediction objectives, specializing in long-horizon dynamic manipulation tasks.
In real robot experiments, WEAVER achieved a policy evaluation correlation of 0.870, a 38% improvement in policy improvement success rate, and a 14% improvement in test-time planning success rate with a 5-10x speedup over prior world models. It also outperformed prior world models in out-of-distribution scenarios.