Gal Fiebelman, Hadar Averbuch-Elor, Sagie Benaim
This paper proposes MV-Forcing, a framework that generates geometrically consistent multi-view videos of dynamic scenes at arbitrary lengths and viewpoint counts using a single diffusion model by introducing a 4D geometric bridge between sequentially generated views.
While recent video diffusion models can generate long single-view videos or short multi-view videos, generating long, multi-view consistent videos of dynamic scenes remains an unsolved challenge.
The framework reconstructs the 3D structure of a completed source view and renders a geometric prior for the next target viewpoint, which the diffusion model then refines into a high-quality video. A joint denoising regime initializes both view slots from noise during training to enable temporally unbounded generation. The model is distilled via Distribution Matching Distillation with Spatio-Temporal Self-Forcing to close the train-inference exposure bias gap.
Extensive experiments on synthetic and real-world data demonstrate that MV-Forcing produces geometrically consistent multi-view videos of dynamic scenes at arbitrary lengths and viewpoint counts using a single few-step student model.