Paul Hyunbin Cho, Jinhyuk Jang, SeokYoung Lee, Joungbin Lee, Siyoon Jin, Heeseong Shin, Jung Yi, Yunjin Park et al.
To overcome the slow inference speed of diffusion-based lip sync models, we propose an autoregressive diffusion method (Lip Forcing) that distills a 14B teacher model to enable real-time lip sync with only two denoising steps.
Existing diffusion-based lip sync models are impractical for real-time inference due to full-sequence bidirectional attention and many denoising steps.
Distill a 14B audio-conditioned bidirectional video diffusion teacher into causal students, and design Sync-Window DMD leveraging a lip-sync-specific teacher-trajectory analysis, a two-step inference schedule without CFG, and a SyncNet-based reward.
The 1.3B student achieves real-time streaming at 31 FPS, 17.6× faster than its same-scale bidirectional model. The 14B student runs 39.8× faster than its teacher at comparable reference fidelity, with sub-millisecond time-to-first-frame.