Qiaowei Miao, Kehan Li, Yawei Luo, Yi Yang
A flexible framework called Align4D is proposed to generate coherent 4D content (video-3D pairs) from any-modal input (X) by aligning it with video and 3D priors.
Existing 4D generation methods are often modality-specific or require expensive dataset construction for arbitrary user-defined X-to-4D generation, limiting their scalability.
Introduces the X4D dataset for benchmarking. Experiments on X4D and Consistent4D demonstrate that Align4D achieves state-of-the-art quality and consistency in X-to-4D generation.