Bardienus Pieter Duisterhof, Deva Ramanan, Jeffrey Ichnowski, Justin Johnson, Keunhong Park
We propose a simple and scalable method for joint image-depth generation using only sparse depth data by leveraging spatial priors from text-to-image (T2I) models.
Existing monocular depth estimation methods require high-quality dense depth data and complex training procedures. Prior attempts to use T2I spatial priors for depth estimation mostly rely on dense depth data and lack scalability.
Modality Forcing assigns different noise levels to each modality (image, depth), enabling conditional and joint generation. This allows training on sparse real-world depth data, with separate decoders for each modality. We also train T2I models of varying sizes (370M to 3.3B parameters) from scratch to verify scalability.
The proposed method reduces AbsRel error by 57% compared to existing joint image-depth generative models and achieves performance competitive with state-of-the-art monocular depth estimators. This strongly suggests that image generation is a scalable pre-training objective for spatial perception.