Pradhaan S Bhat, Naveen Chandra R, Rishubh Parihar, Vaibhav Vavilala, R. Venkatesh Babu, D. A. Forsyth, Anand Bhattad
A method that uses 3D boxes as structured specifications for large 3D edits in real images, leveraging a depth-aligned floor as a global reference.
Existing text or 2D conditioning provides weak control over spatial transformations like large object motions and camera changes. Prior 3D primitive methods only indicate approximate location, not precise transformations.
Users specify input and output 3D boxes, casting editing as a well-posed geometry problem. Each box face is color-coded for 3D orientation. A depth-aligned planar floor serves as a global reference frame with depth-aware shading. An image generator conditioned on this structure produces consistent results. Training is two-stage: on synthetic multi-object scenes and a small set of real videos from Objectron.
The method operates directly on real photographs and substantially outperforms recent state-of-the-art methods on large 3D edits, preserving scene and object identity while recovering unseen regions.