Ipek Oztas, Duygu Ceylan, Aybars Bugra Aksoy, Aysegul Dundar
This paper proposes a method to move objects within an image in a geometry-consistent manner by extending the RoPE positional embeddings of diffusion models to be depth-aware.
Relocating objects in a single image while maintaining scene-level geometric consistency—such as handling occlusions, generating new regions, and preserving coherent shadows and reflections—is challenging for existing methods.
The method extends the RoPE positional embeddings of diffusion transformers into a depth-aware formulation that encodes 3D spatial structure. This enables controlled object motion and scene-aware updates. The model is trained using synthetic data combined with a small set of real images via parameter-efficient fine-tuning.
The approach demonstrates state-of-the-art performance across all evaluation metrics on standard object motion benchmarks. Despite minimal real supervision, it preserves object identity under large spatial displacements, generates plausible content in newly revealed regions, and consistently updates scene-dependent effects like shadows and illumination.