🤖 AI Summary
This work addresses the challenge of anisotropic scaling of objects along their intrinsic coordinate axes in real-world videos, aiming to achieve geometrically plausible, temporally coherent, and background-consistent edits. The authors propose a two-stage progressive training framework that operates without mesh alignment or explicit 3D reconstruction. In the first stage, geometrically perturbed pseudo-source videos are generated through planar transformations guided by 3D deformations centered on the object. The second stage leverages both pseudo-source and original videos in a self-supervised manner to disentangle foreground geometric transformation from background preservation. This approach is the first to enable object scaling in real-scene videos without relying on 3D priors, introduces a dedicated evaluation benchmark, and outperforms existing methods in geometric consistency, foreground fidelity, and background preservation, while offering faster inference and greater practicality.
📝 Abstract
Geometry-aware video object scaling aims to anisotropically resize the object along object-centric axes while preserving geometric plausibility, temporal coherence, and background consistency. Existing text-guided methods mainly operate in the 2D image plane, while depth-guided approaches provide coarse control and mesh-based methods require costly 3D reconstruction. We present a progressive two-stage training framework that decouples geometry-aware foreground transformation from background preservation and realistic video composition, without mesh-pixel alignment and explicit 3D reconstruction at inference. In both stages, geometrically perturbed pseudo-sources are constructed from real videos, while the original complete videos are retained as reconstruction targets. The first stage uses planar transformations to learn robust foreground-background composition, whereas the second introduces object-centric 3D deformation guidance for geometry-aware scaling. This pseudo-source reconstruction formulation enables real-video synthesis without paired real-world scaling targets. We construct complementary paired-geometry and real-background benchmarks and further evaluate on in-the-wild videos. Extensive experiments demonstrate superior geometric consistency, foreground fidelity, and background preservation, together with faster and more practical inference than methods requiring explicit 3D reconstruction.