ScaleVid: Geometry-Aware Video Object Scaling with Mesh-Free Inference

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of anisotropic scaling of objects along their intrinsic coordinate axes in real-world videos, aiming to achieve geometrically plausible, temporally coherent, and background-consistent edits. The authors propose a two-stage progressive training framework that operates without mesh alignment or explicit 3D reconstruction. In the first stage, geometrically perturbed pseudo-source videos are generated through planar transformations guided by 3D deformations centered on the object. The second stage leverages both pseudo-source and original videos in a self-supervised manner to disentangle foreground geometric transformation from background preservation. This approach is the first to enable object scaling in real-scene videos without relying on 3D priors, introduces a dedicated evaluation benchmark, and outperforms existing methods in geometric consistency, foreground fidelity, and background preservation, while offering faster inference and greater practicality.
📝 Abstract
Geometry-aware video object scaling aims to anisotropically resize the object along object-centric axes while preserving geometric plausibility, temporal coherence, and background consistency. Existing text-guided methods mainly operate in the 2D image plane, while depth-guided approaches provide coarse control and mesh-based methods require costly 3D reconstruction. We present a progressive two-stage training framework that decouples geometry-aware foreground transformation from background preservation and realistic video composition, without mesh-pixel alignment and explicit 3D reconstruction at inference. In both stages, geometrically perturbed pseudo-sources are constructed from real videos, while the original complete videos are retained as reconstruction targets. The first stage uses planar transformations to learn robust foreground-background composition, whereas the second introduces object-centric 3D deformation guidance for geometry-aware scaling. This pseudo-source reconstruction formulation enables real-video synthesis without paired real-world scaling targets. We construct complementary paired-geometry and real-background benchmarks and further evaluate on in-the-wild videos. Extensive experiments demonstrate superior geometric consistency, foreground fidelity, and background preservation, together with faster and more practical inference than methods requiring explicit 3D reconstruction.
Problem

Research questions and friction points this paper is trying to address.

geometry-aware scaling
video object scaling
temporal coherence
background consistency
geometric plausibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

geometry-aware scaling
mesh-free inference
video object manipulation
pseudo-source reconstruction
3D deformation guidance
🔎 Similar Papers