BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
BLARM通过混合潜在刚性运动基元,从单目视频中生成3D物体动画,解决了高维顶点运动预测问题。
📝 Abstract
We introduce BLARM, a feed-forward method for video-driven 3D mesh animation. Given a monocular video and a static object mesh, BLARM predicts a temporally coherent animated mesh whose motion follows the video. Rather than relying on explicit rigs or directly regressing high-dimensional vertex motion, we represent animation using a compact set of learned, time-varying rigid motion components and time-invariant vertex-to-component skinning weights. This yields a low-dimensional deformation space without requiring skeletons, cages, skinning weights, or rig annotations. Our architecture conditions geometry-derived deformation latents on video features through factorized spatial-temporal attention, then decodes rigid transformations blended by predicted skinning weights. Trained with trajectory reconstruction, entropy regularization, and motion-aware contrastive learning, BLARM produces accurate and temporally stable animations while recovering compact, interpretable motion structure from monocular video.
Problem

Research questions and friction points this paper is trying to address.

video-driven 3D mesh animation
monocular video
temporally coherent
rigid motion components
Innovation

Methods, ideas, or system contributions that make the work stand out.

video-driven 3D mesh animation
latent rigid motion components
spatial-temporal attention
trajectory reconstruction
motion-aware contrastive learning
🔎 Similar Papers
No similar papers found.