Object Concepts Emerge from Motion

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过从原始视频中学习基于运动边界的对象中心表示来解决现有视觉预训练方法不能保持单个实例身份和一致性的问题。
📝 Abstract
Object-centric visual representations are important for physical-world perception, but existing visual pretraining methods often capture semantic categories without preserving the identity and coherence of individual instances. We present a biologically inspired framework that learns object-centric representations for single images from raw videos. Our approach uses motion boundaries as a source of object-level grouping: off-the-shelf optical flow and clustering produce pseudo-instance masks, which supervise a single-image encoder with pixel-level pairwise metric learning. The framework requires neither human annotations nor camera calibration. We first obtain 195 million pseudo-labeled frames from 7,163 hours of driving and web videos, then expand the supervision to 421 million frames with Motion-Verified Self-Training, which combines model proposals with motion evidence. We train encoders up to Swin-H and distill the learned representations into a family of Swin backbones. Across monocular depth estimation, 3D object detection, 3D occupancy prediction, and end-to-end planning, the resulting models achieve competitive or superior performance relative to supervised and self-supervised pretraining baselines, with particularly strong transfer on geometry- and instance-sensitive tasks. These results show that motion-derived supervision can teach static image encoders to represent visual instances, providing a complementary direction for scalable visual pretraining.
Problem

Research questions and friction points this paper is trying to address.

object-centric representations
visual pretraining
instance identity
physical-world perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

object-centric representations
motion boundaries
pseudo-instance masks
pixel-level pairwise metric learning
Motion-Verified Self-Training
🔎 Similar Papers
No similar papers found.