Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出PLANET,一种端到端的多目标跟踪器,通过嵌入3D场景几何信息和使用双分辨率时间记忆来解决单目视频中深度和空间关系模糊的问题。
📝 Abstract
Monocular videos record 3D scenes as sequences of 2D image-plane projections, obscuring depth and spatial relationships. Multi-object trackers localize and associate objects primarily using appearance and geometry observed only in the image plane, inheriting these ambiguities. To address this limitation, we introduce PLANET, an end-to-end multi-object tracker designed to move beyond the image plane. As an enabling step, we lift existing 2D tracking datasets into 3D. We then form world-grounded queries by embedding reconstructed 3D scene geometry into the features and positional encodings used during query formation. An auxiliary 3D location prediction task further encourages the queries to encode object positions during training. A complementary dual-resolution temporal memory preserves this evidence across longer temporal gaps. As a result, PLANET achieves state-of-the-art performance across three diverse benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Multi-Object Tracking
Image Plane
3D Scene Geometry
Ambiguities
Monocular Videos
Innovation

Methods, ideas, or system contributions that make the work stand out.

world-grounded queries
3D scene geometry
end-to-end multi-object tracker
dual-resolution temporal memory
🔎 Similar Papers
2023-10-27IEEE Transactions on Image ProcessingCitations: 4
💼 Related Jobs
No related jobs found.