MAGneT-3D: Monocular and Domain-Generalizable Temporal 3D Detection

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the poor cross-domain generalization in monocular temporal 3D detection caused by learnable query overfitting. We propose the first domain generalization method for this task, introducing a Domain-Robust Anchor Generator (DRAG) and Temporal Refinement Identity Merging (TRIM) strategy to effectively decouple spatial distribution dependencies and optimize temporal associations. Furthermore, we establish a comprehensive cross-dataset generalization benchmark. Experimental results demonstrate that our approach improves zero-shot cross-domain NDS from 12.1% to 18.6% while simultaneously enhancing in-domain accuracy. By comprehensively outperforming existing baselines, this work significantly strengthens model robustness and generalization capabilities across diverse domains.
📝 Abstract
Monocular temporal 3D detection aims to detect objects in 3D, given a monocular video. Query-based 3D detectors unify detection and cross-view association, but their learnable queries fit the spatial distribution of the training data (e.g., field-of-view). We show that this issue is especially severe when these models are applied to monocular video, hindering generalization to unseen datasets and environments. To address this limitation, we introduce MAGneT-3D, the first method for domain-generalized monocular temporal 3D object detection. Instead of relying on static learnable queries, we propose a Domain-Robust Anchor Generator (DRAG) approach that adaptively derives 3D proposals during inference. To further enable domain generalization, we propose a Temporal Refinement and Identity Merging (TRIM) strategy, reducing dependence on specific 3D proposals. To enable comprehensive domain-generalization evaluation, we establish a cross-dataset benchmark spanning nuScenes, Waymo, Lyft, and ONCE. Under zero-shot domain shifts, MAGneT-3D outperforms all baselines, improving NDS from 12.1% to 18.6% while also increasing in-domain accuracy.
Problem

Research questions and friction points this paper is trying to address.

Monocular Temporal 3D Detection
Domain Generalization
Learnable Queries
Zero-shot Domain Shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Domain-Generalizable Monocular 3D Detection
Domain-Robust Anchor Generator
Temporal Refinement and Identity Merging
Zero-Shot Domain Shift
Cross-Dataset Benchmark
🔎 Similar Papers
No similar papers found.