MVTOP: Multi-View Transformer-based Object Pose-Estimation

📅 2025-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
In multi-view rigid object pose estimation, single-view methods suffer from pose ambiguity caused by occlusion and symmetry. To address this, we propose an end-to-end multi-view fusion framework grounded in line-of-sight (LOS) geometric modeling. Unlike conventional post-hoc fusion or depth-dependent approaches, our method integrates multi-view images early in feature extraction and incorporates camera intrinsics and relative pose priors to guide spatial reasoning via a novel LOS-aware attention mechanism. Crucially, it achieves global multi-view pose estimation without requiring depth supervision—a first in the literature. Extensive experiments on our synthetic dataset and the YCB-Video benchmark demonstrate significant improvements over both single-view baselines and state-of-the-art multi-view methods, validating robustness under ambiguous conditions and strong generalization capability.

Technology Category

Application Category

📝 Abstract
We present MVTOP, a novel transformer-based method for multi-view rigid object pose estimation. Through an early fusion of the view-specific features, our method can resolve pose ambiguities that would be impossible to solve with a single view or with a post-processing of single-view poses. MVTOP models the multi-view geometry via lines of sight that emanate from the respective camera centers. While the method assumes the camera interior and relative orientations are known for a particular scene, they can vary for each inference. This makes the method versatile. The use of the lines of sight enables MVTOP to correctly predict the correct pose with the merged multi-view information. To show the model's capabilities, we provide a synthetic data set that can only be solved with such holistic multi-view approaches since the poses in the dataset cannot be solved with just one view. Our method outperforms single-view and all existing multi-view approaches on our dataset and achieves competitive results on the YCB-V dataset. To the best of our knowledge, no holistic multi-view method exists that can resolve such pose ambiguities reliably. Our model is end-to-end trainable and does not require any additional data, e.g., depth.
Problem

Research questions and friction points this paper is trying to address.

Resolves pose ambiguities in multi-view object estimation
Uses early fusion of view-specific features for accuracy
Handles varying camera parameters without additional data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Early fusion of multi-view features
Transformer-based pose estimation
Lines of sight geometry modeling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Lukas Ranftl
MVTec Software GmbH, Technical University of Munich
F
Felix Brendel
MVTec Software GmbH
B
Bertram Drost
MVTec Software GmbH
Carsten Steger
Carsten Steger
Director of Research, MVTec Software GmbH, and Professor of Computer Science, TU München
Machine VisionComputer VisionPhotogrammetry