Back to the Feature: Zero-Shot 6DoF Pose Estimation via Dense Local Features

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
提出B2TFPose,一种无需训练的零样本6DoF位姿估计方法,使用DINOv3视觉转换器提取密集局部特征,通过合成到真实域泛化,实现精确位姿估计。
📝 Abstract
We present B2TFPose, a training-free zero-shot method for 6DoF pose estimation of unseen objects from RGB images. Using a single frozen DINOv3 vision transformer as its only pretrained component within the pose estimation pipeline, B2TFPose extracts dense patch-level features that generalize across the synthetic-to-real domain gap without any task-specific fine-tuning, revisiting the classical local feature matching paradigm through the lens of large-scale self-supervised foundation models. Three contributions advance the training-free state of the art. A geodesic non-maximum suppression strategy retrieves a viewpoint-diverse template set for coarse-to-fine correspondence matching. Render-guided Re-Correspondence (RRC) synthesizes object-specific views at the estimated pose and re-establishes dense 2D-3D correspondences to sharpen the initial estimate without additional learned parameters. A multi-mask hypothesis selection strategy jointly scores competing segmentation candidates to resolve segmentation ambiguity. On the seven core datasets of the BOP Benchmark, B2TFPose achieves 40.7 mean AR without refinement and 56.4 with refinement, establishing state-of-the-art performance among training-free RGB methods and outperforming trained counterparts including GigaPose and GenFlow, at competitive inference speed.
Problem

Research questions and friction points this paper is trying to address.

6DoF Pose Estimation
Zero-Shot
Dense Local Features
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Shot 6DoF Pose Estimation
Dense Local Features
Geodesic Non-Maximum Suppression
Render-Guided Re-Correspondence (RRC)
Multi-Mask Hypothesis Selection
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
A
Ali Rafiaei
Dept. Electrical & Computer Engineering, Ingenuity Labs, Queen’s University, Kingston, ON, Canada
Michael Greenspan
Michael Greenspan
Professor of Electrical and Computer Engineering, Queen’s University
computer vision