Foundational feature fusion for conditional flow matching in 6D pose estimation

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出FunFlow6D,利用几何与外观基础模型特征解决6D姿态估计问题,通过交叉注意力机制融合特征,减少对特定任务编码器训练的需求,提高性能。
📝 Abstract
Conditional flow matching has enabled a step forward in object 6D pose estimation, achieving state-of-the-art performance by progressively denoising and registering object representations to observed scenes. Existing methods require training task-specific encoders supervised on object-scene overlap and rely on trivial feature fusion strategies to resolve pose ambiguities. We present FunFlow6D, a novel flow matching-based formulation that leverages features from geometric and appearance foundation models for pose estimation, eliminating the need for task-specific encoder training. We also introduce a cross attention-based fusion mechanism that dynamically combines geometric and appearance features to provide richer conditioning for the flow matching module. Experiments on four datasets from the BOP benchmark show that FunFlow6D outperforms the previous state of the art while reducing supervision requirements and memory overhead. Extensive ablations validate the contribution of each proposed component. Project website: https://tev-fbk.github.io/FunFlow6D/.
Problem

Research questions and friction points this paper is trying to address.

6D pose estimation
conditional flow matching
feature fusion
task-specific encoders
Innovation

Methods, ideas, or system contributions that make the work stand out.

flow matching
feature fusion
cross attention
foundation models
6D pose estimation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.