Pose Anything Anywhere:Model-free Object Poses from Arbitrary References
This work addresses the challenging problem of 6D object pose estimation for unseen objects in open-world scenarios, where severe occlusions, large viewpoint variations, and sparse or absent pose annotations in reference images are prevalent. To tackle this, the authors propose PANY, a unified model-free framework that operates with only one or a few pose-annotated reference images. PANY leverages a multi-view Transformer-based geometric backbone to learn view-consistent structures and cross-view alignment cues, and incorporates pose-graph regularized registration to fuse geometric information when auxiliary views are available. Breaking away from conventional pairwise matching paradigms, PANY supports arbitrary reference settings with either RGB or RGB-D inputs. Experiments demonstrate that PANY significantly outperforms existing model-free methods on the YCB-Video and LineMOD-Occlusion benchmarks, achieving pose accuracy improvements of 12% and over 20%, respectively, while exhibiting strong robustness in real-world scenes.