GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding

πŸ“… 2024-11-29
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 1
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Open-vocabulary 3D object affordance localization aims to precisely localize functional regions on 3D objects that enable actions specified by arbitrary natural language instructions. Existing methods suffer from limited semantic priors and struggle to model geometric invariance and implicit interaction intent. To address this, we propose a geometry-intent collaborative reasoning paradigm: (i) an implicit invariant geometric representation that disentangles intrinsic structural properties; (ii) an interaction intent analogy mechanism enabling human-like functional reasoning across objects and instructions; and (iii) PIADv2β€”the largest open-vocabulary 3D affordance understanding dataset to date. Our method integrates point-cloud–image joint representation learning, geometry-guided intent distillation, and contrastive analogy reasoning. It achieves significant improvements over state-of-the-art methods on open-vocabulary affordance localization, demonstrating strong generalization to unseen objects and instructions. Code and dataset are publicly released.

Technology Category

Application Category

πŸ“ Abstract
Open-Vocabulary 3D object affordance grounding aims to anticipate ``action possibilities'' regions on 3D objects with arbitrary instructions, which is crucial for robots to generically perceive real scenarios and respond to operational changes. Existing methods focus on combining images or languages that depict interactions with 3D geometries to introduce external interaction priors. However, they are still vulnerable to a limited semantic space by failing to leverage implied invariant geometries and potential interaction intentions. Normally, humans address complex tasks through multi-step reasoning and respond to diverse situations by leveraging associative and analogical thinking. In light of this, we propose GREAT (GeometRy-intEntion collAboraTive inference) for Open-Vocabulary 3D Object Affordance Grounding, a novel framework that mines the object invariant geometry attributes and performs analogically reason in potential interaction scenarios to form affordance knowledge, fully combining the knowledge with both geometries and visual contents to ground 3D object affordance. Besides, we introduce the Point Image Affordance Dataset v2 (PIADv2), the largest 3D object affordance dataset at present to support the task. Extensive experiments demonstrate the effectiveness and superiority of GREAT. The code and dataset are available at https://yawen-shao.github.io/GREAT/.
Problem

Research questions and friction points this paper is trying to address.

Open-Vocabulary 3D object affordance grounding with arbitrary instructions
Leveraging invariant geometries and interaction intentions for affordance knowledge
Addressing limited semantic space in existing methods through collaborative inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mines object invariant geometry attributes
Performs analogical reasoning in interactions
Combines geometry and visual content knowledge
πŸ’Ό Related Jobs
No related jobs found.
University of Science and Technology of China | Northeastern University
Y
Yawen Shao
University of Science and Technology of China
W
Wei Zhai
University of Science and Technology of China
Y
Yuhang Yang
University of Science and Technology of China
H
Hongcheng Luo
Northeastern University
Y
Yang Cao
University of Science and Technology of China, Institute of Artificial Intelligence, Hefei Comprehensive National Science Center
Z
Zhengjun Zha
University of Science and Technology of China