Finder: Agentic Closed-Loop Object Finding for Embodied Grounding

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Finder通过闭环状态链接查询规划、证据收集等步骤,解决了在部分观察的3D场景中根据语言指示找到物体的问题,提高了物体检索成功率。
📝 Abstract
Finding the object referred to by language in a partially observed 3D scene is a core capability for embodied agents. Existing approaches either couple object search with online exploration, which can be costly when relevant observations have already been captured, or query pre-built open-vocabulary maps and scene graphs in a static, one-shot fashion. We present Finder, an agentic closed-loop object-finding primitive for embodied grounding. Instead of treating grounding as passive retrieval from a fixed scene representation, Finder maintains a typed loop state that links query-conditioned planning, scoped evidence gathering, candidate verification, and accept/continue/abort control. When evidence is incomplete or ambiguous, the loop can redirect subsequent perception and comparison rather than simply returning the top retrieved object. On open-vocabulary embodied Object Retrieval in Habitat/HM3D and real-world RGB-D scenes, Finder improves the averaged 1m success rate by 15.75 points over strong baselines. The same primitive also transfers to sequential object grounding and embodied object-centric question answering, improving spatial and temporal localization without changing the inner grounding protocol. Project page: https://finder-vln.github.io.
Problem

Research questions and friction points this paper is trying to address.

Embodied Agents
Object Finding
3D Scene
Language Grounding
Closed-Loop
Innovation

Methods, ideas, or system contributions that make the work stand out.

agentic closed-loop
typed loop state
query-conditioned planning
scoped evidence gathering
candidate verification
S
Shixiong Xu
Xiaomi Robotics
Z
Zhiyuan Chen
Xiaomi Robotics
S
Song Ding
Xiaomi Robotics
R
Rui Luo
Xiaomi Robotics
X
Xiaowei Liang
Xiaomi Robotics
D
Dongxu Miao
Xiaomi Robotics
Z
Zhiying Du
Xiaomi Robotics