Towards Generalizable Visually Grounded Exploration of Household Devices

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对自主探索设备操作能力不足的问题,提出VGEBench基准,通过逻辑驱动状态机框架促进视觉语言模型的通用视觉基础探索。
📝 Abstract
Recent advancements in Vision-Language Models (VLMs) have demonstrated impressive capabilities in static visual recognition and high-level semantic reasoning. However, current embodied exploration paradigms still heavily rely on imitation learning from human-annotated trajectories, which severely limits agents' generalization ability. The key bottleneck of realizing general autonomous embodied agents lies in Generalizable Visually Grounded Exploration: the ability to operate novel devices without manuals or specific training by actively grounding abstract world knowledge into fine-grained visual affordances. Yet, existing benchmarks fail to evaluate this capability: they generally rely on explicit documents and annotated trajectories, neglecting the dynamic Hypothesis-Interaction-Refinement process essential for functional device operation. To bridge this gap, we introduce VGEBench, a comprehensive benchmark designed to evaluate the generalizable visually grounded exploration capabilities of VLMs. Unlike static datasets, we construct a Logic-Driven State Machine framework. This framework simulates multi-turn interaction loops, compelling agents to achieve goals by active visual perception and feedback-driven correction. Experimental results demonstrate that existing VLMs face significant challenges in translating semantic knowledge into physical execution and maintaining long-horizon state tracking.
Problem

Research questions and friction points this paper is trying to address.

Generalizable Visually Grounded Exploration
Vision-Language Models
embodied exploration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generalizable Visually Grounded Exploration
Logic-Driven State Machine
Active Visual Perception
Feedback-Driven Correction
Multi-turn Interaction Loops
🔎 Similar Papers
No similar papers found.
L
Linhao Zheng
School of Computer Science and Technology, Beijing Institute of Technology
Z
Zeming Liu
School of Computer Science and Engineering, Beihang University
W
Wangke Chen
School of Computer Science and Technology, Beijing Institute of Technology
Li Zeng
Li Zeng
Peking University
LLM training and inferenceVector ComputingGraph Computing
Wanxiang Che
Wanxiang Che
Professor of Harbin Institute of Technology
Natural Language Processing
H
Heyan Huang
School of Computer Science and Technology, Beijing Institute of Technology
Yuhang Guo
Yuhang Guo
Beijing Institute of Technology
Natural Language Processing