AlloEgo-VLM: Disambiguating Allocentric and Egocentric Reference Frames in Vision-Language Models

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of frame-of-reference ambiguity in spatial semantic understanding within vision-language models. We introduce AlloEgo-View, the first dataset designed for dual-frame disambiguation, alongside AlloEgo-VLM, an integrable fine-tuning framework. Through structured spatial annotation and supervised fine-tuning, this approach effectively resolves confusion between allocentric and egocentric reference frames. Validated via NVIDIA Isaac Sim simulations and real-world robotic platforms, the proposed framework significantly enhances spatial reasoning capabilities under ambiguous queries. This work bridges a critical gap in VLM spatial cognition research and demonstrates practical feasibility for open-vocabulary search tasks in embodied intelligence, establishing a robust foundation for resolving spatial ambiguities in downstream robotic applications.
📝 Abstract
This study investigates the challenge of ambiguity faced by Vision-Language Models (VLMs) in understanding spatial semantics. Spatial cognition, shaped by cognitive psychology, spatial science, and cultural context, often assigns directionality to objects. However, natural language descriptions of spatial relations frequently omit explicit reference frames, leading to semantic ambiguity and potentially serious errors for embodied AI robots. Existing VLMs, due to insufficient training on reference frames and object orientations, often produce inconsistent responses. To address this issue, we construct a new dataset, AlloEgo-View, comprising (image, query, view-specific answer) triplets that capture key object relations from both allocentric and egocentric perspectives. The view-specific descriptions follow a structured spatial representation that annotate detailed scene descriptions, reference and target objects, their orientations, reference frames, and view types. Building on AlloEgo-View, we develop AlloEgo-VLM, a framework to disambiguate allocentric and egocentric reference frames, even under ambiguous queries, and to be easily integrated into existing VLMs via supervised fine-tuning. Furthermore, we deploy our framework onto an embodied robotic platform within NVIDIA Isaac Sim to validate its real-world feasibility in open-ended object searching tasks. Experiments highlight the limitations of current VLMs in handling view-specific queries and demonstrate the strong disambiguation ability of AlloEgo-VLM.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Spatial Semantics
Reference Frames
Allocentric
Egocentric
Innovation

Methods, ideas, or system contributions that make the work stand out.

Allocentric-Egocentric Disambiguation
AlloEgo-View Dataset
Spatial Semantics
Vision-Language Models
Embodied AI
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kuan-Lin Chen
Institute of Computer Science and Engineering & College of Artificial Intelligence, National Yang Ming Chiao Tung University, Hsinchu, Taiwan
T
Tzu-Ti Wei
Institute of Computer Science and Engineering & College of Artificial Intelligence, National Yang Ming Chiao Tung University, Hsinchu, Taiwan
C
Chao-Chi Liao
Institute of Computer Science and Engineering & College of Artificial Intelligence, National Yang Ming Chiao Tung University, Hsinchu, Taiwan
Yu-Chee Tseng
Yu-Chee Tseng
College of AI, National Yang Ming Chiao Tung University
mobile computingwireless networkartificial intelligence
J
Jen-Jee Chen
Institute of Computer Science and Engineering & College of Artificial Intelligence, National Yang Ming Chiao Tung University, Hsinchu, Taiwan