🤖 AI Summary
This work addresses whether embodied vision-language models (VLMs) possess reliable spatial understanding and decision-making capabilities in partially observable, safety-critical environments by proposing the Explore, Map, Remember, Decide (EMRD) evaluation framework. EMRD extends theoretical spatial cognition into safety-critical domains for the first time, assessing exploration through environmental coverage and temporal efficiency, while incorporating metrics for spatial mapping fidelity, psychological memory tests, and focal decision-making. Robustness is further evaluated under low-light conditions and texture perturbations. The study reveals that VLMs often rely on textual priors rather than spatial evidence when selecting evacuation points, exhibit significantly degraded spatial reasoning in low light yet remain robust to texture interference, and employ memory mechanisms fundamentally divergent from human cognition—introducing unpredictable alignment risks.
📝 Abstract
Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models (VLMs) under partial observability. As AI techniques are increasingly applied to safety-critical scenarios, it is crucial to understand whether VLMs possess robust spatial memory and make reliable decisions. In this paper, we assess whether VLMs' decisions are based on physical evidence or are corrupted by visual-language biases, if their memory processes align with human cognitive patterns, and how they respond to environmental hazards. We extend the ToS framework into a safety-critical, goal-driven pipeline, named Explore, Map, Remember, and Decide (EMRD). We then quantify Exploration Competence (Explore) through metrics of environmental coverage and temporal efficiency, assess Spatial Fidelity (Map), evaluate, with a suite of psychological metrics, Memory Persistence (Remember), and measure, using focal-point metrics, Cognitive Decision-Making (Decide). Our results show that in terms of decision-making capabilities, VLMs frequently select evacuation points based on pre-trained textual priors while lacking the spatial grounding to justify their choices. We also show that spatial reasoning degrades in low-light conditions, but it is not affected by texture and colour tampering. Our findings suggest that VLM memory fundamentally diverges from human cognition, creating unpredictable risks of misalignment.