Bridging Learned Visual Perception and Symbolic Belief-Space Planning

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过提出VLM作为概率基础的新范式,解决了在部分可观测环境下由于不确定性导致的符号计划不稳健问题。
📝 Abstract
In partially observable settings, agents must act without full knowledge of the world state and rely on uncertain state-estimation pipelines. Obtaining grounded and verifiable symbolic plans under such uncertainty remains a key challenge. Recent work has integrated Vision-Language Models (VLMs) to bridge perception and symbolic reasoning, following two main paradigms. The first, VLM-as-planner, maps images directly to action sequences, and the second, VLM-as-grounder, grounds observations into symbolic predicates used as the initial state by off-the-shelf planners. Both approaches ignore uncertainty in the planning process, compromising robustness. We introduce a third paradigm, VLM-as-probabilistic-grounder, a novel approach that captures the uncertainty of VLM predicate groundings as a probability distribution over symbolic states. This enables planning in belief space and producing robust plans under uncertainty. Experiments in simulated household robot settings show improved robustness and task success over deterministic grounding, underscoring how our approach leverages foundation models for reliable planning under uncertainty.
Problem

Research questions and friction points this paper is trying to address.

partially observable settings
uncertainty
symbolic plans
Innovation

Methods, ideas, or system contributions that make the work stand out.

VLM-as-probabilistic-grounder
belief space planning
uncertainty
robust plans
symbolic predicates
🔎 Similar Papers
G
Guy Azran
Taub Faculty of Computer Science, Technion - Israel Institute of Technology
M
Michael Navat
Faculty of Mathematics, Technion - Israel Institute of Technology
Sarah Keren
Sarah Keren
​The Taub Faculty of Computer Science Technion - Israel Institute of Technology
Artificial IntelligenceGoal recognitionAutomated PlanningReasoning under uncertainty