STAGE: Diagnosing Semantic Transfer at Grounded Execution in Embodied Agents

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了语义与动作之间的转换问题,通过引入VISA接口将恢复的语义转换为具体决策,提高动作执行的一致性和有效性。
📝 Abstract
Embodied language grounding requires more than identifying the referent of an instruction: recovered semantics must also control the action an agent exposes. We study this missing link as a semantic-action gap, where instruction semantics are recoverable but weakly expressed in native continuous actions. We introduce SAT-Bench, a fixed-observation counterfactual benchmark that holds the visual scene and agent state fixed while changing only instruction semantics. On LIBERO target-name and pixel-grounded relation swaps, target recovery reaches 100.0% and 95.8%, whereas OpenVLA action sensitivity remains only 6.8% and 7.7%. The gap persists across 1,000 additional compositional and temporal/procedural counterfactuals, with overall action sensitivity of 6.1%. Hidden-state, threshold-free, cross-policy, and rollout diagnostics further support this semantic-action transfer failure. We introduce VISA, a lightweight execution-time interface that converts recovered semantics into ALLOW, DEFER, target-consistency, and verified-selection decisions. VISA reduces invalid-instruction blind execution from 92.7% to 2.8% while preserving 94.0% of normal commands, and verified selection further improves target-consistent action exposure without updating the underlying policy. Overall, embodied language evaluation should measure semantic-action transfer, not semantic parsing alone.
Problem

Research questions and friction points this paper is trying to address.

embodied agents
semantic-action gap
instruction semantics
Innovation

Methods, ideas, or system contributions that make the work stand out.

VISA
semantic-action transfer
ALLOW and DEFER decisions
target-consistency
verified selection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Baosheng Jin
Center for Data Science, NYU Shanghai
Y
Yushen Liang
Center for Data Science, NYU Shanghai
Hua Shen
Hua Shen
Assistant Professor, NYU Shanghai / New York University
bidirectional human-AI alignmenthuman-AI interactionAI/LLM interpretability and evaluation