Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出TOWN-VLA方法解决视觉-语言-动作操控中因提示形式改变导致性能下降的问题,通过控制干预授权机制提高任务成功率。
📝 Abstract
Retrieval can efficiently and effectively augment a frozen vision--language--action (VLA) policy without retraining, yet retrieved text becomes a control intervention once it enters the executed prompt. In a matched audit, raw appended text reduces mean success from 92.47\% to 3.00\%, while meaningful and length-matched meaningless appends both fail on all 500 states. This result identifies \emph{prompt-form collapse}: changing the instruction form, rather than adding useful semantics, can dominate execution. We introduce TOWN-VLA (Think Only When Needed), a prompt-authority interface that separates candidate generation from permission to alter the policy input. A fixed compatibility rule authorizes a canonical compact instruction; otherwise, the interface restores the original Base prompt exactly. Across 900 audited routes, every route follows this contract: 525 routes recover Base with matching hashes, and all 375 authorized prompts preserve the task signature. On a matched $4\times7$ LIBERO-Plus evaluation with 10{,}030 episodes per method, success rises from 69.5\% to 73.1\% ($+362$ episodes; 95\% CI 1.89--5.45 points), improving on six perturbation axes and all four suites. On a physical PiPER arm with a frozen \pizerofive{} checkpoint, success rises from 52.7\% to 78.7\% over 150 trials per method ($p=3.16\times10^{-6}$). Prompt authority is enforceable for a frozen controller; oracle-free admission calibration is the next deployment target.
Problem

Research questions and friction points this paper is trying to address.

prompt-form collapse
vision-language-action
policy intervention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prompt-Authority Control
Selective Slow-Path Intervention
Vision-Language-Action Manipulation
TOWN-VLA
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Z
Zhiruo Zhou
SIGS, Tsinghua University, China; Wuhan University of Technology, China
Z
Zelin Li
SIGS, Tsinghua University, China; Shanghai Jiao Tong University, China
Xiwen Chen
Xiwen Chen
Clemson University
Deep LearningMultimodalComputer VisionTime Series AnalysisVLM/LLM
J
Jiazhuo Li
University of Michigan - Ann Arbor, USA
C
Chenwei Wang
AiDlab, Hong Kong (China SAR)
H
Huiming Chen
City University of Hong Kong, Hong Kong (China SAR)
Xiaojun Zhu
Xiaojun Zhu
SIGS, Tsinghua University, China