GRAFT: Grounded and Efficient Online Reinforcement Adaptation for Fine-Grained Robot Manipulation

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出GRAFT框架,通过基于区域的监督和缓存视觉-语言前缀重用来解决在线适应细粒度生物医学任务中计算成本高和视觉线索学习难的问题。
📝 Abstract
Pretrained vision-language-action (VLA) policies provide strong priors for robot manipulation, yet adapting them online to fine-grained biomedical tasks remains challenging. Task success often hinges on subtle, view-dependent visual cues, while task-level rewards provide little guidance about which regions matter, making it difficult to learn task-relevant visual grounding from limited real-robot interaction. Online adaptation is further constrained by the computational cost of VLA inference and replay-based updates. We introduce GRAFT (Grounded Reinforcement Adaptation for Fast Task Learning), a framework for efficient online VLA adaptation through grounded perception. GRAFT uses region-level supervision to learn view-specific visual anchors that focus perception on task-relevant local cues without requiring region proposals at deployment. It further combines single-step action generation with cached visual-language prefix reuse to accelerate online learning. Across four biomedical manipulation tasks, GRAFT improves success rates by 25 percentage points under matched adaptation budgets, while reducing the computational overhead of online policy updates.
Problem

Research questions and friction points this paper is trying to address.

Online Reinforcement Adaptation
Fine-Grained Robot Manipulation
Visual-Language-Action Policies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Grounded Perception
Region-level Supervision
Visual Anchors
Single-step Action Generation
Cached Visual-Language Prefix Reuse
💼 Related Jobs
No related jobs found.
Y
Yibo Qiu
Suzhou Institute for Advanced Research, University of Science and Technology of China; School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China
H
Haoliang Ye
Suzhou Institute for Advanced Research, University of Science and Technology of China; School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China
S
Shu'ang Sun
Suzhou Institute for Advanced Research, University of Science and Technology of China; School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China
Z
Zan Huang
Suzhou Institute for Advanced Research, University of Science and Technology of China; School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China
R
Ronald X Xu
Suzhou Institute for Advanced Research, University of Science and Technology of China; School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China
Mingzhai Sun
Mingzhai Sun
University of Science and Technology of China
Biomedical Engineeringdeep learningretinal imaging