Skills in Weights, Memory in Code: Hybrid Learning for Memory-Dependent Robot Manipulation

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing vision-language-action (VLA) policies struggle with non-Markovian tasks due to their inability to retain and reason over critical information from long-horizon interaction histories. This work proposes HyMeS, a hybrid learning framework that decouples memory management from policy execution: it acquires low-level motor skills via gradient-driven imitation learning while employing a heuristically trained code-based agent to orchestrate high-level memory strategies, thereby guiding a Markovian VLA policy through memory-dependent manipulation tasks. By leveraging only a small set of reusable skill demonstrations, HyMeS achieves strong compositional generalization, substantially improving both data efficiency and long-horizon task performance. Evaluated on RoboMemArena, HyMeS boosts the average cumulative success rate from 52.5% to 66.2% and task success rate from 41.3% to 60.1% compared to Pi0.5, outperforming PrediMem by 4.5 and 14.5 percentage points, respectively.
📝 Abstract
Modern vision-language-action (VLA) policies have acquired broad manipulation skills, but typically generate each action chunk from the current observation or a short fixed-length history. However, real-world manipulation is often non-Markovian, requiring robots to retain and reason over task-relevant information from long-horizon interaction histories to determine the next action. To address this challenge, we propose HyMeS, a hybrid learning framework that leverages the reasoning and memory-management capabilities of coding agents to steer a Markovian VLA for memory-dependent manipulation. Specifically, HyMeS learns low-level motor skills through gradient-based imitation learning, while a coding agent acquires high-level memory-management strategies through heuristic learning by iteratively updating an executable heuristic system from rollout feedback. Furthermore, we close the loop between steering and execution through multimodal stage-completion verification, which updates memory using proprioceptive signals and multi-frame VLM judgments. Compared with end-to-end memory-augmented VLAs, HyMeS requires demonstrations only for reusable motor skills rather than for every history-dependent task configuration, enabling data-efficient compositional generalization. On RoboMemArena, HyMeS improves mean cumulative success from 52.5% to 66.2% and mean task success from 41.3% to 60.1% over pi0.5, while outperforming PrediMem by 4.5 points in cumulative success and 14.5 points in task success.
Problem

Research questions and friction points this paper is trying to address.

non-Markovian manipulation
long-horizon memory
vision-language-action policies
memory-dependent tasks
robotic manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

hybrid learning
memory-dependent manipulation
vision-language-action policy
heuristic learning
compositional generalization
💼 Related Jobs
No related jobs found.
Y
Yunhao Zhao
Northwestern University
Z
Zhenyang Ni
Northwestern University
H
Haoyang Chen
University of Minnesota
Ruohan Zhang
Ruohan Zhang
Stanford University
RoboticsCognitive ScienceBrain-Machine InterfaceMachine LearningArt
Q
Qi Zhu
Northwestern University