PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出PonderPounce方法,利用预训练的多模态大语言模型作为机器人控制的情境记忆引擎,通过双系统协作实现高效任务执行。
📝 Abstract
Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and infer behavior from a few examples. Yet vision-language-action (VLA) models generally inherit pretrained representations without using this contextual capacity as episode memory. Memory-dependent policies address this gap through purpose-built history mechanisms. PonderPounce instead reuses an MLLM's native causal context as robot memory. Ponder, a System2 MLLM, accumulates episode observations, demonstrations, and prior cognition in its native causal context and can generate subgoal text and demonstration reasoning for internal use. Pounce, a System1 VLA, receives the current observation, instruction, and proprioception directly; through the Ponder--Pounce interface, it asynchronously receives only the newest continuous cognition token and its age. Both are jointly trained end to end without a purpose-built memory module or separate bridge pretraining. Optimized serving achieves p50 latencies of 78ms for cognition refresh and 25ms for action-model invocation, supporting 20Hz action playback. On RoboMME with base-scale training data, PonderPounce reaches 60.83% with 9B and 50.04% with 0.8B under the same Pounce architecture and interface, versus 44.51% for FrameSamp+Modul and 17.93% for the current-observation π_{0.5}. With 9x data, it reaches 75.54% versus 57.88% for FrameSamp+Modul. On RoboCasa-DC, the same interface learns from action supervision alone and reaches 12.5% versus 11.6% for the strongest published demonstration-conditioned baseline, falling to 8.6% when cognition is replaced by a learned null state.
Problem

Research questions and friction points this paper is trying to address.

Multimodal large language models
Episode memory
Robot control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Large Language Models
Episode Context Engine
Causal Context
End-to-End Training
Action Playback
🔎 Similar Papers
No similar papers found.
S
Suhwan Choi
MAUM.AI, Seoul National University
Jaeyoon Jung
Jaeyoon Jung
MAUM AI Inc, Soongsil University, Republic of Korea
MultimodalEmbodied AI
S
Sungkyung Kim
MAUM.AI, Seoul National University
Yunsung Lee
Yunsung Lee
Head of Research @ WoRV, maum.ai
Y
Youngjae Yu
Seoul National University