UniMPA: A Unified Memory-Prediction-Action Model via Action-Grounded Transition Modeling

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出UniMPA模型,通过共享动作基础的转换接口解决视觉-语言-动作模型中的过渡模糊、预测-执行不匹配和经验-实现不匹配问题。
📝 Abstract
Recent advances in Vision-Language-Action (VLA) models have improved robotic manipulation, yet observation-to-action learning remains limited by a fundamental transition realizability gap, manifested in three tightly coupled problems: (i) Transition ambiguity. Visually similar current observations may correspond to different manipulation phases and imply different subsequent transitions. (ii) Prediction--execution mismatch. A visually plausible predicted future observation does not necessarily correspond to a physically realizable transition. (iii) Experience--realization mismatch. A historically executable action pattern may not necessarily realize the intended transition in the current scene and therefore requires context-aware adaptation. Accordingly, we propose UniMPA, a Unified Memory-Prediction-Action model that addresses these problems through a shared action-grounded transition interface. (i) UniMPA introduces Persistent-Selective Future Prediction to resolve transition ambiguity by modeling the intended future state evolution. A persistent latent stream continuously tracks task-level progress, while a transition-critical pixel stream selectively resolves fine-grained interaction changes through memory-grounded prediction. (ii) To assess the physical executability of the anticipated transition, the predicted transition queries a temporal Visual-Action Memory Bank. The bank retrieves historically realized visual-action experience, grounding future prediction in executable evidence. (iii) To adapt executable experience to the current scene, an Action-Visual Memory Bank retrieves visually grounded action prototypes from historical action evolution. Prototype-Biased Flow then shifts the flow source toward a historically supported action manifold for context-aware refinement.
Problem

Research questions and friction points this paper is trying to address.

Transition Ambiguity
Prediction-Execution Mismatch
Experience-Realization Mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Memory-Prediction-Action Model
Persistent-Selective Future Prediction
Visual-Action Memory Bank
Action-Visual Memory Bank
Prototype-Biased Flow
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
W
Wei Li
School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), China, 518055
Rui Shao
Rui Shao
Professor, Harbin Institute of Technology (Shenzhen)
Computer VisionMultimodal LLMEmbodied AI
J
Jie He
School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), China, 518055
L
Lingsen Zhang
School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), China, 518055
Ziwei Liu
Ziwei Liu
Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics
L
Liqiang Nie
School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), China, 518055