ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对文本到动作生成中检索增强模型的粗粒度检索和融合机制及表示差距问题,提出ReMoMask-2框架,通过结构感知检索和潜空间对齐方法改善生成效果。
📝 Abstract
Text-to-motion (T2M) generation maps natural language to human joint movements, aiding gaming, VR, and robotics. Retrieval-Augmented Text-to-Motion (RAG-T2M) improves generation on complex descriptions by conditioning on retrieved motion-text pairs. However, existing RAG-T2M models face two challenges: coarse-grained retrieval and fusion mechanisms overlook the hierarchical, spatial-temporal topology of human motion, and a representation gap exists because retrieved evidence resides in a semantic space separate from the generator's latents. To address the first, we present ReMoMask, a structure-aware RAG framework coupling Hierarchical Bidirectional Momentum (HBM) contrastive learning to align global and part-level features with text; Semantic Spatial-Temporal Attention (SSTA) for topology-aware fusion; and Topology Structured Masking (TSM) to force robust part-level grounding via adaptive masking. To address the second, we introduce ReMoMask-2, which rebuilds the retrieval database directly within the generator's pre-quantization latent space and aligns text queries via a distilled lightweight projector, allowing the generator to directly consume the retrieved motion's semantic content. Extensive experiments on HumanML3D, KIT-ML, and SnapMoGen demonstrate our retriever achieves state-of-the-art accuracy, while ReMoMask-2 attains the lowest FID on KIT-ML and SnapMoGen; notably, its single mask-transformer stage surpasses ReMoMask's full two-stage pipeline and delivers the fastest inference.
Problem

Research questions and friction points this paper is trying to address.

retrieval-augmented
text-to-motion
spatial-temporal topology
representation gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Bidirectional Momentum (HBM)
Semantic Spatial-Temporal Attention (SSTA)
Topology Structured Masking (TSM)
pre-quantization latent space
lightweight projector
Y
Yiran Wang
University of Sydney, Camperdown 2006, Australia
Zeyu Zhang
Zeyu Zhang
Gaoling School of Artificial Intelligence, Renmin University of China
LLM-based AgentResponsible RecSysCausal Learning
L
Ling Shao
UCAS-Terminus AI Lab, University of Chinese Academy of Sciences, Beijing 101408, China
Hao Tang
Hao Tang
Peking University
computer vision