SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决动作语言建模中语义和运动细节编码问题,提出SeMoCo方法,通过语义优先的运动编解码器及双轴生成器优化了文本到动作的生成。
📝 Abstract
Discrete motion representations have substantially advanced autoregressive text-to-motion generation. However, most motion tokenizers are optimized for reconstruction and do not explicitly allocate capacity according to semantic role. Action-level meaning and fine-grained kinematic detail must therefore be encoded through the same reconstruction-driven hierarchy. We introduce SeMoCo, a semantic-first motion codec, together with a dual-axis motion generator for language-conditioned motion generation. Each motion token contains one semantic token and a residual sequence of kinematic tokens. The generator models semantic progression across time and autoregressively refines the residual entries. We also construct $Ω$-MotionVerse, a large-scale, multi-source human-motion dataset unified under the SOMA representation. Across the reported comparisons, SeMoCo achieves the best reconstruction accuracy among the compared codecs, while strong text-to-motion results demonstrate the effectiveness of its motion tokens for downstream generation.
Problem

Research questions and friction points this paper is trying to address.

discrete motion representations
semantic role
reconstruction
Innovation

Methods, ideas, or system contributions that make the work stand out.

semantic-first motion codec
dual-axis motion generator
language-conditioned motion generation
Ω-MotionVerse
T
Tianlv Huang
Jilin University
H
Hetian Guo
Jilin University
Z
Ziyi Cai
Harbin Institute of Technology, Shenzhen
S
Song Wang
Frontier Robotics
Y
Yanping Zhang
Frontier Robotics
Z
Zipei Fan
Jilin University
Xuan Song
Xuan Song
Dean and Professor, School of Artificial Intelligence, Jilin University; SUSTech
Artificial IntelligenceData MiningUrban Computing
Guangming Wu
Guangming Wu
Frontier Robotics
X
Xin Zheng
Frontier Robotics