MoVT: Video-Augmented Motion Tokenizer for Text-to-Motion Generation

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决文本驱动3D人体动作生成中因训练数据有限导致的问题,提出MoVT框架,利用视频增强的动作tokenizer提高生成质量。
📝 Abstract
Text-driven 3D human motion generation models face significant challenges in responding to diverse and unconstrained textual prompts, primarily due to the limited availability of 3D motion training data. To address this, we introduce MoVT, a novel framework that effectively leverages the extensive range of human action videos to enhance text-to-motion generation. At the core of our approach is the cross-modal augmented motion tokenizer, which projects discrete 3D motion tokens into the 2D domain. This projection allows us to enrich the motion codebook with complex, real-world motion patterns derived from videos. The enriched discrete tokens are then mapped back to the 3D domain, resulting in aligned 3D and 2D codebooks with an enhanced capacity to represent intricate motions. These enhanced codebooks are integrated into a generative masked transformer, which predicts masked motion token indices in a modality-agnostic manner. This enables the use of text-index pairs, generated from the 2D codebook and annotated motion videos, to further enhance the generator. Extensive empirical evaluations show that MoVT performs favorably against prior state-of-the-art methods across multiple key metrics.
Problem

Research questions and friction points this paper is trying to address.

Text-driven 3D human motion
unconstrained textual prompts
limited 3D motion training data
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-modal augmented motion tokenizer
discrete 3D motion tokens
text-to-motion generation
generative masked transformer
B
Beibei Jing
Huazhong University of Science and Technology, School of Computer Science and Technology, Wuhan, Hubei, China
T
Tianle Guo
Huazhong University of Science and Technology, School of Computer Science and Technology, Wuhan, Hubei, China
Youjia Zhang
Youjia Zhang
Huazhong University of Science and Technology
Deep LearningComputer GraphicsComputer Vision
Zikai Song
Zikai Song
Huazhong University of Science and Technology
deep learningmultimedia modelsvisual trackingsocial media analysis
Y
Yawei Luo
Zhejiang University, School of Software Technology, Ningbo, China
Junqing Yu
Junqing Yu
Huazhong University of Science & Technology
Tao Guan
Tao Guan
Huazhong University Of Science And Technology
W
Wei Yang
Huazhong University of Science and Technology, School of Computer Science and Technology, Wuhan, Hubei, China