Spatiotemporally Decoupled Autoregressive Diffusion Model for Human Motion Generation

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决人体运动生成中的表示和架构问题,提出了一种时空解耦框架DeMoDiff,通过时空VAE及掩码注意力机制增强控制能力。
📝 Abstract
Text-driven human motion synthesis has made substantial development with two core modules of motion representation and generative architecture. For representation, Vector Quantization (VQ)-based methods compress motion data into discrete tokens while latent-based models operate directly in continuous space. However, both of these representations exhibit significant limitations. VQ-based methods suffer from inherent information loss, which compromises the quality, diversity, and generalization of generated motions, while continuous representation on holistic whole-body motion hinders part-level flexibility. For architecture, diffusion and autoregressive diffusion models have demonstrated their superiority, yet the fine-grained controllability over individual body parts is also limited. Thus, we propose a unified spatiotemporally decoupled framework named DeMoDiff, which jointly redesigns representation and architecture. To enhance representation extraction capabilities and offer greater part-level controllability, we present a spatial-temporal VAE that encodes each body joint rather than compressing the whole-body motion into a single latent space. Then, we incorporate spatial-temporal masking and attention mechanisms into an autoregressive diffusion generator, achieving both generative capability and controllable editability. Extensive experiments on the HumanML3D and KIT-ML datasets demonstrate that our model achieves state-of-the-art reconstruction performance and compelling motion generation results. Moreover, our framework demonstrates strong temporal and spatial editing capabilities, further validating its effectiveness. Our project page: https://rex0191.github.io/DeMoDiff/
Problem

Research questions and friction points this paper is trying to address.

text-driven human motion synthesis
motion representation
generative architecture
vector quantization
spatiotemporal controllability
Innovation

Methods, ideas, or system contributions that make the work stand out.

spatiotemporally decoupled
spatial-temporal VAE
autoregressive diffusion generator
controllable editability
🔎 Similar Papers
No similar papers found.
C
Chengqun Yang
MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University
Liang Xu
Liang Xu
Shanghai Jiao Tong University
Human-CentricComputer VisionEmbodied AI
Yanping Li
Yanping Li
Shanghai Jiao Tong University
F
Fulong Liu
MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University
Jingnan Gao
Jingnan Gao
Ph.D. student at Shanghai Jiao Tong University
Computer Vision
W
Weili Zeng
MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University
Y
Yichao Yan
MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University