Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion
Text-to-3D human motion generation faces two key challenges: poor generalization—especially to out-of-distribution motions—and coarse-grained control, hindering frame-level precision. To address these, we propose MoMADiff, the first framework integrating masked autoregressive diffusion into continuous-frame motion representation. It enables spatio-temporal fine-grained controllability by allowing users to specify sparse keyframes. Furthermore, we enhance latent-space modeling via a VQVAE-augmented architecture to improve joint discrete-continuous representation learning. Evaluated on two newly constructed sparse keyframe datasets and two standard benchmarks, MoMADiff achieves state-of-the-art performance in motion quality, text instruction fidelity, and keyframe alignment accuracy.