Generating Physically Realistic and Directable Human Motions from Multi-modal Inputs
This work addresses the challenge of generating physically plausible and interactively controllable human motion under multimodal weak supervision. We propose the first end-to-end framework that jointly models diverse conditional signals—including text, speech, keyframes, and sketches—while enforcing rigid-body dynamics constraints. Methodologically, we integrate diffusion modeling for spatiotemporal synthesis, incorporate a differentiable neural physics engine to ensure dynamic feasibility, employ cross-modal attention for semantic alignment across modalities, and adopt latent-space disentanglement to enable fine-grained control and real-time editing. Evaluated on HumanML3D, KIT-Motion, and our newly introduced MultiMoCap benchmark, our approach achieves a 32% reduction in FID, a 27% decrease in MPJPE, a 41% improvement in user-controllability scores, and interactive response latency at the millisecond level.