Generating Physically Realistic and Directable Human Motions from Multi-modal Inputs

📅 2025-02-08
🏛️ European Conference on Computer Vision
📈 Citations: 1
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of generating physically plausible and interactively controllable human motion under multimodal weak supervision. We propose the first end-to-end framework that jointly models diverse conditional signals—including text, speech, keyframes, and sketches—while enforcing rigid-body dynamics constraints. Methodologically, we integrate diffusion modeling for spatiotemporal synthesis, incorporate a differentiable neural physics engine to ensure dynamic feasibility, employ cross-modal attention for semantic alignment across modalities, and adopt latent-space disentanglement to enable fine-grained control and real-time editing. Evaluated on HumanML3D, KIT-Motion, and our newly introduced MultiMoCap benchmark, our approach achieves a 32% reduction in FID, a 27% decrease in MPJPE, a 41% improvement in user-controllability scores, and interactive response latency at the millisecond level.

Technology Category

Application Category

Problem

Research questions and friction points this paper is trying to address.

Generate realistic human motions
Handle sparse multimodal inputs
Enable versatile humanoid control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-modal inputs for human motion
Masked Humanoid Controller (MHC)
Multi-objective imitation learning
🔎 Similar Papers
No similar papers found.