Retrieving and Refining Winning Noise Tickets for Diffusion-Based Motion Generation
Existing diffusion-based text-to-motion generation models often suffer from semantic drift in long-horizon and compositional motions, struggling to balance semantic fidelity with motion coherence. This work proposes WINRO, a novel framework that reveals—for the first time—the decisive role of initial noise in determining the semantics of generated motions. WINRO introduces a training-free, model-agnostic mechanism for noise retrieval and refinement: by retrieving the most text-aligned “winning noise ticket” and applying KL-regularized optimization—optionally enhanced with a single-step forward LoRA adapter—it substantially improves text-motion alignment. Experiments demonstrate that WINRO effectively enhances semantic consistency for both MDM and MotionLCM on HumanML3D, boosts temporal robustness on the MTT benchmark, and generalizes successfully to motion stylization and spatially constrained motion synthesis tasks.