UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing instruction-driven methods for 3D human motion editing suffer from limitations in control precision and data diversity, struggling to simultaneously achieve spatiotemporal accuracy, semantic richness, and fidelity to the source motion. To address these challenges, this work proposes Omni-MoEdit, a unified framework for motion generation and editing, featuring three key contributions: a large-scale, closed-loop synthesized Omni-MoEdit dataset; the UniMoFlow model based on unified latent-space flow matching; and the Source-Anchored Flow Editing (SAFE) inference strategy, accompanied by semantic-aware evaluation metrics. Experiments demonstrate that the proposed approach significantly improves text-motion alignment, editing effectiveness, and cycle consistency, while preserving high-quality source motion fidelity and maintaining strong performance in text-to-motion generation.
📝 Abstract
Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods either resort to training-free adaptation of generative models or rely solely on triplet supervision; however, adaptation often yields suboptimal control, and manually curated triplet datasets remain severely limited in scale and semantic diversity. To overcome this bottleneck, we ground motion editing directly within text-to-motion generation across data, architecture, and inference. At the data level, we develop a closed-loop synthesis-and-verification pipeline that produces Omni-MoEdit, a large-scale dataset spanning body-part, amplitude, temporal, action, and style edits. At the architectural level, we introduce UniMoFlow, a unified latent flow-matching model that shares broad semantic and kinematic knowledge between generation and editing. At the inference level, SAFE (Source-Anchored Flow Editing) complements UniMoFlow with controllable, source-anchored refinement. Furthermore, we augment standard evaluations with semantics-aware metrics to account for valid edits that inherently deviate from a single ground-truth reference. Extensive experiments demonstrate improved target-text alignment, edit effectiveness, and cycle consistency, while maintaining competitive source fidelity and text-to-motion generation quality.
Problem

Research questions and friction points this paper is trying to address.

instruction-driven editing
3D human motion
semantic grounding
spatiotemporal localization
motion generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

instruction-driven motion editing
text-to-motion generation
latent flow matching
Omni-MoEdit dataset
source-anchored refinement
🔎 Similar Papers