High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer
Existing methods struggle to simultaneously preserve fine-grained facial and hand details while maintaining spatiotemporal consistency in long-duration (>3-second) human image animation. To address this, we propose a high-fidelity long-video generation framework based on the Diffusion Transformer (DiT). Our approach introduces three key innovations: (1) a novel hybrid implicit guidance signal combined with a sharpness-aware guidance factor; (2) a temporal-aware positional offset adaptation module enabling arbitrary-length video synthesis; and (3) skeleton-aligned modeling coupled with identity-agnostic data augmentation. These components collectively enhance fine-grained structural modeling and inter-frame coherence. Quantitative and qualitative evaluations demonstrate state-of-the-art performance in critical metrics—including facial expression fidelity, hand motion dynamics, and temporal smoothness—while achieving superior visual quality and strong spatiotemporal consistency.