Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
Addressing the challenges of identity preservation, reliance on fine-tuning, and data scarcity in text-to-video generation, this paper proposes a training-free triple-enhancement framework. First, GPT-4o–driven face-aware prompt enhancement bridges the semantic gap between textual descriptions and visual content. Second, a prompt-aware reference image optimization mechanism improves input consistency. Third, a unified gradient-guided strategy jointly optimizes identity fidelity and spatiotemporal coherence during diffusion model sampling—enabling inference-time refinement without architectural modification. The method requires no model training or fine-tuning. Extensive evaluation on a thousand-video benchmark demonstrates significant improvements in character identity consistency and video quality, outperforming state-of-the-art approaches in both automated metrics and human assessment. It ranked first in the ACM Multimedia 2025 Challenge, validating its strong generalizability and practical applicability.