TBDub: Production-Oriented Visual Dubbing

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出TBDub,通过任务自适应后训练和少步蒸馏解决视觉配音中的同步、身份保持等问题,提高了效率和质量。
📝 Abstract
Visual dubbing must synchronize mouth motion with replacement speech while preserving identity, appearance, and temporal consistency. Although X-Dub provides a strong mask-free video-editing baseline, its application to livestream and generated-video content reveals limitations in production-domain robustness, temporal and motion stability, identity and oral-detail preservation, and inference efficiency. We present \textbf{TBDub}, a production-oriented extension of X-Dub that combines task-adaptive post-training with task-aware few-step distillation. Post-training adapts the video DiT using production-domain data, production-specific conditioning and filtering, and enhanced audio features to obtain a 30-step Teacher. Distillation adapts DMD/DMD2 to conditional video editing and compresses the Teacher into a two-step Student. On 38 TalkVid clips, the Teacher improves all eight reported reconstruction, perceptual, identity, and synchronization metrics over X-Dub. In the MOS evaluation, it improves lip-sync consistency, identity consistency, and visual quality over X-Dub by 0.14, 0.95, and 0.90 points, while the Student achieves the highest lip-sync and visual-quality scores and remains close to the Teacher in identity consistency. In paired end-to-end generation timing from the first VAE encode through the final VAE decode on a single NVIDIA H20 GPU at $512\times512$, the Student reaches 7.13 effective FPS and reduces total latency by $13.93\times$; the DiT stage alone is accelerated by $42.49\times$. The Student largely retains the Teacher's generation quality and audiovisual synchronization. The code is available on GitHub at \https://github.com/TaoLiveAIGC/TBDub, and the 30-step Teacher and two-step Student weights are available on Hugging Face at https://huggingface.co/TaoLiveAIGC/TBDub.
Problem

Research questions and friction points this paper is trying to address.

Visual Dubbing
Temporal Consistency
Identity Preservation
Inference Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

task-adaptive post-training
few-step distillation
production-domain robustness
temporal and motion stability
inference efficiency
🔎 Similar Papers
No similar papers found.
B
Bihan Li
TaoLive AIGC, Taobao & Tmall Group of Alibaba
X
Xinyang Li
TaoLive AIGC, Taobao & Tmall Group of Alibaba
Z
Zeran Xu
TaoLive AIGC, Taobao & Tmall Group of Alibaba
M
Meiguang Jin
TaoLive AIGC, Taobao & Tmall Group of Alibaba
Junfeng Ma
Junfeng Ma
Mississippi State University
Design and ManufacturingLogisticsAI/MLHuman-Technology InteractionSustainability