UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of achieving simultaneous audiovisual identity swapping in speaker videos while preserving lip-sync accuracy, semantic consistency, natural motion dynamics, and temporal alignment. To this end, the authors propose UniSwap, the first diffusion Transformer framework capable of streaming joint audiovisual identity transfer. The method employs a swap-and-reconstruct training strategy to mitigate the scarcity of cross-identity paired data and integrates several key innovations: in-context pretraining, conditional streaming adaptation, an efficient self-enforced Denoising Multi-Decoding (DMD) mechanism, Multi-LoRA switching, and Feature-RoPE decomposition. These components collectively enable high-fidelity identity transfer and stable long-form video generation. Remarkably, UniSwap produces highly synchronized audiovisual outputs with coherent speech-content-preserving facial motions in just three denoising steps.
📝 Abstract
Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer while preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model through In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and Efficient Self-forcing DMD for mitigating exposure bias and reducing sampling from 30 to 3 denoising steps per block. Efficient Multi-LoRA Switching enables the three DMD roles to share a single frozen backbone. Feature-RoPE Decomposition keeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.
Problem

Research questions and friction points this paper is trying to address.

audio-visual identity swapping
talking videos
identity replacement
audio-visual consistency
streaming generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio-visual identity swapping
streaming generation
diffusion transformer
Multi-LoRA
self-forcing DMD
🔎 Similar Papers
No similar papers found.
Yuxuan Zhang
Yuxuan Zhang
The Chinese University of Hong Kong
computer visiondiffusion models
H
Haozhong Xiong
Qwen Applications Business Group of Alibaba
J
Jiayi Song
Qwen Applications Business Group of Alibaba
Jinpeng Yu
Jinpeng Yu
Xiaohongshu, ShanghaiTech University
Computer VisionGenerative AIMultimodal3D
Y
Yang Shi
Qwen Applications Business Group of Alibaba
J
Jiaming Liu
Qwen Applications Business Group of Alibaba
R
Ruihua Huang
Qwen Applications Business Group of Alibaba
L
Liwei Wang
The Chinese University of Hong Kong