ChatPaper.aiChatPaper

UniSwap:面向说话视频的流式音视频身份替换

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

August 13, 2026
作者: Yuxuan Zhang, Haozhong Xiong, Jiayi Song, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang, Liwei Wang
cs.AI

摘要

说话视频人物替换要求在协同迁移外观与声音的同时,保留源动作、场景、语言内容及音视频时序。现有方法对两种模态分别使用独立优化的模型,导致音视频一致性难以保证。我们提出UniSwap——首个在说话视频中实现流式联合音视频身份替换的框架。给定源视频、参考图像和参考语音片段,UniSwap在统一的音视频扩散Transformer中完成参考外观与声音音色的迁移,同时保留源内容与动态。为解决跨身份对齐训练对稀缺的问题,我们引入交换-重建流程,从真实片段中移除视觉与语音身份,并以原始片段作为重建目标。从双向骨干网络出发,我们通过上下文预训练(用于联合替换)、条件流式适配(用于块级因果KV缓存生成)以及高效自强制DMD(用于缓解暴露偏差并将每块去噪步数从30减少至3)逐步适配模型。高效多LoRA切换使三种DMD角色共享同一冻结骨干网络。特征RoPE分解将缓存位置保持在训练范围内,支持稳定的长序列推理。实验结果表明,该方法实现了强音视频同步、具有竞争力的身份保持、高效的流式生成以及稳定的长序列生成。
English
Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer while preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model through In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and Efficient Self-forcing DMD for mitigating exposure bias and reducing sampling from 30 to 3 denoising steps per block. Efficient Multi-LoRA Switching enables the three DMD roles to share a single frozen backbone. Feature-RoPE Decomposition keeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.