ChatPaper.aiChatPaper

UniSwap:用於說話影片的流式音視頻身份交換

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

August 13, 2026
作者: Yuxuan Zhang, Haozhong Xiong, Jiayi Song, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang, Liwei Wang
cs.AI

摘要

說話影片角色替換需要在保留來源動作、場景、語言內容及影音時序的同時,協調轉移外觀與聲音。現有方法針對兩種模態使用各自獨立優化的模型,使得影音一致性難以確保。我們提出UniSwap,這是首個在說話影片中進行串流式聯合影音身分替換的框架。給定來源影片、參考影像與參考語音片段,UniSwap在單一影音擴散Transformer中同時轉移參考外觀與嗓音音色,同時保留來源內容與動態。為了解決跨身分對齊訓練樣本稀缺的問題,我們引入交換-重建流程:從真實片段中移除視覺與語音身分,並以原始片段作為重建目標。我們從雙向骨幹網路出發,逐步調適模型,包括:情境內預訓練以實現聯合替換、條件式串流適應以實現區塊因果KV快取生成,以及高效自強制DMD以緩解曝光偏差並將每個區塊的去噪取樣步數從30步減少至3步。高效多LoRA切換使三種DMD角色得以共享單一凍結骨幹。特徵RoPE分解將快取位置保持在訓練範圍內,支援穩定的長序列推論。實驗結果展現了強勁的影音同步、具競爭力的身分保留、高效的串流能力以及穩定的長序列生成。
English
Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer while preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model through In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and Efficient Self-forcing DMD for mitigating exposure bias and reducing sampling from 30 to 3 denoising steps per block. Efficient Multi-LoRA Switching enables the three DMD roles to share a single frozen backbone. Feature-RoPE Decomposition keeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.