UniSwap: 발화 영상을 위한 스트리밍 오디오-비주얼 신원 교체
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos
August 13, 2026
저자: Yuxuan Zhang, Haozhong Xiong, Jiayi Song, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang, Liwei Wang
cs.AI
초록
토킹 비디오 캐릭터 교체는 소스 동작, 장면, 언어적 내용 및 시청각 타이밍을 보존하면서 외형과 음성의 조율된 전송을 요구한다. 기존 방법들은 두 양식에 대해 개별적으로 최적화된 모델을 사용하므로 시청각 일관성을 확보하기 어렵다. 본 논문은 토킹 비디오에서 스트리밍 방식의 통합 시청각 정체성 교체를 위한 최초의 프레임워크인 UniSwap을 제시한다. 소스 비디오, 참조 이미지, 참조 음성 클립이 주어지면 UniSwap은 단일 시청각 확산 트랜스포머 내에서 참조 외형과 음성 음색을 전송하면서 소스의 내용과 역동성을 보존한다. 정렬된 교차 정체성 학습 쌍의 부족 문제를 해결하기 위해, 실제 클립에서 시각적 및 음성 정체성을 제거하고 원본 클립을 재구성 대상으로 사용하는 스왑-앤-재구성 파이프라인을 도입한다. 양방향 백본에서 출발하여, 통합 교체를 위한 인컨텍스트 사전학습, 블록 인과적 KV 캐시 생성을 위한 조건부 스트리밍 적응, 노출 편향 완화 및 블록당 디노이징 단계를 30회에서 3회로 줄이는 효율적 셀프포싱 DMD를 통해 모델을 점진적으로 적응시킨다. 효율적 멀티-LoRA 스위칭은 세 가지 DMD 역할이 단일 동결 백본을 공유할 수 있게 한다. Feature-RoPE 분해는 캐시된 위치를 학습 범위 내에 유지하여 안정적인 장시간 추론을 지원한다. 실험 결과는 강력한 시청각 동기화, 경쟁력 있는 정체성 보존, 효율적 스트리밍 및 안정적인 장시간 생성을 입증한다.
English
Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer while preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model through In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and Efficient Self-forcing DMD for mitigating exposure bias and reducing sampling from 30 to 3 denoising steps per block. Efficient Multi-LoRA Switching enables the three DMD roles to share a single frozen backbone. Feature-RoPE Decomposition keeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.