ChatPaper.aiChatPaper

UniSwap: トーキングビデオのためのストリーミング音声・映像アイデンティティスワッピング

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

August 13, 2026
著者: Yuxuan Zhang, Haozhong Xiong, Jiayi Song, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang, Liwei Wang
cs.AI

要旨

トーキングビデオのキャラクター置換には、ソースの動き、シーン、言語内容、音声と映像のタイミングを保ちながら、外観と音声の協調的な転送が必要である。既存手法は、2つのモダリティに対して別々に最適化されたモデルを使用するため、音声と映像の一貫性を強制することが困難である。我々は、トーキングビデオにおけるストリーミング型の音声-映像共同アイデンティティ置換のための初のフレームワークであるUniSwapを提案する。ソースビデオ、参照画像、参照音声クリップが与えられると、UniSwapはソースの内容とダイナミクスを保持しながら、単一の音声-映像拡散トランスフォーマー内で参照の外観と声質を転送する。対応付けられたクロスアイデンティティ学習ペアの不足に対処するため、実際のクリップから視覚的および音声的アイデンティティを除去し、元のクリップを再構成ターゲットとして使用するスワップ・アンド・リコンストラクト・パイプラインを導入する。双方向バックボーンから開始し、ジョイント置換のためのインコンテキスト事前学習、ブロック因果KVキャッシュ生成のための条件付きストリーミング適応、露出バイアスの軽減とブロックあたりのデノイジングステップの30から3への削減を実現する効率的なセルフフォーシングDMDを通じて、モデルを段階的に適応させる。効率的なマルチLoRAスイッチングにより、3つのDMDロールが単一の凍結バックボーンを共有できる。フィーチャーRoPE分解は、キャッシュされた位置をトレーニング範囲内に保ち、安定したロングフォーム推論をサポートする。実験では、強力な音声-映像同期、競争力のあるアイデンティティ保存、効率的なストリーミング、安定したロングフォーム生成が実証される。
English
Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer while preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model through In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and Efficient Self-forcing DMD for mitigating exposure bias and reducing sampling from 30 to 3 denoising steps per block. Efficient Multi-LoRA Switching enables the three DMD roles to share a single frozen backbone. Feature-RoPE Decomposition keeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.