ChatPaper.aiChatPaper

EditaLive! 用於直播的統一角色影片編輯

EditaLive! Unified Character Video Editing for Live Streaming

August 27, 2026
作者: Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang, Bo Li, Xiaodong Cun
cs.AI

摘要

傳統影片編輯主要專注於場景層級的內容,而直播則更強調人物主體。然而,將現有影片編輯方法直接應用於以人為中心的直播仍然具有挑戰性,因為這些方法可能引入面部表情不一致的問題,且通常依賴於多個離線推論步驟,因此不適合即時互動。我們提出 EditaLive,一個用於即時串流角色影片編輯的新型框架。具體而言,我們從一個預訓練的圖像動畫模型(Wan-Animate)出發,該模型自然地將外觀與動作解耦,並透過參考幀編輯以及利用我們收集的 CharEdit-50K 資料集進行影片重建,將其重新定位為基於指令的以人為中心的影片編輯基礎模型。此外,我們將模型從離線雙向生成改適為因果串流生成,並設計了一種對齊式自展開蒸餾策略,將模型壓縮為兩步取樣器。其中,固定的 RoPE 與對齊強制(align forcing)可減少訓練與推論之間的差異,而保留首幀的稀疏注意力則可過濾多餘的歷史資訊,以減輕外觀漂移。大量實驗表明,EditaLive 提供了最先進的編輯效能,能忠實保留面部表情,並實現低延遲的即時串流推論。
English
Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decouples appearance from motion, and repurpose it as the base model for instruction-based human-centric video editing by reference frame editing and video reconstruction via the collected CharEdit-50K dataset. Besides, we adapt the model from offline bidirectional to causal streaming generation, and design an aligned self-rollout distillation strategy that compresses the model into a two-step sampler, where fixed RoPE and align forcing reduce training--inference discrepancies, and first-frame preserved sparse attention filters redundant historical information to mitigate appearance drift. Extensive experiments demonstrate that EditaLive delivers state-of-the-art editing performance with faithful preservation of facial expressions and low-latency real-time streaming inference.