ChatPaper.aiChatPaper

EditaLive! 面向直播的统一角色视频编辑

EditaLive! Unified Character Video Editing for Live Streaming

August 27, 2026
作者: Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang, Bo Li, Xiaodong Cun
cs.AI

摘要

传统视频编辑主要关注场景级内容,而直播则更强调人物主体。然而,将现有视频编辑方法直接应用于以人为中心的直播仍具挑战性,因为它们可能引入面部表情不一致问题,且通常依赖多个离线推理步骤,难以适用于实时交互。我们提出了EditaLive,一个用于实时流式角色视频编辑的新型框架。具体而言,我们从预训练图像动画模型(Wan-Animate)出发,该模型天然地将外观与运动解耦,我们通过参考帧编辑和视频重建,利用收集的CharEdit-50K数据集将其重新改造为基于指令的以人为中心的视频编辑基础模型。此外,我们将模型从离线双向生成适配为因果流式生成,并设计了一种对齐的自rollout蒸馏策略,将模型压缩为两步采样器,其中固定的RoPE和对齐强制减少了训练与推理之间的差异,首帧保留的稀疏注意力过滤了冗余的历史信息,以缓解外观漂移。大量实验表明,EditaLive实现了最先进的编辑性能,能够忠实保留面部表情,并支持低延迟的实时流式推理。
English
Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decouples appearance from motion, and repurpose it as the base model for instruction-based human-centric video editing by reference frame editing and video reconstruction via the collected CharEdit-50K dataset. Besides, we adapt the model from offline bidirectional to causal streaming generation, and design an aligned self-rollout distillation strategy that compresses the model into a two-step sampler, where fixed RoPE and align forcing reduce training--inference discrepancies, and first-frame preserved sparse attention filters redundant historical information to mitigate appearance drift. Extensive experiments demonstrate that EditaLive delivers state-of-the-art editing performance with faithful preservation of facial expressions and low-latency real-time streaming inference.