EditaLive!ライブ配信のための統合キャラクター動画編集
EditaLive! Unified Character Video Editing for Live Streaming
August 27, 2026
著者: Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang, Bo Li, Xiaodong Cun
cs.AI
要旨
従来のビデオ編集は主にシーンレベルのコンテンツに焦点を当てるのに対し、ライブ配信では人間の被写体により重点が置かれます。しかし、既存のビデオ編集手法を人間中心のライブ配信に直接適用することは依然として困難です。これらの手法は表情の不整合を引き起こす可能性があり、通常は複数のオフライン推論ステップに依存するため、リアルタイムインタラクションには不向きです。我々は、リアルタイムストリーミングキャラクタービデオ編集のための新しいフレームワークであるEditaLiveを提案します。具体的には、外観と動作を自然に分離する事前学習済み画像アニメーションモデル(Wan-Animate)を起点とし、収集したCharEdit-50Kデータセットを用いた参照フレーム編集とビデオ再構成によって、指示ベースの人間中心ビデオ編集のベースモデルとして転用します。さらに、モデルをオフライン双方向生成から因果的ストリーミング生成へ適応させ、モデルを2ステップサンプラーに圧縮するアライン型自己ロールアウト蒸留戦略を設計します。この戦略では、固定されたRoPEとアラインフォーシングが訓練と推論の乖離を低減し、先頭フレーム保持型スパースアテンションが冗長な履歴情報をフィルタリングして外観ドリフトを軽減します。広範な実験により、EditaLiveが表情の忠実な保持と低遅延リアルタイムストリーミング推論を備えた最先端の編集性能を達成することを実証します。
English
Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decouples appearance from motion, and repurpose it as the base model for instruction-based human-centric video editing by reference frame editing and video reconstruction via the collected CharEdit-50K dataset. Besides, we adapt the model from offline bidirectional to causal streaming generation, and design an aligned self-rollout distillation strategy that compresses the model into a two-step sampler, where fixed RoPE and align forcing reduce training--inference discrepancies, and first-frame preserved sparse attention filters redundant historical information to mitigate appearance drift. Extensive experiments demonstrate that EditaLive delivers state-of-the-art editing performance with faithful preservation of facial expressions and low-latency real-time streaming inference.