ChatPaper.aiChatPaper

EditaLive! 라이브 스트리밍을 위한 통합 캐릭터 비디오 편집

EditaLive! Unified Character Video Editing for Live Streaming

August 27, 2026
저자: Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang, Bo Li, Xiaodong Cun
cs.AI

초록

기존 비디오 편집은 주로 장면 수준의 콘텐츠에 초점을 맞추는 반면, 라이브 스트리밍은 인간 피사체에 더 큰 중점을 둔다. 그러나 기존 비디오 편집 방법을 인간 중심의 라이브 스트리밍에 직접 적용하는 것은 여전히 어려운데, 이는 얼굴 표정 불일치를 초래할 수 있고 일반적으로 여러 번의 오프라인 추론 단계에 의존하여 실시간 상호작용에 적합하지 않기 때문이다. 본 논문에서는 실시간 스트리밍 캐릭터 비디오 편집을 위한 새로운 프레임워크인 EditaLive를 제안한다. 구체적으로, 외관과 모션을 자연스럽게 분리하는 사전 학습된 이미지 애니메이션 모델(Wan-Animate)을 기반으로 하여, 수집된 CharEdit-50K 데이터셋을 통해 참조 프레임 편집과 비디오 재구성을 수행함으로써 이를 지시 기반 인간 중심 비디오 편집을 위한 기본 모델로 재활용한다. 또한, 오프라인 양방향 모델을 인과적 스트리밍 생성 방식으로 변환하고, 모델을 2단계 샘플러로 압축하는 정렬된 자기 롤아웃 증류 전략을 설계한다. 이때 고정 RoPE와 정렬 강제(align forcing)는 학습-추론 간 불일치를 줄이며, 첫 프레임 보존 희소 어텐션은 중복된 과거 정보를 필터링하여 외관 드리프트를 완화한다. 광범위한 실험을 통해 EditaLive가 얼굴 표정의 충실한 보존과 저지연 실시간 스트리밍 추론을 제공하면서 최첨단 편집 성능을 달성함을 입증한다.
English
Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decouples appearance from motion, and repurpose it as the base model for instruction-based human-centric video editing by reference frame editing and video reconstruction via the collected CharEdit-50K dataset. Besides, we adapt the model from offline bidirectional to causal streaming generation, and design an aligned self-rollout distillation strategy that compresses the model into a two-step sampler, where fixed RoPE and align forcing reduce training--inference discrepancies, and first-frame preserved sparse attention filters redundant historical information to mitigate appearance drift. Extensive experiments demonstrate that EditaLive delivers state-of-the-art editing performance with faithful preservation of facial expressions and low-latency real-time streaming inference.