ContextMaster: 고정 예산 희소 컨텍스트 라우팅을 통한 상호작용적 멀티샷 비디오 생성
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
August 5, 2026
저자: Xu Guo, Zhengxuan Wei, Xinghui Li, Hanzhuo Huang, Xinyu Liu, Xiangyang Luo, Min Wei, Yiran Zhu, Qiulin Wang, Yulong Xu, Xintao Wang, Pengfei Wan, Qi Fan, Xiangwang Hou
cs.AI
초록
최근 비디오 모델들은 점차 단일 모델 내에서 생성, 레퍼런스 조건화, 편집을 지원하지만, 일반적으로 이를 고정된 입력에 대한 별개의 연산으로 노출한다. 실제 창작은 여러 샷에 걸쳐 이루어지며, 하나의 모델이 공유 히스토리를 유지하면서 텍스트로부터 생성하고, 레퍼런스를 따르고, 소스 푸티지를 편집할 수 있어야 한다. 우리는 이 설정을 인터랙티브 멀티샷 비디오 생성(IMVC)으로 공식화하고, 이러한 연산을 위한 역할 인식 컨텍스트 표현을 갖춘 통합 모델 ContextMaster를 제안한다. 인터랙티브 모델은 각 디노이징 단계의 컨텍스트 읽기 비용이 증가하지 않으면서 확장되는 히스토리에 대한 접근을 유지해야 한다. ContextMaster는 재사용 가능한 클린 컨텍스트 상태와 고정 예산의 희소 컨텍스트 라우팅을 결합하고, ConstraintSink를 사용하여 작업 제약 조건의 가시성을 유지한다. 희소 컨텍스트 접근과 소수의 디노이징 단계를 통한 추론이라는 이중 과제를 해결하기 위해, 우리는 2단계 특권 컨텍스트 증류 프레임워크를 제안한다. 이는 일관성 증류를 통해 밀집 교사로부터 전체 컨텍스트 동작을 전이한 후, 분포 정합으로 배포 롤아웃을 정제한다. 세 가지 기본 작업에 대한 실험은 특화된 베이스라인 대비 작업 수행 및 샷 간 일관성 개선을 보여준다. 사용자 연구는 유연하게 구성된 워크플로우를 추가로 검증하며, 모델은 단일 GPU에서 16 FPS를 달성한다.
English
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow a reference, or edit source footage while maintaining shared history. We formalize this setting as interactive multi-shot video creation (IMVC) and introduce ContextMaster, a unified model with a role-aware context representation for these operations. An interactive model must retain access to an expanding history without allowing the context read cost at each denoising step to grow. ContextMaster combines reusable clean context states with fixed budget sparse context routing and uses ConstraintSink to keep task constraints visible. To address the dual challenges of sparse context access and inference with few denoising steps, we propose a two-stage privileged context distillation framework, which transfers full context behavior from a dense teacher through consistency distillation and then refines deployment rollouts with distribution matching. Experiments on the three primitive tasks demonstrate improved task fulfillment and consistency across shots over specialized baselines. User studies further validate flexibly composed workflows, while the model reaches 16 FPS on a single GPU.