ContextMaster: 固定予算疎コンテキストルーティングによるインタラクティブなマルチショット動画生成
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
August 5, 2026
著者: Xu Guo, Zhengxuan Wei, Xinghui Li, Hanzhuo Huang, Xinyu Liu, Xiangyang Luo, Min Wei, Yiran Zhu, Qiulin Wang, Yulong Xu, Xintao Wang, Pengfei Wan, Qi Fan, Xiangwang Hou
cs.AI
要旨
近年のビデオモデルは、単一モデル内で生成、参照条件付け、編集をますますサポートするようになっているが、通常これらは固定入力に対する個別の操作として公開されている。実際の制作は複数のショットにまたがって展開され、単一のモデルがテキストからの生成、参照への追従、ソース映像の編集を行いながら、共有履歴を維持することが求められる。我々はこの設定を対話型マルチショットビデオ生成(IMVC)として形式化し、これらの操作のためのロール認識型コンテキスト表現を備えた統合モデルであるContextMasterを提案する。対話型モデルは、拡大し続ける履歴へのアクセスを維持しつつ、各デノイジングステップにおけるコンテキスト読み取りコストを増大させない必要がある。ContextMasterは、再利用可能なクリーンコンテキスト状態と固定予算のスパースコンテキストルーティングを組み合わせ、タスク制約を可視化し続けるためにConstraintSinkを用いる。スパースコンテキストアクセスと少数のデノイジングステップによる推論という二重の課題に対処するため、我々は二段階の特権コンテキスト蒸留フレームワークを提案する。これは、一貫性蒸留を通じて高密度教師から完全なコンテキスト挙動を転移し、その後、分布マッチングによりデプロイメント時のロールアウトを洗練する。三つの基本タスクに関する実験は、専門特化したベースラインと比較して、タスク達成度とショット間の一貫性が向上することを示す。ユーザー研究はさらに、柔軟に構成されたワークフローを検証しており、モデルは単一GPU上で16 FPSを達成する。
English
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow a reference, or edit source footage while maintaining shared history. We formalize this setting as interactive multi-shot video creation (IMVC) and introduce ContextMaster, a unified model with a role-aware context representation for these operations. An interactive model must retain access to an expanding history without allowing the context read cost at each denoising step to grow. ContextMaster combines reusable clean context states with fixed budget sparse context routing and uses ConstraintSink to keep task constraints visible. To address the dual challenges of sparse context access and inference with few denoising steps, we propose a two-stage privileged context distillation framework, which transfers full context behavior from a dense teacher through consistency distillation and then refines deployment rollouts with distribution matching. Experiments on the three primitive tasks demonstrate improved task fulfillment and consistency across shots over specialized baselines. User studies further validate flexibly composed workflows, while the model reaches 16 FPS on a single GPU.