ContextMaster:透過固定預算稀疏上下文路由的互動式多鏡頭影片生成
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
August 5, 2026
作者: Xu Guo, Zhengxuan Wei, Xinghui Li, Hanzhuo Huang, Xinyu Liu, Xiangyang Luo, Min Wei, Yiran Zhu, Qiulin Wang, Yulong Xu, Xintao Wang, Pengfei Wan, Qi Fan, Xiangwang Hou
cs.AI
摘要
近期視頻模型日益支持在單一模型中進行生成、參考條件控制與編輯,但通常將這些功能作為對固定輸入的獨立操作來呈現。實際創作過程跨越多個鏡頭展開,需要模型既能從文本生成、跟隨參考,也能編輯源素材,同時保持共享的歷史信息。我們將這一設定形式化為互動式多鏡頭視頻創作(IMVC),並提出ContextMaster——一個具備角色感知上下文表示的統一模型,以支持上述操作。互動模型必須保持對不斷擴展之歷史信息的訪問能力,同時不允許每次去噪步驟的上下文讀取成本隨之增長。ContextMaster將可重用的乾淨上下文狀態與固定預算的稀疏上下文路由相結合,並利用ConstraintSink使任務約束始終可見。為應對稀疏上下文訪問與少步去噪推論的雙重挑戰,我們提出了一個兩階段特權上下文蒸餾框架,先通過一致性蒸餾從密集教師模型遷移完整上下文行為,再以分佈匹配優化部署階段的推論。在三種基本任務上的實驗表明,相較於專門設計的基線方法,本模型在任務完成度與跨鏡頭一致性上均有提升。用戶研究進一步驗證了靈活組合的工作流程,且該模型可在單張GPU上達到16 FPS的幀率。
English
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow a reference, or edit source footage while maintaining shared history. We formalize this setting as interactive multi-shot video creation (IMVC) and introduce ContextMaster, a unified model with a role-aware context representation for these operations. An interactive model must retain access to an expanding history without allowing the context read cost at each denoising step to grow. ContextMaster combines reusable clean context states with fixed budget sparse context routing and uses ConstraintSink to keep task constraints visible. To address the dual challenges of sparse context access and inference with few denoising steps, we propose a two-stage privileged context distillation framework, which transfers full context behavior from a dense teacher through consistency distillation and then refines deployment rollouts with distribution matching. Experiments on the three primitive tasks demonstrate improved task fulfillment and consistency across shots over specialized baselines. User studies further validate flexibly composed workflows, while the model reaches 16 FPS on a single GPU.