ChatPaper.aiChatPaper

ContextMaster:通过固定预算稀疏上下文路由的交互式多镜头视频生成

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

August 5, 2026
作者: Xu Guo, Zhengxuan Wei, Xinghui Li, Hanzhuo Huang, Xinyu Liu, Xiangyang Luo, Min Wei, Yiran Zhu, Qiulin Wang, Yulong Xu, Xintao Wang, Pengfei Wan, Qi Fan, Xiangwang Hou
cs.AI

摘要

近期视频模型日益支持在单一模型内实现生成、参考条件控制和编辑,但通常将这些操作作为面向固定输入的独立功能。实际创作过程跨越多个镜头展开,要求模型既能根据文本生成,也能遵循参考或编辑源视频,同时保持共享历史。我们将这一设定形式化为交互式多镜头视频创作(IMVC),并提出ContextMaster——一种采用角色感知上下文表示来支持上述操作的统一模型。交互式模型必须持续访问不断扩展的历史,同时避免每次去噪步骤的上下文读取成本随之增长。ContextMaster将可复用的干净上下文状态与固定预算的稀疏上下文路由相结合,并使用ConstraintSink保持任务约束可见。针对稀疏上下文访问与少步去噪推理的双重挑战,我们提出了一种两阶段特权上下文蒸馏框架:首先通过一致性蒸馏从稠密教师模型中迁移完整上下文行为,随后通过分布匹配优化部署时的推理过程。在三个基础任务上的实验表明,与专用基线相比,该方法在任务完成度和跨镜头一致性方面均有提升。用户研究进一步验证了灵活组合的工作流程,且该模型在单块GPU上达到16 FPS。
English
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow a reference, or edit source footage while maintaining shared history. We formalize this setting as interactive multi-shot video creation (IMVC) and introduce ContextMaster, a unified model with a role-aware context representation for these operations. An interactive model must retain access to an expanding history without allowing the context read cost at each denoising step to grow. ContextMaster combines reusable clean context states with fixed budget sparse context routing and uses ConstraintSink to keep task constraints visible. To address the dual challenges of sparse context access and inference with few denoising steps, we propose a two-stage privileged context distillation framework, which transfers full context behavior from a dense teacher through consistency distillation and then refines deployment rollouts with distribution matching. Experiments on the three primitive tasks demonstrate improved task fulfillment and consistency across shots over specialized baselines. User studies further validate flexibly composed workflows, while the model reaches 16 FPS on a single GPU.