ShotPlan: 基于可学习规划令牌的电影级视频生成
ShotPlan: Cinematic Video Generation with Learnable Planning Token
July 20, 2026
作者: Su Guo, Guangce Liu, Haosen Yang, Jiepeng Wang, Cong Liu, Junqi Liu, Haibin Huang, Hongxun Yao, Chi Zhang, Xuelong Li
cs.AI
摘要
当前视频生成模型在单镜头生成方面取得了显著成果,但在电影视频生成方面仍存在局限,因为连贯的叙事和有效的多镜头构图需要明确的镜头规划。为解决这一挑战,我们提出ShotPlan——一个基于视频扩散基础模型构建的显式多镜头电影视频生成框架。该方法引入可学习的规划令牌,用于捕捉镜头级过渡线索,并与原始视频生成令牌无缝集成,以控制过渡时间戳。与标准视频生成令牌不同,所提出的规划令牌配备了分数时间旋转位置嵌入(FRoPE),使得镜头过渡能够在帧级别进行建模。实验表明,ShotPlan显著优于现有电影视频生成方法,提供了更灵活的镜头管理和更强的镜头间一致性。
English
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot planning. To address this challenge, we propose ShotPlan, a framework for explicit multi-shot cinematic video generation built upon a video diffusion foundation model. Our method introduces learnable planning tokens that capture shot-level transition cues and can be seamlessly integrated with the original video generation tokens to control transition timestamps. Unlike standard video generation tokens, the proposed planning tokens are equipped with Fractional Temporal Rotary Position Embedding (FRoPE), enabling shot transitions to be modeled at the frame level. Experiments demonstrate that ShotPlan significantly outperforms existing cinematic video generation methods, offering more flexible shot management and stronger inter-shot consistency.