ShotPlan:基於可學習規劃標記的電影級影片生成
ShotPlan: Cinematic Video Generation with Learnable Planning Token
July 20, 2026
作者: Su Guo, Guangce Liu, Haosen Yang, Jiepeng Wang, Cong Liu, Junqi Liu, Haibin Huang, Hongxun Yao, Chi Zhang, Xuelong Li
cs.AI
摘要
現有的影片生成模型在單次生成上表現出色,但在電影級影片生成中仍有限制,因為連貫的敘事與有效的多鏡頭構圖需要明確的鏡頭規劃。為解決此挑戰,我們提出 ShotPlan,一個基於影片擴散基礎模型、專為明確多鏡頭電影級影片生成設計的框架。我們的方法引入可學習的規劃標記,以捕捉鏡頭轉換的過渡線索,並能與原始影片生成標記無縫整合,以控制轉換的時間點。與標準影片生成標記不同,所提出的規劃標記配備了分數時間旋轉位置嵌入(FRoPE),使鏡頭轉換能在幀層級進行建模。實驗證明,ShotPlan 在現有電影級影片生成方法中表現顯著優異,提供更靈活的鏡頭管理與更強的鏡頭間一致性。
English
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot planning. To address this challenge, we propose ShotPlan, a framework for explicit multi-shot cinematic video generation built upon a video diffusion foundation model. Our method introduces learnable planning tokens that capture shot-level transition cues and can be seamlessly integrated with the original video generation tokens to control transition timestamps. Unlike standard video generation tokens, the proposed planning tokens are equipped with Fractional Temporal Rotary Position Embedding (FRoPE), enabling shot transitions to be modeled at the frame level. Experiments demonstrate that ShotPlan significantly outperforms existing cinematic video generation methods, offering more flexible shot management and stronger inter-shot consistency.