ShotPlan: 学習可能な計画トークンを用いたシネマティックビデオ生成
ShotPlan: Cinematic Video Generation with Learnable Planning Token
July 20, 2026
著者: Su Guo, Guangce Liu, Haosen Yang, Jiepeng Wang, Cong Liu, Junqi Liu, Haibin Huang, Hongxun Yao, Chi Zhang, Xuelong Li
cs.AI
要旨
現在の映像生成モデルは単一ショット生成において優れた結果を達成しているが、一貫したナラティブと効果的なマルチショット構成に明示的なショット計画を必要とする映画的映像生成では限界がある。この課題に対処するため、我々は映像拡散基盤モデルに基づく明示的なマルチショット映画的映像生成のためのフレームワーク、ShotPlanを提案する。本手法では、ショットレベルの遷移キューを捉える学習可能な計画トークンを導入し、元の映像生成トークンとシームレスに統合することで遷移タイムスタンプを制御する。標準的な映像生成トークンとは異なり、提案する計画トークンにはFRoPE(分数時間回転位置埋め込み)が装備されており、フレームレベルでのショット遷移のモデリングが可能となる。実験により、ShotPlanは既存の映画的映像生成手法を大幅に上回り、より柔軟なショット管理と強力なショット間一貫性を提供することが示された。
English
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot planning. To address this challenge, we propose ShotPlan, a framework for explicit multi-shot cinematic video generation built upon a video diffusion foundation model. Our method introduces learnable planning tokens that capture shot-level transition cues and can be seamlessly integrated with the original video generation tokens to control transition timestamps. Unlike standard video generation tokens, the proposed planning tokens are equipped with Fractional Temporal Rotary Position Embedding (FRoPE), enabling shot transitions to be modeled at the frame level. Experiments demonstrate that ShotPlan significantly outperforms existing cinematic video generation methods, offering more flexible shot management and stronger inter-shot consistency.