ShotPlan: 학습 가능한 계획 토큰을 이용한 시네마틱 비디오 생성
ShotPlan: Cinematic Video Generation with Learnable Planning Token
July 20, 2026
저자: Su Guo, Guangce Liu, Haosen Yang, Jiepeng Wang, Cong Liu, Junqi Liu, Haibin Huang, Hongxun Yao, Chi Zhang, Xuelong Li
cs.AI
초록
현재의 비디오 생성 모델은 단일 샷 생성에서 인상적인 결과를 보여주지만, 일관된 내러티브와 효과적인 멀티 샷 구성을 위해 명시적인 샷 계획이 필요한 영화적 비디오 생성에서는 여전히 한계가 있다. 이러한 문제를 해결하기 위해, 우리는 비디오 확산 기초 모델을 기반으로 한 명시적 멀티 샷 영화적 비디오 생성을 위한 프레임워크인 ShotPlan을 제안한다. 우리의 방법은 샷 수준의 전환 신호를 포착하는 학습 가능한 계획 토큰을 도입하며, 이는 원래의 비디오 생성 토큰과 원활하게 통합되어 전환 타임스탬프를 제어할 수 있다. 표준 비디오 생성 토큰과 달리, 제안된 계획 토큰에는 FRoPE(분수 시간 회전 위치 임베딩)가 탑재되어 샷 전환을 프레임 수준에서 모델링할 수 있다. 실험 결과, ShotPlan은 기존의 영화적 비디오 생성 방법보다 훨씬 뛰어난 성능을 보여주며, 더 유연한 샷 관리와 더 강력한 샷 간 일관성을 제공한다.
English
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot planning. To address this challenge, we propose ShotPlan, a framework for explicit multi-shot cinematic video generation built upon a video diffusion foundation model. Our method introduces learnable planning tokens that capture shot-level transition cues and can be seamlessly integrated with the original video generation tokens to control transition timestamps. Unlike standard video generation tokens, the proposed planning tokens are equipped with Fractional Temporal Rotary Position Embedding (FRoPE), enabling shot transitions to be modeled at the frame level. Experiments demonstrate that ShotPlan significantly outperforms existing cinematic video generation methods, offering more flexible shot management and stronger inter-shot consistency.