ChatPaper.aiChatPaper

장편 다중 샷 비디오 생성을 위한 멀티그리드 후속 학습

Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

September 6, 2026
저자: Jiawei Mao, Haoqin Tu, Hardy Chen, Yuhan Wang, Keyang Xu, Jieru Mei, Hongliang Fei, Ruogu Fang, Wei Shao, Cihang Xie, Yuyin Zhou
cs.AI

초록

장편 다중 샷 비디오를 생성하려면 샷 내의 일관된 움직임과 샷 간의 시각적으로 일관된 서사가 필요하다. 기존 비디오 생성기는 연속적인 움직임을 선호하며, 전체 서사를 하나의 시간축에 압축하여 배치할 경우 완전한 샷 집합을 제시하는 데 어려움을 겪는다. 우리는 장편 비디오를 더 짧고 시간 순서대로 정렬된 청크로 분해하고 이를 공간 그리드에 배치하여 공동 모델링하는 Multi-Grid Post-Training 패러다임인 MovieGrid를 제안한다. 이 설계는 각 시간축이 처리하는 샷 수를 줄이는 동시에 청크 간 전역 정보 교환을 가능하게 한다. 우리는 소스 비디오 수집, 계층적 분할, 그리드 비디오 구성, 캐릭터 인식 스토리 주석을 사용하여 1,000개의 장편 비디오로부터 Multi-Grid Long Video (MGLV) 데이터셋을 구축하고, 스토리 프롬프트와 쌍을 이루는 54K개의 그리드 비디오를 생성한다. 우리의 Noise-Free Random-Grid Training은 나머지 청크를 노이즈 제거하기 위한 깨끗한 시각적 문맥으로 무작위 청크 부분집합을 유지한다. Grid Embedding은 그리드 구조를 인코딩하고, 캐릭터 인식 스토리 프롬프트는 반복 등장하는 개체를 연결하며, Grid Boundary Loss는 레이아웃을 안정화한다. 동일한 토큰 예산 하에서 MovieGrid는 1,616프레임 비디오에서 Temporal Packing보다 6.05배 더 많은 샷을 생성한다. 다섯 가지 실제 세계 범주에 걸친 벤치마크에서, 이는 최고 수준의 샷 내 일관성(0.9131 대 HoloCine의 0.8086)과 샷 간 일관성(0.5914 대 StoryMem의 0.5384)을 달성한다. MovieGrid는 단일 또는 다중 생성을 통해 성능 저하를 최소화하면서 비디오 길이를 더욱 확장할 수 있다.
English
Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives across shots. Existing video generators favor continuous motion and struggle to present complete shot sets when an entire narrative is packed along one temporal axis. We propose MovieGrid, a Multi-Grid Post-Training paradigm that decomposes a long video into shorter, temporally ordered chunks and arranges them on a spatial grid for joint modeling. This design reduces the number of shots handled by each temporal axis while enabling global information exchange across chunks. We construct the Multi-Grid Long Video (MGLV) dataset from 1,000 long-form videos using source video collection, hierarchical segmentation, grid video construction, and character-aware story annotation, producing 54K grid videos paired with story prompts. Our Noise-Free Random-Grid Training retains a random subset of chunks as clean visual context for denoising the remaining chunks. Grid Embedding encodes grid structure, character-aware Story Prompts link recurring entities, and Grid Boundary Loss stabilizes layouts. Under the same token budget, MovieGrid generates 6.05 times more shots than Temporal Packing in a 1,616-frame video. On a benchmark spanning five real-world categories, it achieves state-of-the-art intra-shot consistency (0.9131 versus 0.8086 for HoloCine) and inter-shot consistency (0.5914 versus 0.5384 for StoryMem). MovieGrid can further scale video length with minimal compromise through single or multiple generations.