RECAP-Forcing: 장기 비디오 생성을 위한 콘텐츠 외형 유지
RECAP-Forcing: Retaining Content Appearances for Long Video Generation
August 27, 2026
저자: Haiyang Xu, Zheng Ding, Zhuowen Tu
cs.AI
초록
긴 자기회귀 비디오 생성은 근본적인 메모리 문제에 직면한다. 유한한 어텐션 윈도우 안에서 모델은 계속해서 확장되는 과거 정보 중 무엇을 유지할지 결정해야 한다. 기존 방법은 메모리를 시간적으로 구성하여 최근 프레임을 보존하고 오래된 프레임은 압축하거나 폐기한다. 반면 우리는 외관 참신성(appearance novelty)에 따라 메모리를 구성하는 RECAP-Forcing을 제안한다. 긴 비디오는 단순히 프레임의 연속이 아니라, 시간이 지나도 정체성이 일관되게 유지되어야 하는 주체, 객체, 장면들이 진화하며 구성된 집합이다. 우리는 새롭게 나타나는 콘텐츠, 예를 들어 새로 등장하는 주체, 폐색이 해제된 영역, 새로 도입된 장면과 연관된 KV 캐시를 해당 콘텐츠가 처음으로 보이는 순간에 유지함으로써 메모리를 구성한다. 이때 최근성보다는 참신성을 우선한다. 메모리는 비디오 길이가 아니라 새로 도입된 콘텐츠의 양에 비례하여 확장되어야 한다. 이러한 외관 기반 인덱싱 메모리는 장기적 일관성을 메모리 구조의 명시적 속성으로 만든다. 우리의 프레임워크는 이 단일 원칙 아래 두 가지 메커니즘을 통합한다. 비디오의 시작 부분에서는 모든 가시 콘텐츠가 새롭기 때문에 어텐션 싱크(attention sink)가 초기 장면을 보존한다. 비디오가 전개됨에 따라, 옵티컬 플로우 기반 참신성 뱅크(novelty bank)가 새롭게 드러난 콘텐츠를 선택적으로 유지함으로써 동일한 원칙을 확장한다. 학습이 필요 없는 추론 방법으로서 추가적인 학습 가능한 파라미터가 없는 RECAP-Forcing은 여러 강력한 기준선에 걸쳐 시각적 품질과 의미론적 충실도를 일관되게 향상시키며 기존 메모리 방법들을 능가한다.
English
Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model must decide which information from an ever-expanding history to retain. Existing methods organize memory temporally, preserving recent frames while compressing or discarding older ones. We instead propose RECAP-Forcing, organizing memory by appearance novelty. A long video is not merely a sequence of frames, but an evolving cast of subjects, objects, and scenes whose identities must remain consistent over time. We organize memory by retaining the KV cache associated with newly appearing content--such as entering subjects, disoccluded regions, and newly introduced scenes--at the moment it first becomes visible, prioritizing novelty over recency. Memory should scale with the amount of newly introduced content, rather than with video length. This appearance-indexed memory makes long-range consistency an explicit property of the memory structure. Our framework unifies two mechanisms under this single principle. At the beginning of a video, when all visible content is novel, an attention sink preserves the initial scene. As the video evolves, an optical-flow-based novelty bank extends the same principle by selectively retaining newly revealed content. As a training-free inference method with no additional learnable parameters, RECAP-Forcing consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.