ChatPaper.aiChatPaper

LayerRecall: 비디오 생성에서 장기 일관성을 위한 상태 조건부 메모리 라우터

LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation

August 28, 2026
저자: Yixuan Ding, Jiahao Kong, Wei Huang, Ruijie Quan, Yi Yang
cs.AI

초록

자기회귀 비디오 확산은 제한된 최근 컨텍스트에서 청크를 생성함으로써 확장 가능한 장기 비디오 생성을 가능하게 한다. 최신성 기반 캐싱은 지역적 연속성을 보존하지만, 주체, 객체, 장면 또는 속성이 다시 등장할 때 필요한 과거 단서를 제거한다. 기존 메모리 메커니즘은 모델이 비국지적 과거에 접근할 수 있게 하지만, 접근만으로는 효과적 활용이 보장되지 않는다. 우리의 분석에 따르면 비디오 DiT 계층은 현재, 최근, 먼 컨텍스트에 대해 서로 다른 선호도를 보이며, 이는 장거리 메모리가 무엇을 검색할지와 어디에 사용할지 모두 결정해야 함을 시사한다. 우리는 현재-조건부 계층 선택적 메모리 라우터인 LayerRecall을 도입한다. 이는 관련된 과거 K/V 상태를 검색하여 백본 특정 메모리 민감 계층에만 주입하고, 다른 곳에서는 지역적 어텐션을 보존한다. 희소한 고품질 장기 비디오와 명시적 메모리 할당 레이블에 대한 의존도를 줄이기 위해, 우리는 특권적 장기 컨텍스트 참조를 사용하여 예측 공간에서 제한된 메모리 라우터를 감독하는 교차-지평선 예측 정합(CHPM)을 추가로 제안한다. 100개의 멀티샷 평가 프롬프트에서 LayerRecall은 MemoBench와 MovieBench에서 최고의 전체 성능을 달성하고 VBench-Long에서 백본과 동등한 성능을 보여, 지역적 연속성을 희생하지 않으면서 더 강력한 장거리 복원을 입증한다. 정성적 분석은 메모리 유도 자기 수정을 추가로 밝혀낸다. 이는 초기에 불일치했던 지역적 속성이 진행 중인 모션 또는 장면 구조를 재설정하지 않고 과거 외관으로 복귀하는 것이다. 추가 분석은 교차-백본 이식성과 무시할 수 있는 추론 오버헤드를 보여준다.
English
Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent context. While recency-based caching preserves local continuity, it evicts historical cues needed when subjects, objects, scenes, or attributes reappear. Existing memory mechanisms expose models to nonlocal history, but access alone does not ensure effective use. Our analysis reveals that video DiT layers exhibit distinct preferences for current, recent, and distant context, suggesting that long-range memory requires deciding both what to retrieve and where to use it. We introduce LayerRecall, a current-conditioned, layer-selective memory router that retrieves relevant historical K/V states and injects them only into backbone-specific memory-sensitive layers while preserving local attention elsewhere. To reduce reliance on scarce high-quality long-horizon videos and explicit memory-allocation labels, we further propose Cross-Horizon Prediction Matching (CHPM), which uses a privileged long-context reference to supervise the bounded-memory router in prediction space. Across 100 multi-shot evaluation prompts, LayerRecall achieves the best overall results on MemoBench and MovieBench while matching its backbone on VBench-Long, demonstrating stronger long-range recovery without sacrificing local continuity. Qualitative analyses further reveal memory-guided self-correction, whereby initially mismatched local attributes return to their historical appearance without resetting ongoing motion or scene structure. Additional analyses show cross-backbone portability and negligible inference overhead.