ChatPaper.aiChatPaper

자기 기울기 강제: 고유한 장기 비디오 외삽

Self Gradient Forcing: Native Long Video Extrapolation

July 22, 2026
저자: Junhao Zhuang, Shiyi Zhang, Yuxuan Bian, Yaowei Li, Yawen Luo, Yijun Liu, Weiyang Jin, Songchun Zhang, Xianglong He, Xuying Zhang, Haoran Li, Haoyang Huang, Zeyue Xue, Nan Duan
cs.AI

초록

최근 자기회귀 비디오 확산 방법들은 점점 Self Forcing(자기 강제) 방식을 기반으로 구축되고 있는데, 이는 모델(학생)이 실제 비디오 컨텍스트가 아닌 자체 롤아웃에서 생성된 이력을 바탕으로 학습된다. 이는 노출 편향을 줄여주지만, 과거 키-값 캐시는 여전히 미래 프레임에 의해 단순히 고정된 롤아웃 상태로만 사용된다. 그 결과, 미래 손실이 초기 생성된 잠재 변수를 이후 비디오 잠재 생성에 더 유용한 키와 값으로 기록하도록 지도할 수 없다. 우리는 이를 역사적 컨텍스트-그래디언트 격차라고 부른다. 우리는 Self Gradient Forcing (SGF, 자기 그래디언트 강제)를 제안한다. 이는 전체 직렬 롤아웃을 통해 역전파하지 않으면서 누락된 지도 신호를 복원하는 이중 패스 학습 전략이다. 패스 1은 추론과 일치하는 그래디언트 없는 자기회귀 롤아웃을 수행하고, 샘플링된 잡음 제거 종료 단계에서 자체 생성된 컨텍스트와 모델에 입력된 노이즈가 있는 잠재 변수를 모두 기록한다. 패스 2는 기록된 종료 단계에 대해 병렬 컨텍스트-그래디언트 재구성을 수행한다. 생성된 컨텍스트는 그래디언트 정지 클린 잠재 입력으로 사용되는 반면, 모델은 컨텍스트 KV 표현과 미래-컨텍스트 인과 어텐션을 재계산한다. 따라서 SGF는 본질적인 자기회귀 학습 목표 내에서 누락된 메모리 기록 지도를 제공하며, 미래 비디오 잠재 변수에 대한 손실을 사용하여 모델이 컨텍스트를 보다 효과적인 인과 메모리로 인코딩하도록 훈련한다. 다양한 초기화 조건에서 장기 프레임 단위 및 청크 단위의 광범위한 실험을 통해 SGF는 Self Forcing보다 특히 주제 정체성, 배경/레이아웃 일관성 및 시간적 안정성에서 더 뛰어난 본질적 장기 비디오 외삽 성능을 달성한다. 놀랍게도, 단 5초의 훈련 윈도우만 사용하여 SGF는 수 분 길이의 비디오로 외삽할 수 있다. 코드와 모델은 자기회귀 비디오 생성 연구를 발전시키기 위해 공개될 예정이다.
English
Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as frozen rollout state. As a result, future losses cannot supervise how earlier generated latents should be written into more useful keys and values for later video-latent generation. We call this the historical context-gradient gap. We propose Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing supervision signal without backpropagating through the full serial rollout. Pass 1 performs a no-gradient autoregressive rollout matching inference and, at a sampled denoising exit step, records both the self-generated context and the noisy latents fed to the model. Pass 2 performs parallel context-gradient reconstruction for the recorded exit step. The generated context is used as stop-gradient clean-latent input, while the model recomputes the context KV representations and future-to-context causal attention. Thus, SGF provides the missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory. Across extensive long-horizon frame-wise and chunk-wise experiments under different initializations, SGF achieves stronger native long-video extrapolation than Self Forcing, especially in subject identity, background/layout consistency, and temporal stability. Remarkably, using only a 5-second training window, SGF can extrapolate to videos lasting several minutes. Code and models will be released to advance research on autoregressive video generation.