ChatPaper.aiChatPaper

LongE2V: 비디오 확산 모델을 활용한 장기적 이벤트 기반 비디오 복원, 예측 및 프레임 보간

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models

July 9, 2026
저자: Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin, Kun-Ru Wu, Yu-Chee Tseng, Yu-Lun Liu
cs.AI

초록

희소 이벤트 스트림에서 고품질 비디오를 복원하는 것은 까다로운 작업이다. 회귀 방법은 종종 텍스처를 흐리게 만드는 반면, 기존 생성 모델은 장기 안정성에 어려움을 겪는다. 우리는 사전 학습된 비디오 확산 사전을 활용하여 이벤트 기반 비디오 복원, 예측 및 프레임 보간을 공동으로 처리하는 새로운 접근법인 LongE2V를 제안한다. 기반 비디오 모델을 미세 조정함으로써, 우리의 접근 방식은 높은 데이터 효율성과 우수한 지각 품질을 달성한다. 극도로 긴 시퀀스에서 시간적 드리프트를 완화하기 위해 자기회귀 언롤링과 적응형 컨텍스트 전환을 도입한다. 또한 프레임 보간 중 정밀한 양방향 일관성을 보장하기 위해 교차 잔차 보정을 통한 재인코딩 정렬을 제안한다. 더불어 이벤트 복셀 밀도 증강은 다양한 센서 해상도에서 강건성을 보장한다. 실제 벤치마크에 대한 광범위한 실험을 통해 LongE2V가 세 가지 작업 모두에서 최첨단 방법을 능가하며, 뛰어난 시간적 일관성과 제로샷 일반화를 보여줌을 입증한다. 프로젝트 페이지: https://cdfan0627.github.io/LongE2V-page/
English
Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation. By fine-tuning a foundational video model, our approach achieves high data efficiency and superior perceptual quality. We introduce Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences. We also propose Reencoding Alignment with Cross Residual Correction to ensure precise bidirectional consistency during frame interpolation. Furthermore, Event Voxel Density Augmentation ensures robustness across varying sensor resolutions. Extensive experiments on real-world benchmarks demonstrate that LongE2V outperforms state-of-the-art methods across all three tasks, exhibiting exceptional temporal coherence and zero-shot generalization. Project page: https://cdfan0627.github.io/LongE2V-page/