LongE2V:基於事件相機的長時域影片重建、預測與幀插值 - 使用影片擴散模型
LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models
July 9, 2026
作者: Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin, Kun-Ru Wu, Yu-Chee Tseng, Yu-Lun Liu
cs.AI
摘要
從稀疏事件流中恢復高質量視頻是一項具挑戰性的任務。回歸方法常導致紋理模糊,而現有生成模型則難以維持長期穩定性。我們提出 LongE2V,一種新穎方法,利用預訓練的視頻擴散先驗共同處理基於事件的視頻重建、預測與幀插值。透過微調基礎視頻模型,我們的方法實現了高數據效率與卓越的感知質量。我們引入自回歸展開與自適應上下文切換,以減輕極長序列中的時間漂移問題。同時提出跨殘差修正的重新編碼對齊,確保幀插值過程中的精確雙向一致性。此外,事件體素密度增強機制保障了跨不同傳感器解析度的穩健性。在真實世界基準上的大量實驗表明,LongE2V 在三項任務中均優於最先進方法,展現出卓越的時間一致性與零樣本泛化能力。項目頁面:https://cdfan0627.github.io/LongE2V-page/
English
Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation. By fine-tuning a foundational video model, our approach achieves high data efficiency and superior perceptual quality. We introduce Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences. We also propose Reencoding Alignment with Cross Residual Correction to ensure precise bidirectional consistency during frame interpolation. Furthermore, Event Voxel Density Augmentation ensures robustness across varying sensor resolutions. Extensive experiments on real-world benchmarks demonstrate that LongE2V outperforms state-of-the-art methods across all three tasks, exhibiting exceptional temporal coherence and zero-shot generalization. Project page: https://cdfan0627.github.io/LongE2V-page/