ChatPaper.aiChatPaper

LongE2V: ビデオ拡散モデルを用いた長期間イベントベースのビデオ再構成、予測、およびフレーム補間

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models

July 9, 2026
著者: Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin, Kun-Ru Wu, Yu-Chee Tseng, Yu-Lun Liu
cs.AI

要旨

疎なイベントストリームから高品質な映像を復元することは難しい課題である。回帰手法ではしばしばテクスチャがぼやけ、既存の生成モデルは長期的な安定性に課題を抱える。我々はLongE2Vを提案する。これは事前学習されたビデオ拡散事前分布を活用し、イベントベースの映像再構成、予測、フレーム補間を統合的に扱う新しいアプローチである。基本となる映像モデルをファインチューニングすることで、高いデータ効率と優れた知覚品質を実現する。極めて長いシーケンスにおける時間的ドリフトを軽減するために、自己回帰的展開と適応的コンテキスト切り替えを導入する。また、フレーム補間における正確な双方向整合性を保証するために、交差残差補正を伴う再エンコーディングアライメントを提案する。さらに、イベントボクセル密度拡張により、異なるセンサ解像度間でのロバスト性を確保する。実世界ベンチマークにおける広範な実験により、LongE2Vはこれら3つのタスクすべてにおいて最先端手法を上回り、優れた時間的コヒーレンスとゼロショット汎化を示す。プロジェクトページ: https://cdfan0627.github.io/LongE2V-page/
English
Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation. By fine-tuning a foundational video model, our approach achieves high data efficiency and superior perceptual quality. We introduce Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences. We also propose Reencoding Alignment with Cross Residual Correction to ensure precise bidirectional consistency during frame interpolation. Furthermore, Event Voxel Density Augmentation ensures robustness across varying sensor resolutions. Extensive experiments on real-world benchmarks demonstrate that LongE2V outperforms state-of-the-art methods across all three tasks, exhibiting exceptional temporal coherence and zero-shot generalization. Project page: https://cdfan0627.github.io/LongE2V-page/