合成から除去へ:物理に基づく反射シミュレーションと拡散モデルベースのビデオ反射除去
From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
August 12, 2026
著者: Zepeng Wang, Jiagao Hu, Fuhao Li, Yuxuan Chen, Fei Wang, Daiguo Zhou
cs.AI
要旨
ガラス越しに撮影されたビデオには、画質を劣化させ、下流の視覚タスクの妨げとなる反射がしばしば含まれる。単一画像の反射除去は広く研究されてきたが、ビデオ反射除去は、ペアのビデオデータ、時間的に一貫した除去モデル、専用の評価ベンチマークが不足しているため、依然としてほとんど未開拓のままである。本稿では、物理に基づく反射シミュレーション、拡散ベースのビデオ反射除去、ベンチマーク評価を統合した閉ループフレームワークを提案する。我々のS2R-Synthesisパイプラインは、構造空間で物理に基づく拡張を実行し、学習済みビデオ拡散レンダラーで現実的な反射を含むビデオをレンダリングすることにより、反射ありと反射なしのペアビデオを生成する。この拡張は、粗さに起因するぼけ、厚さに起因するゴースト、反射率の変動など、ガラスに関連する主要な効果をモデル化する。合成データに基づき、我々は初の拡散ベースのビデオ反射除去モデルであるS2R-Removalを導入する。これは、反射を考慮した潜在適応と単一ステップの画素・幾何学的精緻化を通じて、事前学習済みのビデオ拡散事前分布を適応させ、単一のノイズ除去ステップでクリーンな透過成分を復元する。さらに、ビデオ反射除去のための初のベンチマークであるS2R-Benchを構築し、フルリファレンス評価と実世界の人間の知覚評価の両方をサポートする。S2R-Benchおよび複数の公開画像ベンチマークでの実験により、最先端の性能と、非拡散ベースラインをも上回る高速な推論を実証し、S2R-Synthesisの有効性を検証する。プロジェクトページ: https://codingwzp.github.io/VideoDereflection_S2R.
English
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, video reflection removal remains largely underexplored due to the lack of paired video data, temporally coherent removal models, and dedicated evaluation benchmarks. We present a closed-loop framework that unifies physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation. Our S2R-Synthesis pipeline generates paired reflected and reflection-free videos by performing physics-grounded augmentation in the structure space and rendering realistic reflected videos with a trained video diffusion renderer; the augmentation models key glass-related effects including roughness-induced blur, thickness-induced ghosting, and reflectance variation. Based on the synthesized data, we introduce S2R-Removal, the first diffusion-based video reflection removal model, which adapts a pretrained video diffusion prior through reflection-aware latent adaptation and one-step pixel-geometric refinement, recovering the clean transmission in a single denoising step. We further build S2R-Bench, the first benchmark for video reflection removal, supporting both full-reference evaluation and real-world human perceptual assessment. Experiments on S2R-Bench and multiple public image benchmarks demonstrate state-of-the-art performance and faster inference than even non-diffusion baselines, and validate the effectiveness of S2R-Synthesis. Project page: https://codingwzp.github.io/VideoDereflection_S2R.