從合成到去除:物理驅動的反射模擬與基於擴散的影片去反射
From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
August 12, 2026
作者: Zepeng Wang, Jiagao Hu, Fuhao Li, Yuxuan Chen, Fei Wang, Daiguo Zhou
cs.AI
摘要
透過玻璃拍攝的影片常含有反射,這些反射會降低視覺品質並干擾下游視覺任務。雖然單張影像的反射移除已獲廣泛研究,但影片反射移除在很大程度上仍未被充分探索,原因在於缺乏成對的影片資料、具時間連貫性的移除模型,以及專門的評測基準。我們提出一個閉環框架,統一了基於物理的反射模擬、基於擴散的影片去反射與基準評測。我們的 S2R-Synthesis 流程在結構空間中執行基於物理的增強,並使用訓練過的影片擴散渲染器渲染逼真的反射影片,從而生成成對的含反射與無反射影片;該增強模擬了與玻璃相關的主要效果,包括粗糙度引起的模糊、厚度引起的重影以及反射率變化。基於這些合成資料,我們提出 S2R-Removal,這是第一個基於擴散的影片反射移除模型。它透過反射感知的潛在空間適應與單步像素-幾何精化,來適應預先訓練的影片擴散先驗,並在單次去噪步驟中恢復乾淨的透射成分。我們進一步建構 S2R-Bench,這是第一個用於影片反射移除的基準,支援全參考評估與真實世界的人類感知評估。在 S2R-Bench 和多個公開影像基準上的實驗顯示,我們的方法達到最先進的效能,且推論速度甚至比非擴散基線更快,並驗證了 S2R-Synthesis 的有效性。專案頁面:https://codingwzp.github.io/VideoDereflection_S2R。
English
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, video reflection removal remains largely underexplored due to the lack of paired video data, temporally coherent removal models, and dedicated evaluation benchmarks. We present a closed-loop framework that unifies physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation. Our S2R-Synthesis pipeline generates paired reflected and reflection-free videos by performing physics-grounded augmentation in the structure space and rendering realistic reflected videos with a trained video diffusion renderer; the augmentation models key glass-related effects including roughness-induced blur, thickness-induced ghosting, and reflectance variation. Based on the synthesized data, we introduce S2R-Removal, the first diffusion-based video reflection removal model, which adapts a pretrained video diffusion prior through reflection-aware latent adaptation and one-step pixel-geometric refinement, recovering the clean transmission in a single denoising step. We further build S2R-Bench, the first benchmark for video reflection removal, supporting both full-reference evaluation and real-world human perceptual assessment. Experiments on S2R-Bench and multiple public image benchmarks demonstrate state-of-the-art performance and faster inference than even non-diffusion baselines, and validate the effectiveness of S2R-Synthesis. Project page: https://codingwzp.github.io/VideoDereflection_S2R.