합성에서 제거까지: 물리 기반 반사 시뮬레이션과 확산 기반 비디오 반사 제거
From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
August 12, 2026
저자: Zepeng Wang, Jiagao Hu, Fuhao Li, Yuxuan Chen, Fei Wang, Daiguo Zhou
cs.AI
초록
유리를 통해 촬영된 비디오에는 시각적 품질을 저하시키고 하위 비전 작업을 방해하는 반사가 자주 포함된다. 단일 이미지 반사 제거는 광범위하게 연구되어 왔지만, 비디오 반사 제거는 쌍을 이루는 비디오 데이터, 시간적으로 일관된 제거 모델, 그리고 전용 평가 벤치마크의 부족으로 인해 여전히 크게 탐구되지 않은 상태로 남아 있다. 본 논문은 물리 기반 반사 시뮬레이션, 확산 기반 비디오 반사 제거, 벤치마크 평가를 통합하는 폐루프 프레임워크를 제시한다. 우리의 S2R-Synthesis 파이프라인은 구조 공간에서 물리 기반 증강을 수행하고 학습된 비디오 확산 렌더러로 실제적인 반사 비디오를 렌더링하여 쌍을 이루는 반사 및 무반사 비디오를 생성한다. 이 증강은 거칠기 유발 블러, 두께 유발 고스팅, 반사율 변화를 포함한 주요 유리 관련 효과를 모델링한다. 합성 데이터를 기반으로, 우리는 반사 인식 잠재 적응과 단일 단계 픽셀-기하 정제를 통해 사전 학습된 비디오 확산 사전 지식을 적응시켜, 단일 노이즈 제거 단계에서 깨끗한 투과 성분을 복구하는 최초의 확산 기반 비디오 반사 제거 모델인 S2R-Removal을 도입한다. 또한, 비디오 반사 제거를 위한 최초의 벤치마크인 S2R-Bench를 구축하여 전체 참조 평가와 실제 인간 지각 평가를 모두 지원한다. S2R-Bench 및 여러 공개 이미지 벤치마크에 대한 실험은 최첨단 성능과 비확산 기준선보다 빠른 추론 속도를 입증하며, S2R-Synthesis의 효과성을 검증한다. 프로젝트 페이지: https://codingwzp.github.io/VideoDereflection_S2R.
English
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, video reflection removal remains largely underexplored due to the lack of paired video data, temporally coherent removal models, and dedicated evaluation benchmarks. We present a closed-loop framework that unifies physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation. Our S2R-Synthesis pipeline generates paired reflected and reflection-free videos by performing physics-grounded augmentation in the structure space and rendering realistic reflected videos with a trained video diffusion renderer; the augmentation models key glass-related effects including roughness-induced blur, thickness-induced ghosting, and reflectance variation. Based on the synthesized data, we introduce S2R-Removal, the first diffusion-based video reflection removal model, which adapts a pretrained video diffusion prior through reflection-aware latent adaptation and one-step pixel-geometric refinement, recovering the clean transmission in a single denoising step. We further build S2R-Bench, the first benchmark for video reflection removal, supporting both full-reference evaluation and real-world human perceptual assessment. Experiments on S2R-Bench and multiple public image benchmarks demonstrate state-of-the-art performance and faster inference than even non-diffusion baselines, and validate the effectiveness of S2R-Synthesis. Project page: https://codingwzp.github.io/VideoDereflection_S2R.