ChatPaper.aiChatPaper

MirrorWorld:駕馭影片擴散模型實現鏡像反射生成

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

August 7, 2026
作者: Youjun Zhao, Alex Warren, Gary K. L. Tam, Rynson W. H. Lau
cs.AI

摘要

近期影片擴散模型的進展已能實現高保真度的影片合成。然而,生成鏡面反射仍然具有挑戰性,因為鏡中的內容必須與周圍場景保持一致。現有的影片擴散模型並非專門針對場景與鏡面之間的關係進行建模,這可能導致反射內容錯誤或空間排列不一致的問題。我們觀察到,鏡面反射生成涉及兩個互補的挑戰:決定應反射哪些場景內容,以及反射內容應如何在鏡面區域內進行空間排列。基於此觀察,我們提出 MirrorWorld,一個具反射感知的影片修補框架,在生成過程中對場景與鏡面之間的關係進行建模。具體而言,我們引入語義關係蒸餾(SRD),從凍結的視覺基礎模型中遷移關係資訊,以促進可見場景內容與鏡面區域之間的語義關聯。我們進一步提出幾何變換對齊(GTA),學習一個引導反射內容空間排列的變換。這兩個元件扮演互補角色,SRD 負責建模應反射什麼內容,而 GTA 負責建模如何排列反射內容。為促進此問題的研究,我們透過將四個現有的影片鏡面資料集重新改造為統一的反射重建任務,建構了影片鏡面反射生成的基準。實驗結果顯示,MirrorWorld 在反射重建品質上優於具代表性的基於影像的反射生成方法,並勝過強大的影片修補基線方法。
English
Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging because the content within a mirror must remain consistent with the surrounding scene. Existing VDMs are not specifically designed to model scene-to-mirror relationships, which can lead to reflections with incorrect content or inconsistent spatial arrangements. We observe that mirror reflection generation involves two complementary challenges: determining what scene content should be reflected and how the reflected content should be spatially arranged within the mirror region. Motivated by this observation, we propose MirrorWorld, a reflection-aware video inpainting framework that models scene-to-mirror relationships during generation. Specifically, we introduce Semantic Relation Distillation (SRD), which transfers relational information from a frozen visual foundation model to encourage semantic associations between visible scene content and mirror regions. We further propose Geometric Transformation Alignment (GTA), which learns a transformation that guides the spatial arrangement of reflected content. The two components play complementary roles, with SRD modeling what should be reflected and GTA modeling how it should be arranged. To facilitate research on this problem, we construct a benchmark for video mirror reflection generation by repurposing four existing video mirror datasets into a unified reflection reconstruction task. Experimental results show that MirrorWorld achieves improved reflection reconstruction quality over representative image-based reflection generation methods and strong video inpainting baselines.