MirrorWorld: 鏡面反射生成のためのビデオ拡散モデルの制御
MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation
August 7, 2026
著者: Youjun Zhao, Alex Warren, Gary K. L. Tam, Rynson W. H. Lau
cs.AI
要旨
近年のビデオ拡散モデル(VDM)の進歩により、高忠実度のビデオ合成が可能になった。しかし、鏡の反射の生成は、鏡内のコンテンツが周囲のシーンと一貫性を保つ必要があるため、依然として困難である。既存のVDMはシーンと鏡の関係をモデル化するように特別に設計されていないため、誤った内容や空間配置の不整合を伴う反射が生成される可能性がある。我々は、鏡の反射生成には、どのシーンコンテンツを反射するかを決定することと、反射コンテンツを鏡領域内でどのように空間配置するかという、2つの相補的な課題が含まれることを観察した。この観察に基づき、我々は生成中にシーンと鏡の関係をモデル化する反射認識型ビデオインペインティングフレームワークであるMirrorWorldを提案する。
具体的には、意味的関係の蒸留(SRD)を導入する。これは、凍結された視覚基盤モデルから関係情報を転送し、可視シーンコンテンツと鏡領域の間の意味的関連付けを促進する。さらに、幾何学的変換の整合(GTA)を提案する。これは、反射コンテンツの空間配置を導く変換を学習する。これら2つの要素は補完的な役割を果たし、SRDは「何を反射すべきか」をモデル化し、GTAは「どのように配置すべきか」をモデル化する。この問題に関する研究を促進するため、既存の4つのビデオ鏡データセットを統一された反射再構成タスクに転用することで、ビデオ鏡反射生成のためのベンチマークを構築する。実験結果は、MirrorWorldが代表的な画像ベースの反射生成手法や強力なビデオインペインティングベースラインよりも優れた反射再構成品質を達成することを示している。
English
Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging because the content within a mirror must remain consistent with the surrounding scene. Existing VDMs are not specifically designed to model scene-to-mirror relationships, which can lead to reflections with incorrect content or inconsistent spatial arrangements. We observe that mirror reflection generation involves two complementary challenges: determining what scene content should be reflected and how the reflected content should be spatially arranged within the mirror region. Motivated by this observation, we propose MirrorWorld, a reflection-aware video inpainting framework that models scene-to-mirror relationships during generation. Specifically, we introduce Semantic Relation Distillation (SRD), which transfers relational information from a frozen visual foundation model to encourage semantic associations between visible scene content and mirror regions. We further propose Geometric Transformation Alignment (GTA), which learns a transformation that guides the spatial arrangement of reflected content. The two components play complementary roles, with SRD modeling what should be reflected and GTA modeling how it should be arranged. To facilitate research on this problem, we construct a benchmark for video mirror reflection generation by repurposing four existing video mirror datasets into a unified reflection reconstruction task. Experimental results show that MirrorWorld achieves improved reflection reconstruction quality over representative image-based reflection generation methods and strong video inpainting baselines.