ChatPaper.aiChatPaper

MirrorWorld: 비디오 확산 모델 제어를 통한 거울 반사 생성

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

August 7, 2026
저자: Youjun Zhao, Alex Warren, Gary K. L. Tam, Rynson W. H. Lau
cs.AI

초록

최근 비디오 확산 모델(VDM)의 발전으로 고충실도 비디오 합성이 가능해졌다. 그러나 거울 내부의 콘텐츠가 주변 장면과 일관성을 유지해야 하므로 거울 반사 생성은 여전히 어렵다. 기존 VDM은 장면-거울 관계를 모델링하도록 특별히 설계되지 않았기 때문에, 잘못된 콘텐츠나 일관되지 않은 공간 배치를 가진 반사를 생성할 수 있다. 우리는 거울 반사 생성이 두 가지 보완적 과제를 수반한다는 점을 관찰한다: 어떤 장면 콘텐츠가 반사되어야 하는지 결정하는 것과, 반사된 콘텐츠가 거울 영역 내에서 어떻게 공간적으로 배치되어야 하는지 결정하는 것이다. 이러한 관찰에 기반하여, 우리는 생성 과정에서 장면-거울 관계를 모델링하는 반사 인지 비디오 인페인팅 프레임워크인 MirrorWorld를 제안한다. 구체적으로, 우리는 의미 관계 증류(SRD)를 도입하여 고정된 시각 기반 모델로부터 관계 정보를 전이함으로써 가시적 장면 콘텐츠와 거울 영역 간의 의미 연관성을 촉진한다. 또한 반사된 콘텐츠의 공간 배치를 안내하는 변환을 학습하는 기하 변환 정렬(GTA)을 제안한다. 이 두 구성 요소는 보완적 역할을 수행하며, SRD는 무엇이 반사되어야 하는지를, GTA는 어떻게 배치되어야 하는지를 모델링한다. 이 문제에 대한 연구를 촉진하기 위해, 우리는 기존의 네 개의 비디오 거울 데이터셋을 통합 반사 재구성 과제로 재구성하여 비디오 거울 반사 생성을 위한 벤치마크를 구축한다. 실험 결과는 MirrorWorld가 대표적인 이미지 기반 반사 생성 방법들과 강력한 비디오 인페인팅 기준 모델들보다 향상된 반사 재구성 품질을 달성함을 보여준다.
English
Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging because the content within a mirror must remain consistent with the surrounding scene. Existing VDMs are not specifically designed to model scene-to-mirror relationships, which can lead to reflections with incorrect content or inconsistent spatial arrangements. We observe that mirror reflection generation involves two complementary challenges: determining what scene content should be reflected and how the reflected content should be spatially arranged within the mirror region. Motivated by this observation, we propose MirrorWorld, a reflection-aware video inpainting framework that models scene-to-mirror relationships during generation. Specifically, we introduce Semantic Relation Distillation (SRD), which transfers relational information from a frozen visual foundation model to encourage semantic associations between visible scene content and mirror regions. We further propose Geometric Transformation Alignment (GTA), which learns a transformation that guides the spatial arrangement of reflected content. The two components play complementary roles, with SRD modeling what should be reflected and GTA modeling how it should be arranged. To facilitate research on this problem, we construct a benchmark for video mirror reflection generation by repurposing four existing video mirror datasets into a unified reflection reconstruction task. Experimental results show that MirrorWorld achieves improved reflection reconstruction quality over representative image-based reflection generation methods and strong video inpainting baselines.