RelightFormer: 다시점 객체 재조명을 위한 피드포워드 생성 트랜스포머
RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting
September 7, 2026
저자: Hejun Wang, Jinxi Li, Junwei Jiang, Shiwei Mao, Hu Cheng, Shouwang Huang, Bo Yang
cs.AI
초록
이미지 재조명은 전통적으로 ill-posed 최적화 문제를 겪는 복잡한 역렌더링 파이프라인이나, 3D 기하 및 재질 상호작용을 이해하는 데 필요한 중요한 다시점 단서를 무시하는 단일 이미지 생성 모델을 통해 다루어져 왔다. 이러한 한계를 해결하기 위해, 우리는 명시적인 고유 속성 추정을 완전히 우회하는 직접적인 단일 및 다시점 이미지 재조명을 위한 피드포워드 생성 트랜스포머를 제안한다. 비디오 파운데이션 모델에서 적응된 우리 아키텍처는 교차 어텐션을 통해 대상 환경 맵을 공간 특징에 동적으로 주입하는 잠재 조명 모듈을 갖추고 있다. 또한 우리는 순차적 편향 없이 순서가 없는 다시점 입력을 대칭적으로 처리하기 위해 순열 불변 위치 인코딩을 사용한다. 이 견고한 데이터 기반 모델을 훈련하기 위해, 우리는 90K 개의 객체와 39K 개의 고유 조명으로 구성된 대규모 Laval Objaverse Dataset(LOD)을 구축한다. 광범위한 실험은 단일 시점, 다시점 및 새로운 시점 재조명 작업 전반에 걸쳐 최첨단 시각적 품질, 사실적인 재조명 품질, 강력한 제로샷 일반화를 입증한다.
English
Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.