RelightFormer:用於多視角物件重照明的前饋生成式 Transformer
RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting
September 7, 2026
作者: Hejun Wang, Jinxi Li, Junwei Jiang, Shiwei Mao, Hu Cheng, Shouwang Huang, Bo Yang
cs.AI
摘要
影像重照明傳統上是透過複雜的逆向渲染管線來處理,而這些管線會遭遇不適定最佳化問題;或是透過單張影像生成模型來處理,而這類模型忽略了理解 3D 幾何與材質交互作用所需的關鍵多視角線索。為了解決這些限制,我們提出一個前饋式生成 Transformer,用於直接的單視角與多視角影像重照明,且完全繞過顯式內在屬性估計。我們的架構改編自視訊基礎模型,具備一個潛在照明模組,可透過交叉注意力將目標環境圖動態注入空間特徵中。此外,我們採用置換不變的位置編碼,以對稱地處理無序的多視角輸入,而不產生序列偏誤。為了訓練這個穩健的資料驅動模型,我們建構了大規模的 Laval Objaverse 資料集(LOD),其中包含 9 萬個物件與 3.9 萬種獨特照明。廣泛的實驗證明其具備最先進的視覺品質、照片擬真的重照明品質,以及在單視角、多視角與新視角重照明任務上強大的零樣本泛化能力。
English
Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.