RelightFormer:多視点物体リライティングのためのフィードフォワード生成トランスフォーマー
RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting
September 7, 2026
著者: Hejun Wang, Jinxi Li, Junwei Jiang, Shiwei Mao, Hu Cheng, Shouwang Huang, Bo Yang
cs.AI
要旨
画像リライティングは伝統的に、不良設定最適化に悩まされる複雑な逆レンダリングパイプライン、あるいは3D幾何と材質の相互作用を理解するために必要な重要な多視点手がかりを無視する単一画像生成モデルによって取り組まれてきた。これらの限界に対処するため、我々は明示的な固有特性推定を完全に回避する、単一視点および多視点画像の直接リライティングのためのフィードフォワード生成Transformerを導入する。動画基盤モデルから適応した我々のアーキテクチャは、クロスアテンションを介してターゲット環境マップを空間特徴に動的に注入する潜在照明モジュールを備える。さらに、置換不変な位置エンコーディングを採用し、順序を持たない多視点入力を系列バイアスなしに対称に処理する。この頑健なデータ駆動モデルを訓練するために、9万個の物体と3.9万個の一意な照明からなる大規模Laval Objaverse Dataset (LOD)を構築する。広範な実験により、単一視点、多視点、新規視点リライティングタスクにわたる最先端の視覚品質、写実的なリライティング品質、強いゼロショット汎化が示される。
English
Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.