RenderFormer-V2: 이종 장면 프리미티브를 이용한 신경 렌더링
RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives
September 4, 2026
저자: Chong Zeng, Yue Dong, Pieter Peers, Lvmin Zhang, Maneesh Agrawala
cs.AI
초록
우리는 최신 물리 기반 렌더링 시스템을 보완하며, 장면별 학습이나 특수 코드 없이 코스틱스, 체적 산란, 환경 조명, 텍스처 및 변위 표면, 분포 외 재질과 같은 다양한 광 수송 효과를 처리할 수 있는 통합된 학습형 트랜스포머 기반 신경 렌더링 모델인 'RenderFormer-V2'를 제시한다. RenderFormer-V2는 전역 광 수송을 시퀀스-투-시퀀스 변환으로 모델링한다. 전작에 이어 RenderFormer-V2 역시 2단계 프로세스를 채택한다: 장면 내 프리미티브 간 수송을 해결하는 시점 독립 단계와 내부 신경 장면 표현을 이미지 픽셀로 변환하는 시점 종속 단계이다. RenderFormer와 달리, 우리 모델은 시점 독립 단계에서 새롭게 결합된 윈도우 어텐션과 렌더링 정보를 반영한 어텐션 싱크를 사용하여 렌더링 정확도를 유지하면서 확장성을 향상시킨다. 범용성을 더욱 향상시키기 위해 RenderFormer-V2는 환경 맵과 참여 매질을 포함한 이종 장면 프리미티브를 지원하며, 새로운 신경 임베딩을 통해 재질 외관을 인코딩하는, 기저 표면 반사 모델과 독립적인 재질 인코딩을 사용한다. 우리는 다양한 장면에서 RenderFormer-V2의 범용성을 입증하고 개선된 어텐션 메커니즘에 대한 광범위한 절제 실험을 수행한다.
English
We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, that can handle diverse light-transport effects such as caustics, volumetric scattering, environment lighting, textured and displaced surfaces and out-of-distribution materials without per-scene training or specialized code. RenderFormer-V2 models global light transport as a sequence-to-sequence transformation. Following its predecessor, RenderFormer-V2 also employs a two stage process: a view-independent stage that resolves intra-scene primitive to primitive transport, and a view-dependent stage that transforms the internal neural scene representation into image pixels. Different from RenderFormer, our model employs a novel combined windowed-attention and rendering-informed attention sink in the view-independent stage to improve scalability while maintaining render accuracy. To further improve versatility, RenderFormerV2 supports heterogeneous scene primitives, including environment maps and participating media, and it employs a material encoding independent of the underlying surface reflectance model that encodes material appearance via a novel neural embedding. We demonstrate the versatility of RenderFormer-V2 on a variety of scenes and perform an extensive ablation of the improved attention mechanism.