RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives
Abstract
RenderFormer-V2 is a transformer-based neural rendering model that handles diverse light-transport effects via a two-stage sequence-to-sequence architecture with improved attention and heterogeneous scene support.
We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, that can handle diverse light-transport effects such as caustics, volumetric scattering, environment lighting, textured and displaced surfaces and out-of-distribution materials without per-scene training or specialized code. RenderFormer-V2 models global light transport as a sequence-to-sequence transformation. Following its predecessor, RenderFormer-V2 also employs a two stage process: a view-independent stage that resolves intra-scene primitive to primitive transport, and a view-dependent stage that transforms the internal neural scene representation into image pixels. Different from RenderFormer, our model employs a novel combined windowed-attention and rendering-informed attention sink in the view-independent stage to improve scalability while maintaining render accuracy. To further improve versatility, RenderFormerV2 supports heterogeneous scene primitives, including environment maps and participating media, and it employs a material encoding independent of the underlying surface reflectance model that encodes material appearance via a novel neural embedding. We demonstrate the versatility of RenderFormer-V2 on a variety of scenes and perform an extensive ablation of the improved attention mechanism.
Community
Project Page: https://renderformer.github.io/v2/
A pretrained transformer that turns a sequence of mixed scene primitives into a globally illuminated image.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- LumiTokens: 3D Relighting via Token-Space Lighting Transformation (2026)
- FF-ProCams: Feed-Forward Gaussian Splatting for Projector-Camera System (2026)
- Luce: Relightable Gaussians for 3D Asset Generation (2026)
- Sparse auto-regressive modeling for scene generation from multi-view images (2026)
- LightBridge: Feed-Forward Generative Relighting for 3D Gaussian Splatting (2026)
- RGBX-Next: Towards Realistic Generative Rendering from G-Buffers (2026)
- USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.05738 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper