Explainability
| Field | Response |
|---|---|
| Intended Task/Domain: | Multi-view 3D scene reconstruction. |
| Model Type: | Transformer |
| Intended Users: | 3D vision, simulation, graphics, and robotics or physical AI researchers and developers. |
| Output | 3D Gaussian Splat representation and rendered novel views. |
| Describe how the model works | Encoder-decoder Transformer with learnable Gaussian tokens directly regresses 3D Gaussian attributes from posed images, trained with rendering and visibility losses. |
| Name the adversely impacted groups this has been tested to deliver comparable outcomes regardless of | Not Applicable |
| Technical Limitations & Mitigation | TokenGS may miss fine-grained geometric details. Quality depends on camera pose quality and multiview coverage, so users should validate outputs and provide sufficient view diversity and accurate camera metadata. |
| Verified to have met prescribed NVIDIA quality standards | Yes |
| Performance Metrics | PSNR, SSIM, LPIPS; additional comparisons under view extrapolation and camera-noise robustness. |
| Potential Known Risks | Reconstruction failures or incomplete geometry may produce misleading renderings or assets. |
| Licensing | The use of the model is governed by the NVIDIA Internal Scientific Research and Development Model License. |