**Explainability** | Field | Response | | :---- | :---- | | Intended Task/Domain: | Multi-view 3D scene reconstruction. | | Model Type: | Transformer | | Intended Users: | 3D vision, simulation, graphics, and robotics or physical AI researchers and developers. | | Output | 3D Gaussian Splat representation and rendered novel views. | | Describe how the model works | Encoder-decoder Transformer with learnable Gaussian tokens directly regresses 3D Gaussian attributes from posed images, trained with rendering and visibility losses. | | Name the adversely impacted groups this has been tested to deliver comparable outcomes regardless of | Not Applicable | | Technical Limitations & Mitigation | TokenGS may miss fine-grained geometric details. Quality depends on camera pose quality and multiview coverage, so users should validate outputs and provide sufficient view diversity and accurate camera metadata. | | Verified to have met prescribed NVIDIA quality standards | Yes | | Performance Metrics | PSNR, SSIM, LPIPS; additional comparisons under view extrapolation and camera-noise robustness. | | Potential Known Risks | Reconstruction failures or incomplete geometry may produce misleading renderings or assets. | | Licensing | The use of the model is governed by the [NVIDIA Internal Scientific Research and Development Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-internal-scientific-research-and-development-model-license/). |