Instructions to use kornia/dinov3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- TensorRT
How to use kornia/dinov3 with TensorRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
| license: other | |
| license_name: dinov3-license | |
| license_link: LICENSE.md | |
| base_model: facebook/dinov3-vits16-pretrain-lvd1689m | |
| tags: | |
| - dinov3 | |
| - onnx | |
| - tensorrt | |
| - jetson | |
| - image-feature-extraction | |
| library_name: vision-rt | |
| pipeline_tag: image-feature-extraction | |
| # kornia/dinov3 | |
| ONNX export of **DINOv3 ViT-S/16** for [`vision-rt`](https://github.com/kornia/vision-rt), | |
| producing a global image descriptor (the CLS token) for retrieval, visual place | |
| recognition and scene-change detection on NVIDIA Jetson. | |
| **All model credit belongs to Meta AI** (Siméoni et al.). This repository contains only a | |
| format conversion of | |
| [`facebook/dinov3-vits16-pretrain-lvd1689m`](https://huggingface.co/facebook/dinov3-vits16-pretrain-lvd1689m) | |
| — no retraining, no fine-tuning, no architectural change. Use of these files is governed | |
| by the **DINOv3 License**, reproduced in full in [`LICENSE.md`](LICENSE.md); by | |
| downloading them you agree to its terms, including its restrictions on military, | |
| weapons-related and other prohibited uses. | |
| ## Files | |
| | File | Notes | | |
| |------|-------| | |
| | `dinov3-vits16-336.onnx` | Graph only, ~1.2 MB | | |
| | `dinov3-vits16-336.onnx.data` | **Required** external weights, ~86 MB | | |
| The `.onnx.data` sidecar is mandatory — torch's exporter externalizes weights regardless | |
| of the 2 GB protobuf limit, so the `.onnx` alone contains no weights. Keep the two files | |
| in the same directory; the ONNX parser resolves the sidecar next to the graph. | |
| ## Model I/O | |
| Input `input` `[1, 3, 336, 336]`, ImageNet-normalized CHW (mean `[0.485, 0.456, 0.406]`, | |
| std `[0.229, 0.224, 0.225]`), 336 = 21 × 16. | |
| | Output | Shape | Meaning | | |
| |--------|-------|---------| | |
| | `descriptor` | `[1, 384]` | CLS token (`pooler_output`) — the global descriptor | | |
| | `tokens` | `[1, 446, 384]` | Full sequence: 1 CLS + 4 registers + 441 patches (21×21) | | |
| Verified at export: `cos(pooler_output, last_hidden_state[:,0]) = 1.000000`, so the | |
| descriptor *is* the CLS token. Note the 4 **register tokens** between CLS and the patch | |
| grid — any hand-rolled pooling must skip indices 1–4 or it averages in artifact tokens. | |
| Exported with opset 18 from torch 2.11 / transformers 4.57.6 by | |
| [`crates/vrt-dinov3/scripts/export_dinov3.py`](https://github.com/kornia/vision-rt). | |
| ## ⚠️ Build BF16, not FP16 | |
| **FP16 produces all-NaN for this model.** Measured on Jetson Orin (SM87, TensorRT | |
| 10.3.0.30), descriptor cosine against the PyTorch reference: | |
| | Precision | GPU compute | Cosine vs PyTorch | | | |
| |-----------|------------:|------------------:|--| | |
| | **bf16** | **7.6 ms** | **0.999568** | ✅ recommended | | |
| | fp32 | 13.4 ms | 0.999998 | ✅ | | |
| | fp16 | 5.2 ms | all NaN | ❌ | | |
| | fp16 + bf16 | 5.2 ms | all NaN | ❌ TRT picks fp16 for attention | | |
| Cause: exactly one tensor leaves fp16 range — the attention logits `Q·Kᵀ`, max |value| | |
| **2.1e6** (2.5e6 on a natural image) against fp16's 65,504 ceiling. They overflow to `inf` | |
| in the first block, `softmax(inf − inf)` is NaN, and it reaches every output. Every other | |
| tensor stays under 1800. These are already-scaled logits — DINOv3 genuinely produces | |
| massive attention values (the "attention sink" its register tokens absorb). It is a | |
| dynamic-**range** problem, so bf16 fixes it by keeping fp32's exponent. | |
| **Do not try to pin fp16 with `--layerPrecisions`.** TensorRT fuses the attention block | |
| into a myelin kernel, so ONNX node names are absent from the engine; both wildcard and | |
| explicit pins match nothing and are silently ignored — the build succeeds, runs at | |
| identical speed, and still emits NaN. | |
| ```bash | |
| trtexec --onnx=dinov3-vits16-336.onnx --saveEngine=dinov3.engine --bf16 \ | |
| --memPoolSize=workspace:2048 | |
| ``` | |
| ## Usage | |
| ```rust | |
| let dino = DinoV3::from_hub(stream.clone())?; // or from_engine_file / from_onnx | |
| let mut r = dino.alloc_result()?; | |
| dino.submit(&img, &mut r)?; // enqueue, no sync | |
| stream.synchronize()?; | |
| let d = r.descriptor_host()?; // L2-normed [384], cosine-ready | |
| ``` | |
| Measured separation (real images): two views of one place score **~0.96**; unrelated | |
| scenes **−0.001 … 0.11**. Live on a 1280×720 RTSP camera: 15 fps, source-gated, with the | |
| GPU using 6.8 ms of the 66 ms frame interval. | |
| ## Citation | |
| ```bibtex | |
| @article{simeoni2025dinov3, | |
| title={DINOv3}, | |
| author={Sim{\'e}oni, Oriane and others}, | |
| journal={arXiv preprint arXiv:2508.10104}, | |
| year={2025} | |
| } | |
| ``` | |