dinov2 / README.md
AMD-PAVS-AI's picture
Update image references to HF URLs
25cc419 verified
|
Raw
History Blame Contribute Delete
4.46 kB
---
library_name: onnx
license: apache-2.0
tags:
- foundation
- amd
- rocm
- image-classification
pipeline_tag: image-classification
---
![](https://huggingface.co/AMD-PAVS-AI/dinov2/resolve/main/dinov2.png)
# DINOv2: Optimized for AMD ROCm
DINOv2 is a self-supervised Vision Transformer evaluated here as an ImageNet linear classifier (ViT-S/B/L/G with a 1000-class head). This repository packages evaluation and inference for image classification using **ONNX Runtime**, exported and validated for **AMD ROCm** (via the MIGraphX execution provider) so it runs efficiently on AMD GPUs, CPUs, and NPUs.
This is based on the implementation of DINOv2 found [here](https://github.com/facebookresearch/dinov2).
This repository contains configurations and scripts optimized for **AMD® ROCm™** platforms. You can use the [dinov2 AMD scripts](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/dinov2) to reproduce results or export with custom configurations. More details on model performance can be found [here](#accuracy-pipeline).
---
## Task Overview
**Task:** Image classification
**Dataset:** ImageNet validation (~50,000 images, 1,000 classes)
**Output metrics:** Top-1 accuracy, Top-5 accuracy
> **Model variants:** Default is **base** (`dinov2_vitb14_lc`). Override with `MODEL_SIZE=small|base|large|giant` (hub names: `dinov2_vits14_lc`, `dinov2_vitb14_lc`, `dinov2_vitl14_lc`, `dinov2_vitg14_lc`).
---
## AMD ROCm Optimization
This model export has been adapted and validated for **AMD Instinct™ / Radeon™ GPUs** running **ROCm**, as well as AMD CPUs. Key points:
- Validated backends: **ONNX Runtime** across **CPU** (native CPU execution provider), **GPU** (MIGraphX execution provider, ROCm-based), and **NPU** (VitisAI execution provider).
- CPU execution path supported independently of GPU/NPU availability, for environments without ROCm-capable GPU or NPU hardware.
| Runtime | Precision | Backend | Hardware | Notes |
|---|---|---|---|---|
| ONNX Runtime | FP32, FP16, BF16, INT8 | CPU Execution Provider | AMD CPU | INT8 available via runtime auto-quantization flags |
| ONNX Runtime | FP32, FP16, BF16, INT8 | MIGraphX Execution Provider (ROCm) | AMD GPU | First run may take longer due to graph compilation |
| ONNX Runtime | FP32, FP16, BF16, INT8 | VitisAI Execution Provider | AMD NPU (Ryzen AI) | INT8 requires AMD Quark calibration (`make quantize-npu-int8`) |
---
## Getting Started
For setup instructions, evaluation scripts, and custom configuration options, see the [dinov2 on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/dinov2).
---
## Model Details
**Model Type:** Self-supervised Vision Transformer (ViT) image classifier with a linear-probe head
**Base Model:** `facebookresearch/dinov2` (DINOv2 ViT-B/14 with linear classification head)
**Model Stats:**
- Model variant: base (default) — small, large, giant also supported
- Precision tested: FP32, FP16, BF16, INT8
---
## Accuracy Pipeline
Higher Top-1 accuracy means a larger fraction of images are classified with the correct label as the top prediction — 100% is perfect, 0% is chance-level for random guessing. Top-5 allows credit when the true class appears anywhere in the model's five highest-scoring labels; values above ~80% Top-1 on ImageNet val are considered strong for this linear-probe setup.
### Metrics Explained
| Metric | Description |
|--------|-------------|
| Top-1 | Primary classification metric — the fraction of images where the highest-scoring class matches ground truth. Strictest single-label score; a wrong top prediction counts as a full miss even if the true class ranked second. |
| Top-5 | Fraction of images where the true class appears in the model's top five predictions. Looser than Top-1 and typically higher; useful when near-miss rankings still indicate the model recognized the object category. |
---
## Dig Deeper
Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?
📂 **[View the full project on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/dinov2)**
The GitHub repository includes:
- Makefile-driven benchmark, profile, and evaluation targets for CPU, GPU, and NPU
- Per-device precision configs (FP32/FP16/BF16/INT8) and ONNX export scripts
- ImageNet dataset staging and full-validation evaluation pipeline
- Additional DINOv2 model variants (small, large, giant)