File size: 4,474 Bytes
96cce16
 
 
 
 
 
 
 
 
 
 
f86104b
96cce16
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
---
library_name: onnx
license: apache-2.0
tags:
  - foundation
  - amd
  - rocm
  - image-classification
pipeline_tag: image-classification
---

![](https://huggingface.co/AMD-PAVS-AI/efficientnetv2/resolve/main/efficientnetv2.png)

# EfficientNetV2: Optimized for AMD ROCm

EfficientNetV2 (`tf_efficientnetv2_s`) is a convolutional image classification model trained on ImageNet-1k. This repository packages inference for image classification using **ONNX Runtime** with a FastAPI serving layer, exported and validated for **AMD ROCm** so it runs efficiently on AMD GPUs, CPUs, and NPUs.

This is based on the implementation of EfficientNetV2 found [here](https://github.com/huggingface/pytorch-image-models).
This repository contains configurations and scripts optimized for **AMD® ROCm™** platforms. You can use the [efficientnetv2 AMD scripts](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/efficientnetv2) to reproduce results or export with custom configurations. More details on model performance can be found [here](#accuracy-pipeline).

---

## Task Overview

**Task:** ImageNet-1k classification (1000 classes), evaluated on the Imagenette 10-class validation subset

**Dataset:** Imagenette validation set (10 classes, mapped to ImageNet-1k 1000-way ordering)

**Output metrics:** Top-1 accuracy, Top-5 accuracy

> **NPU note:** VitisAI EP accepts FP32 ONNX and performs quantization internally to run on NPU hardware; expect small Top-1/Top-5 differences vs CPU/GPU FP32.

---

## AMD ROCm Optimization

This model export has been adapted and validated for **AMD Instinct™ / Radeon™ GPUs** running **ROCm**, as well as AMD CPUs and AMD Ryzen AI NPUs. Key points:

- Validated backends: **ONNX Runtime** across CPU (FP32 only), GPU (MIGraphX execution provider — FP32/FP16/BF16/INT8 via runtime quantization), and NPU (VitisAI execution provider, FP32 ONNX with internal quantization).
- All GPU precision benchmarks use a single FP32 ONNX model with runtime quantization via MIGraphX EP options — no pre-quantized model files required.
- No code changes required versus the upstream timm implementation — only environment/runtime configuration differs.

| Runtime | Precision | Backend | Hardware | Notes |
|---|---|---|---|---|
| ONNX Runtime | FP32 | CPU Execution Provider | AMD CPU | — |
| ONNX Runtime | FP32 / FP16 / BF16 / INT8 | MIGraphX Execution Provider | AMD Instinct™ / Radeon™ GPU (ROCm) | Runtime quantization, single FP32 ONNX model |
| ONNX Runtime | FP32 | VitisAI Execution Provider | AMD Ryzen AI NPU | VitisAI quantizes internally |

---

## Getting Started

For setup instructions, evaluation scripts, and custom configuration options, see the [efficientnetv2 on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/efficientnetv2).

---

## Model Details

**Model Type:** Image classification (convolutional neural network)

**Base Model:** `tf_efficientnetv2_s` (timm, ImageNet-1k pretrained)

**Model Stats:**
- Export size: 384×384 (native eval resolution for `tf_efficientnetv2_s`), overridable via `EFFNET_IMG_SIZE`
- ONNX model size: ~85 MB
- Precision tested: FP32 (CPU/NPU); FP32, FP16, BF16, INT8 (GPU)

---

## Accuracy Pipeline

Top-1 measures the fraction of images where the model's highest-confidence prediction matches the ground-truth class; Top-5 is looser, counting a match if the true class appears among the five highest-confidence predictions. Preprocessing matches the timm eval pipeline for `tf_efficientnetv2_s` (resize shorter side to 384, crop_pct=1.0, Inception mean/std normalization, bicubic interpolation).

### Metrics Explained

| Metric | Description |
|--------|-------------|
| Top-1 | Fraction of images where the model's highest-confidence prediction matches the ground-truth class. |
| Top-5 | Fraction of images where the ground-truth class appears among the model's five highest-confidence predictions. |

---

## Dig Deeper

Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?

📂 **[View the full project on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/efficientnetv2)**

The GitHub repository includes:
- Benchmark, profile, and evaluation Makefile targets for CPU, GPU, and NPU
- Runtime quantization details for MIGraphX (FP16/BF16/INT8)
- Imagenette dataset staging and full-validation evaluation pipeline
- ONNX export scripts producing a device-agnostic FP32 model