EfficientNetV2: Optimized for AMD ROCm
EfficientNetV2 (tf_efficientnetv2_s) is a convolutional image classification model trained on ImageNet-1k. This repository packages inference for image classification using ONNX Runtime with a FastAPI serving layer, exported and validated for AMD ROCm so it runs efficiently on AMD GPUs, CPUs, and NPUs.
This is based on the implementation of EfficientNetV2 found here. This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the efficientnetv2 AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.
Task Overview
Task: ImageNet-1k classification (1000 classes), evaluated on the Imagenette 10-class validation subset
Dataset: Imagenette validation set (10 classes, mapped to ImageNet-1k 1000-way ordering)
Output metrics: Top-1 accuracy, Top-5 accuracy
NPU note: VitisAI EP accepts FP32 ONNX and performs quantization internally to run on NPU hardware; expect small Top-1/Top-5 differences vs CPU/GPU FP32.
AMD ROCm Optimization
This model export has been adapted and validated for AMD Instinct™ / Radeon™ GPUs running ROCm, as well as AMD CPUs and AMD Ryzen AI NPUs. Key points:
- Validated backends: ONNX Runtime across CPU (FP32 only), GPU (MIGraphX execution provider — FP32/FP16/BF16/INT8 via runtime quantization), and NPU (VitisAI execution provider, FP32 ONNX with internal quantization).
- All GPU precision benchmarks use a single FP32 ONNX model with runtime quantization via MIGraphX EP options — no pre-quantized model files required.
- No code changes required versus the upstream timm implementation — only environment/runtime configuration differs.
| Runtime | Precision | Backend | Hardware | Notes |
|---|---|---|---|---|
| ONNX Runtime | FP32 | CPU Execution Provider | AMD CPU | — |
| ONNX Runtime | FP32 / FP16 / BF16 / INT8 | MIGraphX Execution Provider | AMD Instinct™ / Radeon™ GPU (ROCm) | Runtime quantization, single FP32 ONNX model |
| ONNX Runtime | FP32 | VitisAI Execution Provider | AMD Ryzen AI NPU | VitisAI quantizes internally |
Getting Started
For setup instructions, evaluation scripts, and custom configuration options, see the efficientnetv2 on GitHub.
Model Details
Model Type: Image classification (convolutional neural network)
Base Model: tf_efficientnetv2_s (timm, ImageNet-1k pretrained)
Model Stats:
- Export size: 384×384 (native eval resolution for
tf_efficientnetv2_s), overridable viaEFFNET_IMG_SIZE - ONNX model size: ~85 MB
- Precision tested: FP32 (CPU/NPU); FP32, FP16, BF16, INT8 (GPU)
Accuracy Pipeline
Top-1 measures the fraction of images where the model's highest-confidence prediction matches the ground-truth class; Top-5 is looser, counting a match if the true class appears among the five highest-confidence predictions. Preprocessing matches the timm eval pipeline for tf_efficientnetv2_s (resize shorter side to 384, crop_pct=1.0, Inception mean/std normalization, bicubic interpolation).
Metrics Explained
| Metric | Description |
|---|---|
| Top-1 | Fraction of images where the model's highest-confidence prediction matches the ground-truth class. |
| Top-5 | Fraction of images where the ground-truth class appears among the model's five highest-confidence predictions. |
Dig Deeper
Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?
📂 View the full project on GitHub
The GitHub repository includes:
- Benchmark, profile, and evaluation Makefile targets for CPU, GPU, and NPU
- Runtime quantization details for MIGraphX (FP16/BF16/INT8)
- Imagenette dataset staging and full-validation evaluation pipeline
- ONNX export scripts producing a device-agnostic FP32 model
