PaDiM: Optimized for AMD ROCm

PaDiM (Patch Distribution Modeling) models each spatial patch of a CNN backbone's feature map as a multivariate Gaussian fit only on defect-free training images, then flags anomalies via Mahalanobis distance to that patch's distribution at inference time — no anomalous training examples are needed. This repository packages training, export, and inference for anomaly detection and localization using PyTorch and ONNX Runtime, exported and validated for AMD ROCm so it runs efficiently on AMD GPUs, CPUs, and NPUs.

This is based on the implementation of PaDiM found here. This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the padim AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.


Task Overview

Task: Anomaly detection and localization

Dataset: MVTec AD (15 object/texture classes, ~5,354 images total; trains/evaluates on bottle by default)

Output metrics: Image AUROC, Pixel AUROC, optimal threshold, F1 (pixel-level)

Model variants: PaDiM has no MODEL_SIZE variants — configurable knobs are backbone architecture (resnet18 default / wide_resnet50_2) and covariance mode (diagonal default / full), both passed via ARGS.


AMD ROCm Optimization

This model export has been adapted and validated for AMD Instinct™ / Radeon™ GPUs running ROCm, as well as AMD CPUs and AMD Ryzen AI NPUs. Key points:

  • Validated backends: ONNX Runtime across CPU (FP32), GPU (MIGraphX execution provider — FP32/FP16/BF16/INT8), and NPU (VitisAI execution provider, auto-quantized internally).
  • Unlike models that ship pretrained weights, PaDiM must be trained (fitting Gaussian parameters per class) before export, benchmark, profile, or eval can run.
  • GPU and NPU targets carry a first-run compilation/tuning cost (MIGraphX kernel tuning, VitisAI graph compilation) that can take 30+ minutes; subsequent runs are faster.
  • NPU inference is auto-quantized internally by VitisAI — expect some AUROC drop relative to CPU/GPU FP32.
Runtime Precision Backend Hardware Notes
ONNX Runtime FP32 CPU Execution Provider AMD CPU
ONNX Runtime FP32 / FP16 / BF16 / INT8 MIGraphX Execution Provider AMD Instinct™ / Radeon™ GPU (ROCm) First-run kernel tuning can take 30+ minutes
ONNX Runtime Auto VitisAI Execution Provider AMD Ryzen AI NPU Auto-quantized internally

Getting Started

For setup instructions, evaluation scripts, and custom configuration options, see the padim on GitHub.


Model Details

Model Type: Anomaly detection and localization (patch distribution modeling over a CNN backbone)

Base Model: ResNet-18 backbone (default) — Wide ResNet-50-2 also supported

Model Stats:

  • Backbone: resnet18 (default) — wide_resnet50_2 also supported
  • Covariance mode: diagonal (default, fastest) — full covariance also supported (~1-2% better AUROC, ~55x slower Mahalanobis)
  • ONNX export levels: backbone (11 MB), full (11 MB, recommended), base/mahalanobis (132 MB)
  • Default trained/evaluated class: bottle (of 15 MVTec AD classes)

Accuracy Pipeline

Higher AUROC means the model ranks anomalous samples above normal ones more consistently — 1.0 is a perfect ranking, 0.5 is random chance. The PaDiM paper reports ~96.7% Image AUROC / 96.0% Pixel AUROC with ResNet18, and ~97.5%/97.5% with Wide ResNet-50-2, on MVTec AD; on-device numbers noticeably below that suggest a training/data issue or (for NPU) quantization-induced accuracy loss.

Metrics Explained

Metric Description
Image AUROC Probability that a randomly chosen anomalous image scores higher (via its max per-image patch anomaly score) than a randomly chosen normal image — measures whole-image anomaly classification skill, independent of any threshold choice.
Pixel AUROC Same ranking measure computed pixel-by-pixel against the ground-truth defect masks — captures localization quality.
Optimal threshold The pixel anomaly score cutoff that maximizes pixel-level F1 on the test set. Used only to binarize the heatmap for saved visualizations.
F1 (pixel-level) Harmonic mean of pixel precision and recall at the optimal threshold — a single fixed-operating-point score, unlike AUROC which integrates over every threshold.

No measured on-device results are included in the source README yet — by default only the bottle class is trained/evaluated; additional classes can be trained and evaluated with make train-cpu ARGS="--classes <class1> <class2> ...".


Dig Deeper

Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?

📂 View the full project on GitHub

The GitHub repository includes:

  • Training and ONNX export scripts (four export levels trading off latency vs. pipeline coverage)
  • MVTec AD dataset staging across all 15 classes
  • Image/Pixel AUROC + F1 evaluation pipeline with anomaly visualization overlays
  • Benchmarking and reproduction instructions for CPU, GPU, and NPU
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for AMD-PAVS-AI/padim