CLIP: Optimized for AMD ROCm
CLIP (Contrastive Language-Image Pre-training) performs zero-shot image classification by comparing image embeddings against text prompt embeddings. This repository packages evaluation/inference for zero-shot image classification using PyTorch and vLLM, exported and validated for AMD ROCm so it runs efficiently on AMD GPUs and CPUs.
This is based on the implementation of CLIP found here. This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.
Task Overview
Task: Zero-shot image classification
Dataset: CIFAR-10 test split (10,000 images, 10 classes)
Output metrics: Zero-shot accuracy (%), image throughput (img/s), text throughput (prompts/s)
Model variants: Default is base32 (
openai/clip-vit-base-patch32/ViT-B/32). Override withMODEL_VARIANT=base16|large14|large14-336.
vLLM note: vLLM's CLIP backend embeds one modality per request — text prompts and images are sent in separate API calls, then cosine similarity is computed client-side (same approach as the original evaluation scripts).
AMD ROCm Optimization
This model export has been adapted and validated for AMD Instinct™ / Radeon™ GPUs running ROCm, as well as AMD CPUs. Key points:
- Validated backends: PyTorch (native ROCm HIP kernels via
torch.cuda) and vLLM (ROCm-enabled server build with pooling runner). - No code changes required versus the upstream CLIP implementation — only environment/runtime configuration differs.
- CPU fallback path supported via a lightweight
venv_cpuenvironment for environments without a ROCm-capable GPU.
| Runtime | Precision | Backend | Hardware | Notes |
|---|---|---|---|---|
| GPU | FP32, FP16 | PyTorch | AMD RYZEN AI MAX+ 395 w/ Radeon 8060S | ROCm PyTorch, native HIP kernels |
| GPU | FP32, FP16 | vLLM | AMD RYZEN AI MAX+ 395 w/ Radeon 8060S | ROCm-enabled server, pooling runner |
| CPU | FP32, FP16 | PyTorch | AMD RYZEN AI MAX+ 395 w/ Radeon 8060S | Lightweight venv_cpu environment |
Getting Started
For setup instructions, evaluation scripts, and custom configuration options, see the CLIP on GitHub.
Model Details
Model Type: Zero-shot image classification (vision-language dual encoder)
Base Model: openai/clip-vit-base-patch32 (CLIP ViT-B/32)
Model Stats:
- Model variant: base32 (default) — base16, large14, large14-336 also supported
- Number of parameters:
151M - Precision tested: FP32, FP16
Accuracy Pipeline
Higher zero-shot accuracy means more test images are assigned the correct CIFAR-10 class via CLIP's image–text similarity — 100% is perfect, 10% is chance level for 10 classes. Values above ~85% on CIFAR-10 with ViT-B/32 are typical for this benchmark.
Metrics Explained
| Metric | Description |
|---|---|
| Zero-shot accuracy (%) | Fraction of CIFAR-10 test images whose highest-scoring text prompt matches the ground-truth label after softmax over 10 class prompts. Primary accuracy metric; sensitive to both image and text embedding quality. |
Accuracy Results
Full Dataset Evaluation (CIFAR-10 test) — filled from evaluation_results/; run make metrics to refresh:
| Device | Backend | Precision | Variant | Accuracy (%) |
|---|---|---|---|---|
| CPU | PyTorch | FP16 | clip-vit-base-patch32 | 88.79 |
| CPU | PyTorch | FP32 | clip-vit-base-patch32 | 88.80 |
| GPU | PyTorch | FP16 | clip-vit-base-patch32 | 88.75 |
| GPU | PyTorch | FP32 | clip-vit-base-patch32 | 88.80 |
| GPU | vLLM | FP16 | clip-vit-base-patch32 | 88.78 |
| GPU | vLLM | FP32 | clip-vit-base-patch32 | 88.80 |
Dig Deeper
Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?
📂 View the full project on GitHub
The GitHub repository includes:
- Setup and prerequisites for ROCm environments
- Scripts for the supported runners
- Additional model variants and datasets
- Benchmarking and reproduction instructions
