| --- |
| library_name: pytorch |
| license: mit |
| tags: |
| - foundation |
| - amd |
| - rocm |
| - zero-shot-image-classification |
| pipeline_tag: zero-shot-image-classification |
| --- |
| |
|  |
|
|
| # CLIP: Optimized for AMD ROCm |
|
|
| CLIP (Contrastive Language-Image Pre-training) performs zero-shot image classification by comparing image embeddings against text prompt embeddings. This repository packages evaluation/inference for zero-shot image classification using PyTorch and vLLM, exported and validated for **AMD ROCm** so it runs efficiently on AMD GPUs and CPUs. |
|
|
| This is based on the implementation of CLIP found [here](https://github.com/openai/CLIP). |
| This repository contains configurations and scripts optimized for **AMD® ROCm™** platforms. You can use the [AMD scripts](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/clip) to reproduce results or export with custom configurations. More details on model performance can be found [here](#accuracy-pipeline). |
|
|
| --- |
|
|
| ## Task Overview |
|
|
| **Task:** Zero-shot image classification |
|
|
| **Dataset:** CIFAR-10 test split (10,000 images, 10 classes) |
|
|
| **Output metrics:** Zero-shot accuracy (%), image throughput (img/s), text throughput (prompts/s) |
|
|
| > **Model variants:** Default is **base32** (`openai/clip-vit-base-patch32` / `ViT-B/32`). Override with `MODEL_VARIANT=base16|large14|large14-336`. |
| |
| > **vLLM note:** vLLM's CLIP backend embeds **one modality per request** — text prompts and images are sent in separate API calls, then cosine similarity is computed client-side (same approach as the original evaluation scripts). |
| |
| --- |
| |
| ## AMD ROCm Optimization |
| |
| This model export has been adapted and validated for **AMD Instinct™ / Radeon™ GPUs** running **ROCm**, as well as AMD CPUs. Key points: |
| |
| - Validated backends: **PyTorch** (native ROCm HIP kernels via `torch.cuda`) and **vLLM** (ROCm-enabled server build with pooling runner). |
| - No code changes required versus the upstream CLIP implementation — only environment/runtime configuration differs. |
| - CPU fallback path supported via a lightweight `venv_cpu` environment for environments without a ROCm-capable GPU. |
|
|
| | Runtime | Precision | Backend | Hardware | Notes | |
| |---|---|---|---|---| |
| | GPU | FP32, FP16 | PyTorch | AMD RYZEN AI MAX+ 395 w/ Radeon 8060S | ROCm PyTorch, native HIP kernels | |
| | GPU | FP32, FP16 | vLLM | AMD RYZEN AI MAX+ 395 w/ Radeon 8060S | ROCm-enabled server, pooling runner | |
| | CPU | FP32, FP16 | PyTorch | AMD RYZEN AI MAX+ 395 w/ Radeon 8060S | Lightweight `venv_cpu` environment | |
|
|
| --- |
|
|
| ## Getting Started |
|
|
| For setup instructions, evaluation scripts, and custom configuration options, see the [CLIP on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/clip). |
|
|
| --- |
|
|
| ## Model Details |
|
|
| **Model Type:** Zero-shot image classification (vision-language dual encoder) |
|
|
| **Base Model:** `openai/clip-vit-base-patch32` (CLIP ViT-B/32) |
|
|
| **Model Stats:** |
| - Model variant: base32 (default) — base16, large14, large14-336 also supported |
| - Number of parameters: `151M` |
| - Precision tested: FP32, FP16 |
|
|
| --- |
|
|
| ## Accuracy Pipeline |
|
|
| Higher zero-shot accuracy means more test images are assigned the correct CIFAR-10 class via CLIP's image–text similarity — 100% is perfect, 10% is chance level for 10 classes. Values above ~85% on CIFAR-10 with ViT-B/32 are typical for this benchmark. |
|
|
| ### Metrics Explained |
|
|
| | Metric | Description | |
| |--------|-------------| |
| | Zero-shot accuracy (%) | Fraction of CIFAR-10 test images whose highest-scoring text prompt matches the ground-truth label after softmax over 10 class prompts. Primary accuracy metric; sensitive to both image and text embedding quality. | |
|
|
| ### Accuracy Results |
|
|
| **Full Dataset Evaluation (CIFAR-10 test)** — filled from `evaluation_results/`; run `make metrics` to refresh: |
|
|
| <!-- accuracy-table-start --> |
| | Device | Backend | Precision | Variant | Accuracy (%) | |
| |--------|---------|-----------|---------|--------------| |
| | CPU | PyTorch | FP16 | clip-vit-base-patch32 | 88.79 | |
| | CPU | PyTorch | FP32 | clip-vit-base-patch32 | 88.80 | |
| | GPU | PyTorch | FP16 | clip-vit-base-patch32 | 88.75 | |
| | GPU | PyTorch | FP32 | clip-vit-base-patch32 | 88.80 | |
| | GPU | vLLM | FP16 | clip-vit-base-patch32 | 88.78 | |
| | GPU | vLLM | FP32 | clip-vit-base-patch32 | 88.80 | |
| <!-- accuracy-table-end --> |
|
|
| --- |
|
|
| ## Dig Deeper |
|
|
| Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples? |
|
|
| 📂 **[View the full project on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/clip)** |
|
|
| The GitHub repository includes: |
| - Setup and prerequisites for ROCm environments |
| - Scripts for the supported runners |
| - Additional model variants and datasets |
| - Benchmarking and reproduction instructions |
|
|