CLIP / README.md
AMD-PAVS-AI's picture
Update image references to HF URLs
b8a3961 verified
|
Raw
History Blame Contribute Delete
4.78 kB
---
library_name: pytorch
license: mit
tags:
- foundation
- amd
- rocm
- zero-shot-image-classification
pipeline_tag: zero-shot-image-classification
---
![](https://huggingface.co/AMD-PAVS-AI/CLIP/resolve/main/clip.png)
# CLIP: Optimized for AMD ROCm
CLIP (Contrastive Language-Image Pre-training) performs zero-shot image classification by comparing image embeddings against text prompt embeddings. This repository packages evaluation/inference for zero-shot image classification using PyTorch and vLLM, exported and validated for **AMD ROCm** so it runs efficiently on AMD GPUs and CPUs.
This is based on the implementation of CLIP found [here](https://github.com/openai/CLIP).
This repository contains configurations and scripts optimized for **AMD® ROCm™** platforms. You can use the [AMD scripts](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/clip) to reproduce results or export with custom configurations. More details on model performance can be found [here](#accuracy-pipeline).
---
## Task Overview
**Task:** Zero-shot image classification
**Dataset:** CIFAR-10 test split (10,000 images, 10 classes)
**Output metrics:** Zero-shot accuracy (%), image throughput (img/s), text throughput (prompts/s)
> **Model variants:** Default is **base32** (`openai/clip-vit-base-patch32` / `ViT-B/32`). Override with `MODEL_VARIANT=base16|large14|large14-336`.
> **vLLM note:** vLLM's CLIP backend embeds **one modality per request** — text prompts and images are sent in separate API calls, then cosine similarity is computed client-side (same approach as the original evaluation scripts).
---
## AMD ROCm Optimization
This model export has been adapted and validated for **AMD Instinct™ / Radeon™ GPUs** running **ROCm**, as well as AMD CPUs. Key points:
- Validated backends: **PyTorch** (native ROCm HIP kernels via `torch.cuda`) and **vLLM** (ROCm-enabled server build with pooling runner).
- No code changes required versus the upstream CLIP implementation — only environment/runtime configuration differs.
- CPU fallback path supported via a lightweight `venv_cpu` environment for environments without a ROCm-capable GPU.
| Runtime | Precision | Backend | Hardware | Notes |
|---|---|---|---|---|
| GPU | FP32, FP16 | PyTorch | AMD RYZEN AI MAX+ 395 w/ Radeon 8060S | ROCm PyTorch, native HIP kernels |
| GPU | FP32, FP16 | vLLM | AMD RYZEN AI MAX+ 395 w/ Radeon 8060S | ROCm-enabled server, pooling runner |
| CPU | FP32, FP16 | PyTorch | AMD RYZEN AI MAX+ 395 w/ Radeon 8060S | Lightweight `venv_cpu` environment |
---
## Getting Started
For setup instructions, evaluation scripts, and custom configuration options, see the [CLIP on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/clip).
---
## Model Details
**Model Type:** Zero-shot image classification (vision-language dual encoder)
**Base Model:** `openai/clip-vit-base-patch32` (CLIP ViT-B/32)
**Model Stats:**
- Model variant: base32 (default) — base16, large14, large14-336 also supported
- Number of parameters: `151M`
- Precision tested: FP32, FP16
---
## Accuracy Pipeline
Higher zero-shot accuracy means more test images are assigned the correct CIFAR-10 class via CLIP's image–text similarity — 100% is perfect, 10% is chance level for 10 classes. Values above ~85% on CIFAR-10 with ViT-B/32 are typical for this benchmark.
### Metrics Explained
| Metric | Description |
|--------|-------------|
| Zero-shot accuracy (%) | Fraction of CIFAR-10 test images whose highest-scoring text prompt matches the ground-truth label after softmax over 10 class prompts. Primary accuracy metric; sensitive to both image and text embedding quality. |
### Accuracy Results
**Full Dataset Evaluation (CIFAR-10 test)** — filled from `evaluation_results/`; run `make metrics` to refresh:
<!-- accuracy-table-start -->
| Device | Backend | Precision | Variant | Accuracy (%) |
|--------|---------|-----------|---------|--------------|
| CPU | PyTorch | FP16 | clip-vit-base-patch32 | 88.79 |
| CPU | PyTorch | FP32 | clip-vit-base-patch32 | 88.80 |
| GPU | PyTorch | FP16 | clip-vit-base-patch32 | 88.75 |
| GPU | PyTorch | FP32 | clip-vit-base-patch32 | 88.80 |
| GPU | vLLM | FP16 | clip-vit-base-patch32 | 88.78 |
| GPU | vLLM | FP32 | clip-vit-base-patch32 | 88.80 |
<!-- accuracy-table-end -->
---
## Dig Deeper
Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?
📂 **[View the full project on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/clip)**
The GitHub repository includes:
- Setup and prerequisites for ROCm environments
- Scripts for the supported runners
- Additional model variants and datasets
- Benchmarking and reproduction instructions