File size: 4,654 Bytes
0785ae0 4958570 0785ae0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 | ---
library_name: pytorch
license: bsd-3-clause
tags:
- foundation
- amd
- rocm
- visual-question-answering
pipeline_tag: visual-question-answering
---

# BLIP-2: Optimized for AMD ROCm
BLIP-2 (Bootstrapping Language-Image Pre-training) is a vision-language model for visual question answering. This repository packages evaluation/inference for VQA using PyTorch (CPU and GPU) and in-process vLLM (GPU), exported and validated for **AMD ROCm** so it runs efficiently on AMD GPUs and CPUs.
This is based on the implementation of BLIP-2 found [here](https://huggingface.co/Salesforce/blip2-opt-2.7b).
This repository contains configurations and scripts optimized for **AMD® ROCm™** platforms. You can use the [BLIP-2 AMD scripts](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/blip2) to reproduce results or export with custom configurations. More details on model performance can be found [here](#accuracy-pipeline).
---
## Task Overview
**Task:** Visual question answering (VQA)
**Dataset:** VQAv2 validation split (214,354 questions on 40,504 COCO val2014 images; default eval uses 2,000 questions)
**Output metrics:** Soft VQA accuracy (%), throughput (samples/s), latency (ms/sample)
> **Model variants:** Default is **opt-2.7b** (`Salesforce/blip2-opt-2.7b`). Override with `MODEL_VARIANT=opt-6.7b|flan-t5-xl`.
> **vLLM note:** BLIP-1 architectures are unsupported by vLLM; this example uses BLIP-2 (`Blip2ForConditionalGeneration`) only.
---
## AMD ROCm Optimization
This model export has been adapted and validated for **AMD Instinct™ / Radeon™ GPUs** running **ROCm**, as well as AMD CPUs. Key points:
- Validated backends: **PyTorch** (CPU and GPU, native ROCm HIP kernels) and **vLLM** (in-process, GPU only).
- No code changes required versus the upstream BLIP-2 implementation — only environment/runtime configuration differs.
- CPU fallback path supported (FP32 only) for environments without a ROCm-capable GPU.
- On ROCm gfx1151, `VLLM_ROCM_USE_SKINNY_GEMM=0` is set in the Makefile.
| Runtime | Precision | Backend | Hardware | Notes |
|---|---|---|---|---|
| GPU | FP32 / FP16 / BF16 | PyTorch | AMD Eng Sample: 100-000002199-50_Y | Native ROCm HIP kernels |
| GPU | FP32 / FP16 / BF16 | vLLM | AMD Eng Sample: 100-000002199-50_Y | In-process vLLM server |
| CPU | FP32 | PyTorch | AMD Eng Sample: 100-000002199-50_Y | Smoke-test path only |
---
## Getting Started
For setup instructions, evaluation scripts, and custom configuration options, see the [BLIP-2 on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/blip2).
---
## Model Details
**Model Type:** Vision-language model for visual question answering
**Base Model:** `Salesforce/blip2-opt-2.7b` (BLIP-2 OPT-2.7b)
**Model Stats:**
- Model variant: opt-2.7b (default) — opt-6.7b, flan-t5-xl also supported
- Architecture: `Blip2ForConditionalGeneration`
- Precision tested: FP32, FP16, BF16
---
## Accuracy Pipeline
Higher soft VQA accuracy means more generated answers overlap ground-truth answer tokens — 100% would be perfect, ~0% is chance-level.
### Metrics Explained
| Metric | Description |
|--------|-------------|
| Soft VQA accuracy (%) | Fraction of questions where the model's answer shares normalized tokens with any human-provided answer (VQAv2 soft match). Primary quality metric; sensitive to answer phrasing and temperature. |
| Throughput (samples/s) | VQA question–image pairs processed per second end-to-end. Higher is faster; compare at the same sample count and dtype. |
### Accuracy Results
**Full Dataset Evaluation (VQAv2 val, 2,000 questions)** — `Salesforce/blip2-opt-2.7b`:
<!-- accuracy-table-start -->
| Device | Backend | Precision | VQA Accuracy (%) | Throughput (samples/s) |
|--------|---------|-----------|-------------------|-------------------------|
| CPU | PyTorch | FP32 | 62.50 | 0.43 |
| GPU | PyTorch | FP32 | 62.50 | 2.61 |
| GPU | PyTorch | FP16 | 62.50 | 4.49 |
| GPU | PyTorch | BF16 | 62.50 | 4.33 |
| GPU | vLLM | FP32 | 41.67 | 2.42 |
| GPU | vLLM | FP16 | 41.67 | 5.22 |
| GPU | vLLM | BF16 | 45.83 | 4.50 |
<!-- accuracy-table-end -->
---
## Dig Deeper
Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?
📂 **[View the full project on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/blip2)**
The GitHub repository includes:
- Setup and prerequisites for ROCm environments
- Scripts for the supported runners
- Additional model variants and datasets
- Benchmarking and reproduction instructions
|