BLIP-2: Optimized for AMD ROCm

BLIP-2 (Bootstrapping Language-Image Pre-training) is a vision-language model for visual question answering. This repository packages evaluation/inference for VQA using PyTorch (CPU and GPU) and in-process vLLM (GPU), exported and validated for AMD ROCm so it runs efficiently on AMD GPUs and CPUs.

This is based on the implementation of BLIP-2 found here. This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the BLIP-2 AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.


Task Overview

Task: Visual question answering (VQA)

Dataset: VQAv2 validation split (214,354 questions on 40,504 COCO val2014 images; default eval uses 2,000 questions)

Output metrics: Soft VQA accuracy (%), throughput (samples/s), latency (ms/sample)

Model variants: Default is opt-2.7b (Salesforce/blip2-opt-2.7b). Override with MODEL_VARIANT=opt-6.7b|flan-t5-xl.

vLLM note: BLIP-1 architectures are unsupported by vLLM; this example uses BLIP-2 (Blip2ForConditionalGeneration) only.


AMD ROCm Optimization

This model export has been adapted and validated for AMD Instinct™ / Radeon™ GPUs running ROCm, as well as AMD CPUs. Key points:

  • Validated backends: PyTorch (CPU and GPU, native ROCm HIP kernels) and vLLM (in-process, GPU only).
  • No code changes required versus the upstream BLIP-2 implementation — only environment/runtime configuration differs.
  • CPU fallback path supported (FP32 only) for environments without a ROCm-capable GPU.
  • On ROCm gfx1151, VLLM_ROCM_USE_SKINNY_GEMM=0 is set in the Makefile.
Runtime Precision Backend Hardware Notes
GPU FP32 / FP16 / BF16 PyTorch AMD Eng Sample: 100-000002199-50_Y Native ROCm HIP kernels
GPU FP32 / FP16 / BF16 vLLM AMD Eng Sample: 100-000002199-50_Y In-process vLLM server
CPU FP32 PyTorch AMD Eng Sample: 100-000002199-50_Y Smoke-test path only

Getting Started

For setup instructions, evaluation scripts, and custom configuration options, see the BLIP-2 on GitHub.


Model Details

Model Type: Vision-language model for visual question answering

Base Model: Salesforce/blip2-opt-2.7b (BLIP-2 OPT-2.7b)

Model Stats:

  • Model variant: opt-2.7b (default) — opt-6.7b, flan-t5-xl also supported
  • Architecture: Blip2ForConditionalGeneration
  • Precision tested: FP32, FP16, BF16

Accuracy Pipeline

Higher soft VQA accuracy means more generated answers overlap ground-truth answer tokens — 100% would be perfect, ~0% is chance-level.

Metrics Explained

Metric Description
Soft VQA accuracy (%) Fraction of questions where the model's answer shares normalized tokens with any human-provided answer (VQAv2 soft match). Primary quality metric; sensitive to answer phrasing and temperature.
Throughput (samples/s) VQA question–image pairs processed per second end-to-end. Higher is faster; compare at the same sample count and dtype.

Accuracy Results

Full Dataset Evaluation (VQAv2 val, 2,000 questions)Salesforce/blip2-opt-2.7b:

Device Backend Precision VQA Accuracy (%) Throughput (samples/s)
CPU PyTorch FP32 62.50 0.43
GPU PyTorch FP32 62.50 2.61
GPU PyTorch FP16 62.50 4.49
GPU PyTorch BF16 62.50 4.33
GPU vLLM FP32 41.67 2.42
GPU vLLM FP16 41.67 5.22
GPU vLLM BF16 45.83 4.50

Dig Deeper

Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?

📂 View the full project on GitHub

The GitHub repository includes:

  • Setup and prerequisites for ROCm environments
  • Scripts for the supported runners
  • Additional model variants and datasets
  • Benchmarking and reproduction instructions
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support