File size: 5,497 Bytes
a24e3b0 c373bec a24e3b0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 | ---
library_name: onnx
license: apache-2.0
tags:
- foundation
- amd
- rocm
- video-segmentation
pipeline_tag: mask-generation
---

# SAM 2: Optimized for AMD ROCm
SAM 2 is Meta's video object segmentation model that tracks and segments objects across video frames using promptable masks. This repository packages inference for video object segmentation using **ONNX Runtime**, exported and validated for **AMD ROCm** so it runs efficiently on AMD GPUs, CPUs, and NPUs.
This is based on the implementation of SAM 2 found [here](https://github.com/facebookresearch/sam2).
This repository contains configurations and scripts optimized for **AMD® ROCm™** platforms. You can use the [sam2 AMD scripts](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/sam2) to reproduce results or export with custom configurations. More details on model performance can be found [here](#accuracy-pipeline).
---
## Task Overview
**Task:** Video object segmentation (semi-supervised VOS)
**Dataset:** DAVIS 2017 val (30 videos; smoke runs use `EVAL_MAX_VIDEOS`)
**Output metrics:** J (region similarity), F (boundary), J&F (mean)
> **Model variants:** Four SAM 2 Hiera sizes are supported — `tiny`, `small` (default), `base_plus`, `large`. Set `MODEL_SIZE` on the make command line.
---
## AMD ROCm Optimization
This model export has been adapted and validated for **AMD Instinct™ / Radeon™ GPUs** running **ROCm**, as well as AMD CPUs and AMD Ryzen AI NPUs. Key points:
- Validated backends: **ONNX Runtime** across CPU (FP32/FP16/BF16/INT8), GPU (MIGraphX execution provider — FP32/FP16/BF16), and NPU (VitisAI execution provider — FP32/FP16; BF16/INT8 disabled by default).
- No code changes required versus the upstream SAM 2 implementation — only environment/runtime configuration differs.
- INT8 is disabled on GPU (gfx1151) due to a rocBLAS kernel gap.
- `benchmark-gpu-bf16` fails at runtime — MIGraphX throws a `RuntimeException` in the compiled bf16 kernel; fp16 and fp32 GPU paths work correctly.
- `profile-gpu-int8` fails during MIGraphX compilation — the MIGraphX ONNX parser does not support the resize mode used by SAM2 in the INT8 path.
| Runtime | Precision | Backend | Hardware | Notes |
|---|---|---|---|---|
| ONNX Runtime | FP32 / FP16 / BF16 / INT8 | CPU Execution Provider | AMD CPU | — |
| ONNX Runtime | FP32 / FP16 / BF16 | MIGraphX Execution Provider | AMD Instinct™ / Radeon™ GPU (ROCm) | INT8 disabled on gfx1151 (rocBLAS gap); BF16 fails at runtime |
| ONNX Runtime | FP32 / FP16 | VitisAI Execution Provider | AMD Ryzen AI NPU | BF16/INT8 disabled by default (`ENABLE_NPU_BF16_INT8=1` to enable) |
---
## Getting Started
For setup instructions, evaluation scripts, and custom configuration options, see the [sam2 on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/sam2).
---
## Model Details
**Model Type:** Video object segmentation (promptable mask tracking) — three-stage ONNX pipeline (frame encoder, track-step init, track-step propagation)
**Base Model:** SAM 2 Hiera-small (default) — `tiny`, `base_plus`, `large` also supported
**Model Stats:**
- Input image: `(1, 3, 1024, 1024)` float32 per-frame encoder input
- GT mask (eval): `(H, W)` binary, frame 0 prompt only
- Output mask: `(H, W)` binary, per-frame segmentation
- Precision tested: FP32, FP16, BF16, INT8 (CPU); FP32, FP16, BF16 (GPU); FP32, FP16 (NPU)
---
## Accuracy Pipeline
Higher J&F means predicted masks agree more closely with ground truth across the video — 100.0 would be perfect segmentation, 0.0 means no overlap. Published SAM 2 Hiera-small J&F on DAVIS 2017 is ~89.6.
### Metrics Explained
| Metric | Description |
|--------|-------------|
| J | Region similarity (Jaccard index averaged over frames) — measures mask overlap quality. |
| F | Boundary F-measure — rewards precise mask boundaries, not just region overlap. |
| J&F | Mean of J and F — primary DAVIS semi-supervised VOS metric. |
### Accuracy Results
**Partial Dataset Evaluation (DAVIS 2017 val, 5 of 30 videos)** — default variant `MODEL_SIZE=small`; run `EVAL_MAX_VIDEOS=10 make eval-onnx-*` for full-val numbers:
<!-- accuracy-table-start -->
| Device | Precision | Videos | J | F | J&F |
|--------|-----------|--------|-----|-----|------|
| CPU | FP32 | 5 (partial) | 88.55 | 91.50 | 90.03 |
| CPU | FP16 | 5 (partial) | 88.55 | 91.50 | 90.03 |
| CPU | BF16 | 5 (partial) | 88.55 | 91.50 | 90.03 |
| CPU | INT8 | 5 (partial) | 88.55 | 91.50 | 90.03 |
| GPU | FP32 | 5 (partial) | 87.22 | 90.35 | 88.79 |
| GPU | FP16 | 5 (partial) | 87.21 | 90.36 | 88.78 |
| GPU | BF16 | 5 (partial) | 87.15 | 90.27 | 88.71 |
| NPU | FP32 | 5 (partial) | 88.55 | 91.50 | 90.03 |
| NPU | FP16 | 5 (partial) | 88.55 | 91.50 | 90.03 |
<!-- accuracy-table-end -->
These are partial DAVIS runs (5 of 30 videos); J&F is indicative only.
---
## Dig Deeper
Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?
📂 **[View the full project on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/sam2)**
The GitHub repository includes:
- Setup and prerequisites for ROCm environments
- Three-stage ONNX export (frame encoder, track-step init, track-step propagation) for CPU/GPU/NPU
- DAVIS 2017 dataset staging and J/F/J&F evaluation pipeline
- Benchmarking and reproduction instructions across CPU, GPU, and NPU
|