sam3 / README.md
AMD-PAVS-AI's picture
Update image references to HF URLs
dacd847 verified
|
Raw
History Blame Contribute Delete
6.09 kB
---
library_name: pytorch
license: other
tags:
- foundation
- amd
- rocm
- image-segmentation
pipeline_tag: mask-generation
---
![](https://huggingface.co/AMD-PAVS-AI/sam3/resolve/main/sam3.png)
# SAM3: Optimized for AMD ROCm
SAM3 (Segment Anything Model 3) is a promptable instance segmentation model that generates high-quality masks from text, box, or point prompts. This repository packages inference for instance segmentation using **PyTorch**, exported and validated for **AMD ROCm** so it runs efficiently on AMD GPUs and CPUs.
This is based on the implementation of SAM3 found [here](https://huggingface.co/facebook/sam3).
This repository contains configurations and scripts optimized for **AMD® ROCm™** platforms. You can use the [sam3 AMD scripts](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/sam3) to reproduce results or export with custom configurations. More details on model performance can be found [here](#accuracy-pipeline).
---
## Task Overview
**Task:** Instance segmentation
**Dataset:** COCO val2017 (5,000 images, 80 categories)
**Output metrics:** Mask mAP@0.5:0.95, Mask mAP@0.5, Mask mAP@0.75, Mean Mask IoU, Mask Precision, Mask Recall, Mask F1
> **PyTorch note:** CPU runs FP32 only; GPU supports FP32/FP16/BF16 via `torch.autocast` (no INT8 — SAM3 is PyTorch-only, not ONNX/MIGraphX). NPU is not available since ONNX export is unsupported upstream. Hugging Face access is required — SAM3 weights are gated; you must accept the license at [facebook/sam3](https://huggingface.co/facebook/sam3) and export `HF_TOKEN`.
---
## AMD ROCm Optimization
This model export has been adapted and validated for **AMD Instinct™ / Radeon™ GPUs** running **ROCm**, as well as AMD CPUs. Key points:
- Validated backend: **PyTorch** (native ROCm HIP kernels via `torch.autocast`), FP32/FP16/BF16 on GPU, FP32 on CPU.
- No code changes required versus the upstream SAM3 implementation — only environment/runtime configuration differs.
- No NPU path — SAM3 ONNX export is unsupported upstream, and VitisAI requires an ONNX graph.
| Runtime | Precision | Backend | Hardware | Notes |
|---|---|---|---|---|
| PyTorch | FP32 | HIP (ROCm) | AMD CPU | — |
| PyTorch | FP32 / FP16 / BF16 | HIP (ROCm) | AMD Instinct™ / Radeon™ GPU | Via `torch.autocast`; no INT8 path |
---
## Getting Started
For setup instructions, evaluation scripts, and custom configuration options, see the [sam3 on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/sam3).
---
## Model Details
**Model Type:** Promptable instance segmentation (text, box, or point prompts)
**Base Model:** `facebook/sam3` (gated)
**Model Stats:**
- Precision tested: FP32 (CPU); FP32, FP16, BF16 (GPU)
---
## Accuracy Pipeline
Higher Mask mAP means the model's predicted masks agree more closely with the ground-truth segmentation annotations — 1.0 would be perfect mask overlap, 0.0 means no correct masks. Because SAM3 is a promptable segmentation model (not a detector), these metrics measure pure mask quality independent of detection accuracy. Since full-dataset inference is too slow for a complete 5,000-image COCO val2017 pass, evaluation uses three nested random subsets (N=100, 300, 500) drawn once with a fixed seed for reproducibility.
### Metrics Explained
| Metric | Description |
|--------|-------------|
| Mask mAP@0.5:0.95 | Primary COCO mask metric — mean AP averaged across IoU thresholds 0.5–0.95. The strictest, most holistic mask accuracy number; higher means predicted masks overlap ground truth tightly across a range of thresholds. |
| Mask mAP@0.5 | AP at a single, looser IoU threshold of 0.5 — a predicted mask only needs 50% overlap to count as correct, so this is typically higher than mAP@0.5:0.95 and reflects "did it segment the object" more than "how precisely." |
| Mask mAP@0.75 | AP at a stricter IoU threshold of 0.75 — the predicted mask must overlap ground truth by 75%, rewarding precise segmentation boundaries over merely finding the object. |
| Mean Mask IoU | Average intersection-over-union across all predicted masks — a direct geometric measure of how well each individual mask covers its ground-truth region, complementing the AP-based aggregate metrics. |
| Mask Precision | Of all masks the model produced, what fraction were correct — high precision means few false-positive masks (segments that don't correspond to real objects). |
| Mask Recall | Of all ground-truth objects, what fraction did the model successfully segment — high recall means few missed objects (false negatives). |
| Mask F1 | Harmonic mean of Mask Precision and Mask Recall — balances the two failure modes (false positives vs. false negatives) into a single number. |
### Accuracy Results
**Sampled Evaluation (COCO val2017)** — nested subsets, seed=42:
<!-- accuracy-table-start -->
| Device | Precision | N | Mask mAP@0.5:0.95 | Mask mAP@0.5 | Mean IoU | Mask Precision | Mask Recall | Mask F1 |
|--------|-----------|---|-------------------|-------------|----------|----------------|-------------|---------|
| CPU | FP32 | 100 | 0.6269 | 0.8982 | 0.7933 | 0.5895 | 0.6166 | 0.6027 |
| CPU | FP32 | 300 | 0.5991 | 0.8937 | 0.7878 | 0.5525 | 0.6026 | 0.5765 |
| CPU | FP32 | 500 | 0.5813 | 0.9045 | 0.7886 | 0.5390 | 0.5938 | 0.5651 |
| GPU | FP32 | 100 | 0.6269 | 0.8982 | 0.7933 | 0.5895 | 0.6166 | 0.6027 |
| GPU | FP32 | 300 | 0.5991 | 0.8937 | 0.7878 | 0.5525 | 0.6026 | 0.5765 |
| GPU | FP32 | 500 | 0.5813 | 0.9045 | 0.7886 | 0.5390 | 0.5938 | 0.5651 |
<!-- accuracy-table-end -->
---
## Dig Deeper
Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?
📂 **[View the full project on GitHub](https://github.com/AMD-PAVS/physical_ai_sdk/blob/main/models/sam3)**
The GitHub repository includes:
- Setup and prerequisites for ROCm environments
- Text/box/point prompt-based inference scripts
- COCO val2017 sampled evaluation pipeline with fixed-seed reproducibility
- Benchmarking and reproduction instructions for CPU and GPU