SAM 2.1 hiera-tiny β split ONNX (w8a8), multimask decoder
Repackaged for on-device, CPU-only tap-to-segment in DreamUI. Derived from Qualcomm AI Hub's Segment-Anything-Model-2 w8a8 ONNX release (v0.59.0), which is itself an export of Meta's SAM 2.1 hiera-tiny.
sam2-split-w8a8.zip (86,835,041 bytes) contains six flat files β each .onnx
references its .data by bare filename, so they must extract to one directory:
| File | Inputs β outputs | Runs |
|---|---|---|
trunk.onnx / .data |
image β embeddings, high-res features, pix_feat |
once per image |
prompt.onnx / .data |
unnorm_coords, labels β sparse_embedding |
once per tap (3 KB) |
decoder.onnx / .data |
embeddings + sparse_embedding β masks, scores |
once per tap |
Two changes from the AI Hub release
1. The encoder is split into trunk + prompt. AI Hub fuses the prompt encoder into the
image encoder, so a naive pipeline re-runs a 33.5M-parameter trunk on every click. The
trunk's outputs are bit-identical across clicks β only sparse_embedding varies β so
cutting at that seam turns ~817 ms per tap into ~817 ms per image plus ~28 ms per tap
(measured, Snapdragon 8 Elite, CPU EP, 4 threads).
2. The decoder emits all four mask tokens. The upstream export computes
masks [1,4,256,256] and slices [0:1]. Token 0 is the single-mask head, which blends
the competing interpretations of an ambiguous point prompt and visibly bleeds past object
boundaries; tokens 1β3 are the real granularity candidates. The Slice node's ends is
widened 1 β 4. The trailing QuantizeLinear is untouched, so masks remain uint8 on
the same scale/zero-point (0.3612250089645386 / 165) β only the channel count changes.
Measured on a 512Γ512 fixture, five clicks: token 3 was tighter than token 0 at every click (25β55% smaller area) while the model's own IoU head rated the two within 0.04.
Verified, not assumed
- Decoder channel 0 is bit-identical to the stock AI Hub decoder at all five clicks.
trunk + prompt + decoderis bit-identical toencoder + decoderat all five clicks.
Usage notes
- β
unnorm_coordswants NORMALISED coordinates, in[0, 1], despite the name. Pixel coordinates return a confidently misplaced mask. - β Label values are ignored by this export β only whether the second label is
-1(marking the second point slot unused) matters. There are no negative/background points. - β Blur the logits before thresholding. Raw thresholding produces heavy salt-and-pepper stipple along soft boundaries β checkerboard artifacting from the mask decoder's transposed convolutions, present in the float export too. A 3Γ3 box blur on the 256Γ256 logit field cuts it ~10Γ.
Quantization parameters for every tensor are in the AI Hub bundle's metadata.json.
Licence
Apache 2.0, inherited from facebookresearch/sam2. Redistributed with attribution to
Meta AI (original model) and Qualcomm AI Hub (the ONNX export these files derive from).
Model tree for AbrahamPJ/sam2-tiny-split-onnx
Base model
facebook/sam2.1-hiera-tiny