SAM3.1 β CoreML (Apple Silicon)
Real SAM3.1 weights exported to CoreML for on-device video segmentation on Apple Silicon: a stateless per-frame vision backbone plus a stateful multiplex tracker (track_step), designed to run together.
backbone/β per-frame image encoder (stateless)tracker/β per-frametrack_stepas a stateful CoreML package (multiplex_count=16)
Contract
The backbone feeds the tracker:
image [1,3,1008,1008] fp32 (normalized to [-1,1])
-> vis72 [5184,1,256]
-> hires0 [1,32,288,288] # post sam_mask_decoder.conv_s0
-> hires1 [1,64,144,144] # post conv_s1
Performance
Backbone, per-frame encode, 16GB Apple Silicon (M-series laptop):
| Engine | Mean/frame |
|---|---|
CoreML (CPU_AND_GPU) |
~2.05s |
| MLX | ~2.38s (published M3 Max ViT: ~0.8s β MLX scales better on faster silicon) |
Tracker, track_step steady-state (CoreML CPU_AND_GPU): 0.7s/frame (7s first call, JIT). CoreML fp16 vs eager PyTorch: mask sign-agree 1.0, worst relative error 3.1e-3 across 16 frames β parity holds.
Tracker: tracker/
Stateful track_step, multiplex_count=16, real SAM3.1 weights (457 tensors, tracker.model. prefix stripped from sam3.1_multiplex.pt).
models/dense_sam3_trackstep.mlpackage(fp16, 86M) β deploy this,compute_units=CPU_AND_GPU.models/dense_sam3_trackstep_fp32.mlpackage(fp32, 170M) β precision reference only;CPU_ONLYpredict at mux=16 OOMs on 16GB.scripts/β export pipeline (dense_wrapper.py,common.py,convert_fp16.py,export_coreml.py) plus parity/repro scripts (verify_coreml_lean.py,toy_*,ane_compile_test.py).patches/optimize_state.patchβ works around a coremltools 9.0 bug (optimize_state.py::canonicalize_inplace_patterndeletes ops mid-iteration; hits any graph with β₯3 rolling state buffers via shift+cat). Repro:scripts/toy_bisect.py.triton_stub/β stub sosam3/model/edt.py's bareimport tritondoesn't fail on Mac; the real EDT kernel is CUDA-only and unused here.
ANE: stateful CoreML models fail ANE compile on macOS 26.5.2 / coremltools 9.0 (platform limitation, not this graph). Deploy on CPU_AND_GPU.
Not wired: only one conditioning frame is used; the model supports max_cond_frames_in_attn=4.
Backbone: backbone/
Per-frame image encoder, stateless. Feeds the tracker with vis72 / hires0 / hires1.
models/sam3_vision_backbone.mlpackage(~875MB fp16)
Weights & license
Real SAM3.1 weights, pulled at export time from the community mirror AEmotionStudio/sam3.1 (Meta gates the original on facebook/sam3; override with SAM3_HF_REPO / SAM3_CHECKPOINT). That mirror is unofficial β verify against Meta's release before trusting it.
Derivative of Meta's SAM Materials β subject to Meta's SAM License (share-alike, field-of-use, attribution). Not permissive β redistributing the weight binaries means honoring those terms.
Env
Python 3.9, Torch 2.7.0, torchvision 0.22.0, coremltools 9.0 (with patches/optimize_state.patch) β this is a coremltools 9.0 ceiling; newer deps will drift the export path.
- Downloads last month
- -
Model tree for sdkv2/sam3.1-coreml
Base model
facebook/sam3