license: other
license_name: sam-license
license_link: https://github.com/adambasha0/SAM3-for-Insects-segementation/blob/main/LICENSE
pipeline_tag: mask-generation
library_name: sam3-insect
tags:
- instance-segmentation
- object-detection
- insects
- arthropods
- entomology
- biodiversity-monitoring
- sam3
- segment-anything
SAM 3 for Insects — checkpoint_18
Meta's SAM 3 image model, fine-tuned for detection and instance segmentation of terrestrial arthropods and served through flat-bug-style pyramid tiling so that small specimens in large trap images are recovered rather than lost to downsampling.
One instance mask, one bounding box and one confidence per insect, as COCO polygons. It is a single-class detector: it finds arthropods, it does not identify them.
| Code, CLI and Colab notebook | github.com/adambasha0/SAM3-for-Insects-segementation |
| Try it in Colab | notebook |
| Base model | SAM 3 image model |
| Training data | The flat-bug aggregate — 23 insect datasets |
| Fine-tuning run | flatbug_medium_ft, epoch 18 of 20, 1008 px, lr 8e-5 |
| File | checkpoint_18_inference.pt — 3.14 GB, fp32, unmodified |
| SHA-256 | dd8a6ce0402a6c2d00b2849a3e08becc6f3aa4ececdc526580a54539c9c41829 |
Usage
pip install -e "git+https://github.com/adambasha0/SAM3-for-Insects-segementation.git#egg=sam3-insect"
from sam3_insect import InsectPredictor, annotations_to_coco, resolve_checkpoint
predictor = InsectPredictor(resolve_checkpoint("hf", hf_repo="tea98/sam3-for-insects-segmentation"))
result = predictor.predict("trap_photo.jpg")
strong = [a for a in result.annotations if a["score"] >= 0.4]
print(f"{len(strong)} insects found")
predictor.render(result, score_threshold=0.4).save("overview.jpg")
coco = annotations_to_coco(result)
Or fetch the file directly:
from huggingface_hub import hf_hub_download
path = hf_hub_download("tea98/sam3-for-insects-segmentation", "checkpoint_18_inference.pt")
Inference needs no HuggingFace token and no licence acceptance: this checkpoint carries every
weight the model uses, so the gated facebook/sam3 repository is never touched.
Choosing a threshold
SAM 3's decoder spends a fixed budget of 200 object queries in full on every tile and has no per-query "nothing here" output — a query declines only by scoring low. So on a tile holding N insects you get ~200 detections above 0.005 and roughly N above 0.5, at the same recall. That is an operating-point property of DETR-style detectors, not a defect from fine-tuning.
In the fine-tuning domain real detections usually score above 0.8. The library runs inference at 0.02 and the demo app filters its display at 0.4. Fix a threshold before reporting a count.
Known limitations
- Single class — no taxonomy.
- Masks run tight, most visibly on hairy or translucent specimens and on legs and antennae.
Lower
MASK_THRESHOLDtowards 0.3 when mask area matters. - Very small objects still need help; raise
SCALE_BEFOREto 1.5–2.0. - Domain shift costs precision more than recall — on unfamiliar imagery, raise the threshold.
- Dense clusters are merged or split inconsistently; NMS at IoU 0.2 will suppress a genuinely overlapping pair.
- fp16 does not work.
text_projectionholds values up to 9.58e18, past fp16's 65504 ceiling, so casting turns those weights intoinf. Keep fp32 weights; activations are autocast at runtime.
The full model card, including how the weights' fidelity was verified, is in the repository.
Licence and citation
Derivative work of the SAM 3 materials, governed by the SAM License — redistribution must carry a copy of that licence and stay within Meta's Acceptable Use Policy. Cite SAM 3 and flat-bug alongside this work.
Fine-tuning by Adam Basha, Karlsruhe Institute of Technology, 2026.