| --- |
| license: apache-2.0 |
| pipeline_tag: image-segmentation |
| library_name: onnx |
| tags: |
| - anylearning |
| - onnx |
| - onnxruntime |
| - segment-anything |
| - segment-anything-2 |
| - segment-anything-2.1 |
| - image-segmentation |
| - object-detection |
| - instance-segmentation |
| - promptable-segmentation |
| - edge-ai |
| authors: |
| - Neural Research Lab |
| - Viet-Anh Nguyen |
| --- |
| |
| # AnyLearning labeling models |
|
|
| Versioned ONNX model bundles used by |
| [AnyLearning](https://github.com/nrl-ai/anylearning-oss) for local object |
| detection, instance segmentation, and prompt-guided image segmentation. |
|
|
| This repository contains the seventeen models currently offered by AnyLearning: |
|
|
| | File | Model | Download size | SHA-256 | |
| | --- | --- | ---: | --- | |
| | `mobile_sam_20230629.zip` | MobileSAM | 36,655,105 bytes | `41aff2660b7531becfee21fb257c49933ddc892c554507bdb775bf504d443942` | |
| | `efficientvit_sam_l0.zip` | EfficientViT-SAM-L0 | 129,322,579 bytes | `ef48c8f4c6d72da3df0690d9f16aa794035978d23a637acbd24a984220020072` | |
| | `efficientvit_sam_l1.zip` | EfficientViT-SAM-L1 | 177,189,481 bytes | `e7f6a5a38a00b53c6706ea45e70c5601b2140c57729ad6412c0600fc719634a2` | |
| | `efficientvit_sam_l2.zip` | EfficientViT-SAM-L2 | 228,296,139 bytes | `be208e948e445b2d3743d17c2cab83564a403b698e90c9cbdfb1ea6273ab823e` | |
| | `efficientvit_sam_xl0.zip` | EfficientViT-SAM-XL0 | 434,729,250 bytes | `98ccd8ca3a4ff339cdaa29f6f8ff6dad259516f38e7fb093c8b5cbd575d604ea` | |
| | `efficientvit_sam_xl1.zip` | EfficientViT-SAM-XL1 | 756,188,273 bytes | `316c92d1900b09e102cfbb3585f8da25b6df20c6c6fe0f0ef6552ec3fc0aab6d` | |
| | `sam2_hiera_tiny.zip` | SAM 2 Hiera-Tiny | 143,293,170 bytes | `f96e764cf19ba63e7870f8d588cfbfe772cda40a13180f19e73100e28f8c92da` | |
| | `sam2_hiera_small.zip` | SAM 2 Hiera-Small | 169,638,840 bytes | `90aa65e8b9c98cca73916c0b579d0b63139ef60e8b508946c389204819195261` | |
| | `sam2_hiera_base_plus.zip` | SAM 2 Hiera-Base+ | 302,582,623 bytes | `c30ce8030d9be4ee68afb825285d0d50b706571a671847ba0d7cb36e7aab75c3` | |
| | `sam2_hiera_large.zip` | SAM 2 Hiera-Large | 843,636,551 bytes | `97d5841ca0827acb3ae54ed896adedb23cc77a4fa23118402efca3e169d20542` | |
| | `sam2_1_hiera_tiny.zip` | SAM 2.1 Hiera-Tiny | 116,527,977 bytes | `b15c3d3265392901f6dbe45a42e2f64709bf265b15fde899b2439bc04496ba04` | |
| | `sam2_1_hiera_small.zip` | SAM 2.1 Hiera-Small | 142,921,559 bytes | `5802f9e3b05d41b9a5ce03f32d3935fe6cc94294b77956f37c58b7f2581471ce` | |
| | `sam2_1_hiera_base_plus.zip` | SAM 2.1 Hiera-Base+ | 272,032,266 bytes | `389aae8f1e552275326e13a10973e822583989dba95aa7c3ffc80022460ac7eb` | |
| | `sam2_1_hiera_large.zip` | SAM 2.1 Hiera-Large | 805,293,551 bytes | `15f74c68530b0bc9f37d2189394a100b81541f55362b34f55a309ba12b9e1fa4` | |
| | `rfdetr_nano_detection_1_9_4.zip` | RF-DETR Nano detection | 99,769,055 bytes | `5b130a1c2eb01be3bfbda703367b5d993821690de26f4abbb14a94ee3660c5fe` | |
| | `rfdetr_nano_segmentation_1_9_4.zip` | RF-DETR Segmentation Nano | 113,602,026 bytes | `133fdb5aed76233a6959addbdb7d64f5132f3cf2b2705a1994c2cbdfc2e93f4d` | |
| | `dfine_n_coco_956d170.zip` | D-FINE-N COCO detection | 13,979,719 bytes | `f753f6e552632ef1696ec53c92ccb1ca374dbf2f014c6e4ab005ab98efa3a6e8` | |
|
|
| Each promptable-segmentation ZIP contains a small AnyLearning model |
| configuration and one encoder plus one decoder ONNX model. Each RF-DETR and |
| D-FINE ZIP contains one static ONNX graph plus its provenance, checksums, and |
| verbatim upstream license. `MANIFEST.json` records the exact source revision, archive |
| size, checksum, and expected members. Every transformed or exported bundle pins |
| each extracted member's size and SHA-256 as well as its source artifact |
| identity. |
|
|
| ## Provenance |
|
|
| The source graph pairs originate from checksum-pinned ONNX exports published by |
| Viet-Anh Nguyen: |
|
|
| - SAM 2 bundles: source revision |
| `071f58077599431edd0e5d2ac52ecca4c78f1cab` from |
| `vietanhdev/segment-anything-2-onnx-models`. |
| - SAM 2.1 bundles: source revision |
| `6a3ac868340a3196a349050a6efae22a5acc0330` from |
| `vietanhdev/segment-anything-2.1-onnx-models`. |
| - MobileSAM bundle: source revision |
| `9effc01a9e135621d710d49159f1ffb0b6f724dc` from |
| `vietanhdev/segment-anything-onnx-models`. |
| - EfficientViT-SAM encoders and source decoders: immutable model revision |
| `a2f0c5929196a13ef1ffc9338d3f9e482e1e0e68` from |
| `mit-han-lab/efficientvit-sam`. The published decoders were transformed by |
| [AnyLearning's checksum-gated ONNX-only tool](https://github.com/nrl-ai/anylearning-oss/blob/450e4fc7e291d7166b5ac650b7b72c42fac70203/scripts/prepare_efficientvit_sam_decoder.py) |
| to expose the four native mask tokens before the official graph's lossy |
| single-mask selection. The transform does not load a native checkpoint or |
| alter the learned tensors. |
| - RF-DETR Nano detection and instance segmentation: official checkpoints and |
| exporter from `rfdetr==1.9.4`, source revision |
| `9b009fa928d6218320439803d1da01869a85c072`. Both graphs use static batch-one |
| float32 contracts at opset 17. Native/export parity evidence is recorded |
| inside each archive. |
| - D-FINE-N detection: the official COCO-only checkpoint and exporter at source |
| revision `956d1709314c2c6a4df6f34de232054578a7449f`. The checkpoint was loaded |
| only by the included restricted weights-only conversion helper. The static |
| opset-16 graph has exact native label parity and less than 0.005-pixel box |
| drift on the retained landscape/portrait corpus. Objects365-derived weights |
| are excluded. |
|
|
| The SAM 2 and SAM 2.1 encoders were transformed by |
| [AnyLearning's checksum-gated ONNX-only tool](https://github.com/nrl-ai/anylearning-oss/blob/7575ab52cc0a31e5ec3d71b6ec7157dbe042c7d7/scripts/prepare_sam2_encoder.py). |
| It repairs exact stale shape metadata, removes only unreachable initializers, |
| requires strict ONNX type/shape inference, and records the source digest. It |
| does not load a native checkpoint or change reachable learned tensors. The |
| decoders are unchanged mirrors. |
|
|
| Original model projects: |
|
|
| - [Segment Anything 2](https://github.com/facebookresearch/sam2) |
| - [Segment Anything](https://github.com/facebookresearch/segment-anything) |
| - [MobileSAM](https://github.com/ChaoningZhang/MobileSAM) |
| - [EfficientViT-SAM](https://github.com/mit-han-lab/efficientvit/tree/de7d7733cc0329f391b33f1f459271562ec27bd5/applications/efficientvit_sam) |
| - [RF-DETR](https://github.com/roboflow/rf-detr/tree/9b009fa928d6218320439803d1da01869a85c072) |
| - [D-FINE](https://github.com/Peterande/D-FINE/tree/956d1709314c2c6a4df6f34de232054578a7449f) |
|
|
| The SAM and MobileSAM files are unchanged mirrors. EfficientViT-SAM encoders |
| are unchanged mirrors; their decoders preserve the original learned tensors and |
| compute graph while exposing the four pre-selection outputs and recording |
| transform provenance in ONNX metadata. SAM 2 and SAM 2.1 prepared encoders are |
| paired with unchanged source decoders in deterministic archives. |
|
|
| ## Secure and reproducible download |
|
|
| Pin a repository revision and verify the SHA-256 value from `MANIFEST.json` |
| before extracting or loading a model. Consumers should reject absolute paths, |
| parent traversal, links, encrypted or duplicate entries, unexpected archive |
| members, and files whose exact sizes or digests differ from the manifest. |
|
|
| AnyLearning performs the full image encoding, prompt conversion, mask decoding, |
| and editable-shape conversion. These archives are not standalone applications. |
|
|
| ## Intended use and limitations |
|
|
| - Intended for interactive point/rectangle-prompt segmentation in AnyLearning. |
| - RF-DETR bundles are intended for human-reviewed COCO object detection and |
| instance-segmentation suggestions. |
| - The D-FINE bundle is intended for human-reviewed COCO object-detection |
| suggestions. |
| - Results require human review before becoming dataset labels. |
| - Quality and latency vary with image content, hardware, execution provider, and |
| model size. |
| - These models can reproduce biases and limitations of their original training |
| data. |
| - Do not use segmentation output as the sole basis for safety-critical, |
| medical, legal, or similarly consequential decisions. |
|
|
| ## License |
|
|
| The model code and weights are distributed under Apache License 2.0 by their |
| respective upstream projects. See `LICENSES.md` for source and attribution links. |
|
|
| ## Citation |
|
|
| For SAM 2 and SAM 2.1, cite: |
|
|
| ```bibtex |
| @article{ravi2024sam2, |
| title={SAM 2: Segment Anything in Images and Videos}, |
| author={Ravi, Nikhila and others}, |
| journal={arXiv:2408.00714}, |
| year={2024} |
| } |
| ``` |
|
|
| For SAM, cite: |
|
|
| ```bibtex |
| @article{kirillov2023segment, |
| title={Segment Anything}, |
| author={Kirillov, Alexander and others}, |
| journal={arXiv:2304.02643}, |
| year={2023} |
| } |
| ``` |
|
|
| For EfficientViT-SAM, cite: |
|
|
| ```bibtex |
| @inproceedings{cai2023efficientvit, |
| title={EfficientViT: Lightweight Multi-Scale Attention for High-Resolution Dense Prediction}, |
| author={Cai, Han and Li, Junyan and Hu, Muyan and Gan, Chuang and Han, Song}, |
| booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision}, |
| year={2023} |
| } |
| |
| @article{zhang2024efficientvit, |
| title={EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss}, |
| author={Zhang, Zhuoyang and Cai, Han and Han, Song}, |
| journal={arXiv preprint arXiv:2402.05008}, |
| year={2024} |
| } |
| ``` |
|
|
| For RF-DETR, use the citation requested by the |
| [official project](https://github.com/roboflow/rf-detr/tree/9b009fa928d6218320439803d1da01869a85c072). |
|
|
| For D-FINE, use the citation requested by the |
| [official project](https://github.com/Peterande/D-FINE/tree/956d1709314c2c6a4df6f34de232054578a7449f). |
|
|