| --- |
| license: apache-2.0 |
| library_name: onnxruntime |
| pipeline_tag: image-segmentation |
| base_model: |
| - AnnotateIt/edgecrafter-ecseg-s-onnx |
| - AnnotateIt/edgecrafter-ecseg-m-onnx |
| tags: |
| - onnx |
| - onnxruntime |
| - onnxruntime-web |
| - instance-segmentation |
| - fp16 |
| - int8 |
| - quantization |
| - webassembly |
| - experimental |
| - negative-results |
| --- |
| |
| # ECSeg ONNX optimization experiments |
|
|
| > [!WARNING] |
| > The ONNX files in this repository are experimental reproducibility artifacts. They are **not recommended |
| > for production inference or annotation workflows**. AnnotateIt continues to use the original FP32 |
| > ECSeg-S and ECSeg-M checkpoints. |
|
|
| ## Result |
|
|
| None of the tested graph-optimized, FP16, dynamic-INT8, or static-INT8 variants improved size, browser |
| latency, segmentation quality, and runtime stability simultaneously on ONNX Runtime Web 1.24.3 CPU/WASM. |
|
|
| | Model | Decision | Production artifact | FP32 warm p50 | |
| |---|---|---|---:| |
| | ECSeg-S | **REJECT optimized variants** | [Original FP32](https://huggingface.co/AnnotateIt/edgecrafter-ecseg-s-onnx) | 1299 ms | |
| | ECSeg-M | **REJECT optimized variants** | [Original FP32](https://huggingface.co/AnnotateIt/edgecrafter-ecseg-m-onnx) | 2093 ms | |
|
|
| Measurements were made on an Apple M4 Max using single-thread WASM, the configuration used by the |
| AnnotateIt Desktop and iOS targets. Chrome was the primary engine and WebKit was used as a cross-browser |
| check. See `results/` for the machine-readable data. |
|
|
| ## Tested variants |
|
|
| | Variant | Storage result | Browser/runtime result | Quality result | |
| |---|---|---|---| |
| | Graph optimization | 1β2% larger | Warm latency unchanged | Bit-identical | |
| | FP16 | 49β51% smaller | 4β5% slower in Chrome WASM | Mean mask IoU 0.979β0.988, but individual masks reached IoU 0 | |
| | Dynamic INT8 | 58β59% smaller | Session creation fails with `ShapeInferenceError` | Not runnable | |
| | Static INT8 | 68β71% smaller | 12β17% slower | Zero instances at the production threshold | |
| | Selective static INT8 | 65β68% smaller | 12β15% slower | Zero instances at the production threshold | |
|
|
| The principal incompatibility is the DETR-style deformable-attention segmentation head, including |
| `GridSample`, `Einsum`, dynamic `TopK`, and `GatherElements`. The tested graph also failed to initialize on |
| threaded WASM and WebGPU in this runtime version. |
|
|
| ## Included experimental artifacts |
|
|
| Only the two tested FP16 variants are retained as reproducibility artifacts. Broken INT8 variants and the |
| larger graph-optimized variants are deliberately omitted. |
|
|
| | File | Size | SHA-256 | Status | |
| |---|---:|---|---| |
| | `artifacts/ecseg-s.fp16.onnx` | 21,565,662 bytes | `508c79e144fa3d9a4691970798529ae03243992aff94b0c28794b4c877c4766c` | Experimental; do not deploy | |
| | `artifacts/ecseg-m.fp16.onnx` | 41,031,365 bytes | `e0d8865bc52e0c38f14638ece649a175ccd239dfa200274f6de38efa2e10eb90` | Experimental; do not deploy | |
|
|
| Both retain the production I/O contract: |
|
|
| - Input: `images`, float32 `[1,3,640,640]` |
| - Outputs: `labels` int64 `[1,300]`, `boxes` float32 `[1,300,4]`, `scores` float32 `[1,300]`, |
| `masks` float32 `[1,300,160,160]` |
| - Opset 17, IR version 8 |
|
|
| The FP16 conversion keeps public I/O in float32 with boundary casts. The smaller file size does not |
| translate to faster CPU/WASM execution because this target has no native FP16 compute path for the graph. |
|
|
| ## Quality warning |
|
|
| Average agreement hides the failure tail. On the held-out comparison set: |
|
|
| - ECSeg-S FP16: mean mask IoU 0.988, worst-image mean 0.667, minimum individual-instance IoU 0.000. |
| - ECSeg-M FP16: mean mask IoU 0.979, worst-image mean 0.613, minimum individual-instance IoU 0.000. |
|
|
| For an annotation product, occasional fully incorrect masks are more important than the favorable average. |
| The artifacts must not be presented as drop-in optimized replacements. |
|
|
| ## Repository contents |
|
|
| - `artifacts/`: the exact FP16 binaries used in the reported measurements and `SHA256SUMS`. |
| - `results/`: CSV and JSON size, latency, and correctness results. |
| - `images/`: qualitative benchmark comparisons, including the FP16 tail and collapsed static-INT8 output. |
| - `report/ecseg-quantization-report.md`: full methodology, limitations, tables, and verdicts. |
| - `scripts/`: reproducible conversion, validation, and browser-benchmark tooling. |
|
|
| The source tooling is also maintained in the |
| [AnnotateIt repository](https://github.com/yvolokitin/AnnotateIt_npm/tree/main/scripts/model-optimization). |
|
|
| ## Reproduction |
|
|
| The exact commands are documented in `scripts/README.md`. In outline: |
|
|
| 1. Download the immutable FP32 revisions and verify their pinned SHA-256 values. |
| 2. Create an isolated Python 3.12 environment from `scripts/requirements.txt`. |
| 3. Generate variants with `scripts/optimize.py`. |
| 4. Compare raw outputs and final masks with `scripts/correctness.py`. |
| 5. Benchmark in a real browser with `scripts/bench/bench_browser.mjs` and ONNX Runtime Web 1.24.3. |
|
|
| To verify the published binaries: |
|
|
| ```bash |
| sha256sum -c artifacts/SHA256SUMS |
| ``` |
|
|
| On macOS, use `shasum -a 256` against the values in `artifacts/SHA256SUMS`. |
|
|
| ## Limitations |
|
|
| - Replacement quality was evaluated as agreement against FP32 on 47 held-out images, not as full COCO AP. |
| - Static INT8 calibration used a separate 50-image COCO val2017 subset. |
| - Performance was measured on one Apple M4 Max; Chrome and WebKit were tested. |
| - Multi-thread WASM session creation hung for every ECSeg variant, including FP32, so latency comparisons use |
| the production single-thread Desktop/iOS configuration. |
| - These results apply to ONNX Runtime Web 1.24.3. A future runtime or a redesigned/exported ECSeg head may |
| justify rerunning the study. |
|
|
| ## License and provenance |
|
|
| The ECSeg checkpoints are published under Apache-2.0. These files are numerical conversions of the |
| immutable AnnotateIt ECSeg-S and ECSeg-M ONNX releases; they are not newly trained models. The original |
| model repositories remain the authoritative production artifacts. |
|
|