homr ONNX checkpoints
This repository is a mirror of the ONNX checkpoints that liebharc/homr downloads at runtime, published here so that client-side (browser) builds can fetch them.
It exists because the models on homr's GitHub Release tag are served without
Access-Control-Allow-Origin, which blocks a browser from downloading them. Hugging
Face serves these files with CORS enabled. See HuggingFaceModels.md in the homr
repository for the full rationale and workflow.
The source of truth is the GitHub release tag, not this repository:
https://github.com/liebharc/homr/releases/download/onnx_checkpoints/
If the two ever disagree, the GitHub release wins.
Models
Three models, each in fp32 and fp16:
| File | Role |
|---|---|
segnet_308-3296ccd40960f90ca6ab9c035cca945675d30a0f.onnx |
Segmentation (staff lines, note heads, stems/rests, bar lines, clefs/keys) |
segnet_308-3296ccd40960f90ca6ab9c035cca945675d30a0f_fp16.onnx |
…fp16 |
encoder_pytorch_model_465-597144cab54c8f6d0f6c9619df5c5312694eadd6.onnx |
TrOMR transformer encoder (ConvNeXt) |
encoder_pytorch_model_465-597144cab54c8f6d0f6c9619df5c5312694eadd6_fp16.onnx |
…fp16 |
decoder_pytorch_model_465-597144cab54c8f6d0f6c9619df5c5312694eadd6.onnx |
TrOMR transformer decoder, with KV cache |
decoder_pytorch_model_465-597144cab54c8f6d0f6c9619df5c5312694eadd6_fp16.onnx |
…fp16 |
Exact sizes:
| File | Size |
|---|---|
segnet_308-….onnx |
54.66 MB |
segnet_308-…_fp16.onnx |
27.34 MB |
encoder_pytorch_model_465-….onnx |
50.41 MB |
encoder_pytorch_model_465-…_fp16.onnx |
25.24 MB |
decoder_pytorch_model_465-….onnx |
45.12 MB |
decoder_pytorch_model_465-…_fp16.onnx |
89.61 MB |
Note that fp16 is not smaller for the decoder (89.61 MB vs 45.12 MB). A browser build targeting WebGPU should use fp16 segnet + fp16 encoder + fp32 decoder (97.7 MB total) rather than all-fp16 (292.4 MB).
Usage in a browser
const url =
'https://huggingface.co/ngbcoder/Homr-onnx/resolve/main/' +
'segnet_308-3296ccd40960f90ca6ab9c035cca945675d30a0f_fp16.onnx';
const session = await ort.InferenceSession.create(url);
Cache the bytes (caches.open(...) / Cache Storage) keyed by the full URL: at ~98 MB
per visitor, re-downloading on every page load is not acceptable.
Prefer to pin a specific revision? Replace main with a commit SHA from the
Files tab of this repo.
Intended input/output shapes
| Model | Input | Output |
|---|---|---|
| segnet | [batch, 3, 320, 320] float32 or float16 (matches the fp16 variant) |
[batch, 6, 320, 320] logits, argmax over axis 0 |
| encoder | [1, 1, H, W] float32/float16, H ≤ 256, W ≤ 1280 |
context |
| decoder | rhythms, pitchs, lifts, articulations, slurs, context, cache_len, cache_in0..31 |
out_rhythms, out_pitchs, out_lifts, out_positions, out_articulations, out_slurs, attention, cache_out0..31 |
The decoder expects 32 KV-cache tensors because it has decoder_depth = 8 layers with
4 cache tensors each.
Normalisation for the encoder input is (pixel / 255 - 0.7931) / 0.1738, with a single
channel.
Provenance
These weights come from homr's training pipeline, which builds on:
- the segmentation models of oemer
- the transformer model of Polyphonic-TrOMR
homr itself is licensed AGPL-3.0. Please confirm the licensing terms of the upstream model weights before redistributing them.
Updating
Re-export from homr's training pipeline, publish to the onnx_checkpoints release tag,
then mirror here. Never change a model in only one of the two places.