homr ONNX checkpoints

This repository is a mirror of the ONNX checkpoints that liebharc/homr downloads at runtime, published here so that client-side (browser) builds can fetch them.

It exists because the models on homr's GitHub Release tag are served without Access-Control-Allow-Origin, which blocks a browser from downloading them. Hugging Face serves these files with CORS enabled. See HuggingFaceModels.md in the homr repository for the full rationale and workflow.

The source of truth is the GitHub release tag, not this repository:

https://github.com/liebharc/homr/releases/download/onnx_checkpoints/

If the two ever disagree, the GitHub release wins.

Models

Three models, each in fp32 and fp16:

File Role
segnet_308-3296ccd40960f90ca6ab9c035cca945675d30a0f.onnx Segmentation (staff lines, note heads, stems/rests, bar lines, clefs/keys)
segnet_308-3296ccd40960f90ca6ab9c035cca945675d30a0f_fp16.onnx …fp16
encoder_pytorch_model_465-597144cab54c8f6d0f6c9619df5c5312694eadd6.onnx TrOMR transformer encoder (ConvNeXt)
encoder_pytorch_model_465-597144cab54c8f6d0f6c9619df5c5312694eadd6_fp16.onnx …fp16
decoder_pytorch_model_465-597144cab54c8f6d0f6c9619df5c5312694eadd6.onnx TrOMR transformer decoder, with KV cache
decoder_pytorch_model_465-597144cab54c8f6d0f6c9619df5c5312694eadd6_fp16.onnx …fp16

Exact sizes:

File Size
segnet_308-….onnx 54.66 MB
segnet_308-…_fp16.onnx 27.34 MB
encoder_pytorch_model_465-….onnx 50.41 MB
encoder_pytorch_model_465-…_fp16.onnx 25.24 MB
decoder_pytorch_model_465-….onnx 45.12 MB
decoder_pytorch_model_465-…_fp16.onnx 89.61 MB

Note that fp16 is not smaller for the decoder (89.61 MB vs 45.12 MB). A browser build targeting WebGPU should use fp16 segnet + fp16 encoder + fp32 decoder (97.7 MB total) rather than all-fp16 (292.4 MB).

Usage in a browser

const url =
  'https://huggingface.co/ngbcoder/Homr-onnx/resolve/main/' +
  'segnet_308-3296ccd40960f90ca6ab9c035cca945675d30a0f_fp16.onnx';

const session = await ort.InferenceSession.create(url);

Cache the bytes (caches.open(...) / Cache Storage) keyed by the full URL: at ~98 MB per visitor, re-downloading on every page load is not acceptable.

Prefer to pin a specific revision? Replace main with a commit SHA from the Files tab of this repo.

Intended input/output shapes

Model Input Output
segnet [batch, 3, 320, 320] float32 or float16 (matches the fp16 variant) [batch, 6, 320, 320] logits, argmax over axis 0
encoder [1, 1, H, W] float32/float16, H ≤ 256, W ≤ 1280 context
decoder rhythms, pitchs, lifts, articulations, slurs, context, cache_len, cache_in0..31 out_rhythms, out_pitchs, out_lifts, out_positions, out_articulations, out_slurs, attention, cache_out0..31

The decoder expects 32 KV-cache tensors because it has decoder_depth = 8 layers with 4 cache tensors each.

Normalisation for the encoder input is (pixel / 255 - 0.7931) / 0.1738, with a single channel.

Provenance

These weights come from homr's training pipeline, which builds on:

homr itself is licensed AGPL-3.0. Please confirm the licensing terms of the upstream model weights before redistributing them.

Updating

Re-export from homr's training pipeline, publish to the onnx_checkpoints release tag, then mirror here. Never change a model in only one of the two places.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support