omr-weights — Turkish (makam) optical music recognition
ONNX graphs for KomaVision, an optical music recognition model for Classical Turkish (makam) music: a photo or screenshot of sheet music in, notes out — including the microtonal accidentals (koma, küçük mücennep, bakiye, büyük mücennep) that Western OMR models have no vocabulary for.
Source code: https://github.com/EmreDikimen/Turkish_note_to_solfeggio_converter
What is here
Int8-quantized ONNX exports of a Donut-style vision-encoder-decoder (~143M parameters). The current
graphs are the project's Round 4 control checkpoint (r4-ctl-stage2-last), live since
2026-09-28.
| File | Role |
|---|---|
encoder_model.onnx |
image encoder |
decoder_model.onnx |
first decode step |
decoder_with_past_model.onnx |
subsequent steps, with KV cache |
Input is a 409×583 grayscale strip — one staff line's worth of music, not a whole page. The application slices a page into strips before decoding. Output is a LilyPond-flavoured token stream with AEU accidental tokens.
The application reads pages with these graphs on a CPU server (onnxruntime-node). The same decode
code also runs in a browser via onnxruntime-web — used in development and tests, not by the live
site's visitors. There is one decode implementation, shared.
Licence and attribution
Apache-2.0, inherited from the base model.
Fine-tuned from Flova/omr_transformer
(Apache-2.0) — a pretrained OMR transformer. That model is the reason this project did not need to
train an OMR system from scratch, and its licence and attribution travel with these weights as
Apache-2.0 §4 requires.
Training data
- Self-rendered synthetic strips. Turkish scores engraved by the project's own VexFlow renderer, with augmentation aimed at what users actually upload (screenshots more than photos). The pixels and the labels come from one code path, so a label can never disagree with its image.
- Real printed pages, hand-labelled, from freely-published Turkish score archives. These are used as training and evaluation data locally and are not redistributed here or anywhere.
- Score metadata for the synthetic renders derives from SymbTr (Karaosmanoğlu et al.), which is licensed CC BY-NC-SA 4.0 — attributed here accordingly.
No Western rehearsal data was used in fine-tuning; coverage comes from self-rendered Turkish strips.
Intended use and limits
Intended for reading Classical Turkish music notation. It is not a general-purpose OMR model — it was fine-tuned on a Turkish token vocabulary and will not do anything sensible with orchestral or piano scores.
Known limits, stated plainly:
- Accuracy on clean synthetic strips is effectively solved; accuracy on real printed pages is the open problem and is substantially lower. Treat every decode as a draft to be corrected — the application ships an editor for exactly that reason.
- On real pages the largest error groups are the key signature, note height and note length; the koma / küçük mücennep sharp distinction inside the key signature is the hardest accidental case.
- Long or dense staff lines can overrun the decoder's token budget.
- Handwritten manuscript is out of scope.
Citation
If the base model is useful to you, cite that first —
Flova/omr_transformer. For SymbTr:
M. K. Karaosmanoğlu, "A Turkish makam music symbolic database for music information retrieval: SymbTr", ISMIR 2012.
Model tree for Beyaban/omr-weights
Base model
Flova/omr_transformer