| --- |
| license: mit |
| library_name: pytorch |
| pipeline_tag: image-to-text |
| tags: [optical-music-recognition, omr, music, jianpu, sargam, image-to-text] |
| --- |
| |
| # omr-synth — a small OMR model for Western staff, Jianpu (简谱) & Sargam |
|
|
| A ~26M-parameter image→sequence Transformer (CNN encoder + autoregressive decoder) that reads a page of |
| music notation and writes out a compact symbolic score (OMRDSL). Trained **entirely on synthetic data, |
| overnight, on a Mac mini (M4, Apple-Silicon MPS)** — no cloud GPU. **A weekend learning project: early and |
| honest, not production OMR.** |
|
|
| - 💻 Code: https://github.com/abhid1234/omr-synth |
| - ▶️ Interactive demo: https://omr-synth.vercel.app |
| - 📚 Dataset: https://huggingface.co/datasets/abhid1234/omr-synth |
|
|
| ## What it does |
| Reads **Western staff, Jianpu, and Sargam** — three ways of writing music — into ONE shared symbolic target. |
| General vision models can't read Jianpu/Sargam and there's almost no labeled data for them; a render-forward |
| engine manufactures it (write the score first — that's the label — then render it into an image). |
|
|
| ## Honest held-out results (pages it never saw) |
| Scoped to **level-0** (clean, monophonic) to be learnable in one overnight run: |
|
|
| | Notation | Token accuracy | Content-only (the actual notes) | |
| |---|---|---| |
| | Western | ~30% | ~11% | |
| | Jianpu | ~28% | ~9% | |
| | Sargam | ~28% | ~8% | |
| | **Overall** | **~29%** | **~9%** | |
|
|
| *Token accuracy* counts every output token, including the easy structural scaffolding (clef, meter, |
| bar/voice layout) it almost always gets right. *Content-only* counts just the actual notes and durations — |
| the honest "can it read the music" number: about one note in ten. Exact-match ~0%. Balanced across all three. |
|
|
| ## Architecture |
| CNN stem (1→64→128→384) + 6-layer Transformer encoder (d=384) reading a grayscale, ink-positive, |
| aspect-padded page; 6-layer autoregressive Transformer decoder emitting OMRDSL. ~26M params |
| (`src/model/omr.py`). This checkpoint is the best of the overnight run (gentler-LR rerun, epoch 7). |
|
|
| ## Usage |
| ```bash |
| git clone https://github.com/abhid1234/omr-synth && cd omr-synth |
| pip install torch torchvision |
| python predict.py --checkpoint omr-synth-e7.pt --image page.png |
| ``` |
|
|
| ## Limitations |
| Level-0 only (no polyphony/dynamics/ornaments, no real manuscripts). Reads content partially. Training was |
| unstable early (fixed with a warmup+cosine schedule). The reusable contribution is the render-forward |
| **data engine**, not this early model. MIT · Abhi Das. |
|
|