Ensomi R2

Ensomi R2 turns a song into an osu!mania 4K chart. Give it an audio file and a star rating from 2 to 6, and a few seconds later it writes a complete .osz that opens in osu!. It is the first Ensomi release that works from audio alone.

SaiGetsu (Cut Ver.), Otokaze, 3.5 stars: a chart generated by Ensomi R2

Yellow notes are taps and cyan notes are long notes; each panel reads bottom to top, then continues in the panel to its right. The gallery shows four songs at seed 0 with their generated .osu files. They illustrate behaviour; they are not a quality benchmark.

The full report, with the project's aims, the decisions behind R2 and the evidence for every claim below, is docs/research/r2_system.md in the code repository.

What it is for

Ensomi wants listening to music, playing it and creating for it to become one activity: charts that follow music already playing around you, candidate choreography for a mapper's selected section, and practice charts built around a movement a player wants to train. Those need charts that are playable at the requested difficulty, follow the music's rhythm, keep one character through the whole song, and respond to requests.

R2 is the first step: a whole song, generated while you wait. It does not yet play along in real time or regenerate one section of an existing chart.

How it works

audio โ”€> timing system โ”€> rows model โ”€> corpus prior โ”€> R2 with controls โ”€> .osz
         beat grid        head rows     chart-level      lanes, taps,
                                        targets          long notes
  • Timing system. The BeatThis beat tracker finds the beats; each tempo segment is then refitted to the song's own note onsets, which sit much closer to where mappers put their beats. No trained weights of its own.
  • Rows model (rows/model.pt, 2.1M parameters). Beat by beat, it chooses where notes start from what it hears at each position in the beat, and its copy heads repeat earlier figures where the music sounds the same.
  • Corpus prior (prior.json). From the 64 human charts whose head rows look most like these, it picks the chart's character: long-note level, chord density, difficulty. This stands in for the mapper's choice for the whole chart.
  • R2 (r2/model.pt, 2.4M parameters). It arranges the lanes, taps and long notes on every head row; one decision per row places its new notes and the releases of the long notes it closes. A restoring force holds the chart's character to the end, playability limits measured on human charts (calibration.json) keep it playable, and the requests tilt its choices. R2 does not hear the audio.

What you can ask for

Request Status
Star rating for the whole song works in blind screens
Jumpstream and chordstream sections work in blind screens
Stream and jack sections reach their target in measurement; not screened on their own
Chordjack sections not offered as working
Long-note amount (`--ln tap half
Section difficulty in stars not yet screened

What has been checked

Every chart passes legality and replay checks and has whole-millisecond times. Quality was judged by one experienced player in blind pairs, one generation seed per side, so a single pair is weak evidence:

  • the new rows model beat the earlier head generator 10 to 3, one undecided;
  • whole-song difficulty was called the right way in both pairs;
  • the human chart was chosen over R2's in all three reference pairs.

On a holdout of 519 human charts, 85 % have 90 % of their heads within 20 ms of the timing system's grid. The four example charts regenerate note for note from this repository's files.

Known limits: long notes are the weakest part; easier sections land short of their target; long-song stability rests on a decode-time rule; every verdict is one listener's.

Quick start

Checked on macOS arm64 (Apple M5, Python 3.10.20, torch 2.11.0, CPU, one thread). Linux and CUDA (--extra cuda) were not run.

git clone https://github.com/ensomi-labs/ensomi-model.git
cd ensomi-model
git checkout --detach r2-1.0
uv sync --python 3.10 --extra mps
hf download sed-i/ensomi-r2 --local-dir artifacts/hf-r2

uv run --python 3.10 --extra mps python -m ensomi_model.system.make \
  --model-dir artifacts/hf-r2 --audio /path/to/song.mp3 --star 3.5 \
  --title "Song title" --artist "Artist" --seed 0 --out artifacts/my-chart.osz

A 2-3.3 minute song takes 8-12 s. The first run also downloads BeatThis's final0 checkpoint (77 MB) from its authors. Decoding audio other than WAV needs ffmpeg. --help lists every request, for example --stream 60000,90000,0.7,jumpstream.

Files

File Bytes SHA-256
rows/model.pt 9,622,005 46b6ca37870120fda4d5de632949921fc26e020cf80abe44b4ba7f6f9557dda4
r2/model.pt 9,655,631 8dedefe4b11d238ad0b3186e6dfefe31f52cb333091aafe2e8efda8ad2ee3c6f
prior.json 1,958,310 be5741d520120aaeb7ecf9126c2c12401015095b107b025584802f8e715590af
calibration.json 443,193 d39cd16fda451163d00d70cd6a913eb4bf64326ecd577feab08c25d8c22455ab
scopes.parquet 2,447,837 61f8648523cddb86e7316812b46896410f12b1a6ae179dcbf082c8b8615e3134
system.json 209 333d1c59a5ca199e07369b980992c2bb9683851246807426f1152260502e9718
LICENSE 20,850 e66c269d4819aaab34b49ef5220c4ddab6756f21bb5180761a4eb8561f2b7bbd

SHA256SUMS lists every file, including examples/ (four generated charts with their receipts) and previews/. r2/model.pt holds R2's weights and model configuration only; the example receipts record the hash of the training checkpoint, which also held optimizer state. The .pt files are PyTorch pickles, loaded by the code repository.

onnx/ holds the same networks as ONNX step graphs (float32, opset 17) for in-browser inference by the TypeScript runtime in ensomi-web: BeatThis final0, the rows encoder and step, and R2 predict, release and commit, with runtime.json (graph interfaces, carried state and decode constants). Replaying the same random draws, they reproduce the PyTorch charts of the four examples note for note.

Component Trained from Code tag
Timing system no training timing-1.0
Rows model commit cf753aa, 5,000 steps rows-1.0
R2 commit 3954031, 64M decisions; checkpoint at 56,000,574 r2-1.0

R2 and the rows model were trained on osu!'s ranked and loved 4K charts from 2 to 6 stars (11,368 charts of 4,167 songs). Neither those charts nor their audio are distributed here. The prior, calibration and scope table hold numbers derived from them, without titles, artists, mappers or file paths.

Licence

Everything in this repository is licensed CC BY-NC-SA 4.0: you may use, share and adapt it with credit, not for commercial purposes, and adaptations must carry the same licence, except onnx/beatthis.onnx, an ONNX export of BeatThis final0 (CPJKU), which keeps its MIT licence (onnx/LICENSE-beat_this). The Python system downloads BeatThis from its authors instead. The model code is AGPL-3.0-only; the browser runtime is Apache-2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support