- midisimx
- Greatly improved, enhanced, and streamlined fork of midisim for calculating, searching, and analyzing MIDI-to-MIDI similarity at scale
- What's new
- π midisimx vs midisim β comparison table
- Main features
- Pre-trained model
- Pre-computed embeddings sets
- Installation
- Basic use guide
- Creating custom MIDI corpus embeddings
- Music discovery pipeline
- Install midisimx and discovermidi PyPI packages
- Download and unzip Discover MIDI Dataset
- Prepare midisimx model and desired corresponding embeddings set
- Create Master MIDI dataset directory and upload your source/master MIDIs in it
- Initialize midisimx, download and load midisimx model and embeddings set
- Create Master MIDI dataset files list
- Launch the search
- MIDI Representation Encoding
- Documentation
- Project Structure
- Limitations
- Citations
- Greatly improved, enhanced, and streamlined fork of midisim for calculating, searching, and analyzing MIDI-to-MIDI similarity at scale
midisimx
Greatly improved, enhanced, and streamlined fork of midisim for calculating, searching, and analyzing MIDI-to-MIDI similarity at scale
What's new
π midisimx vs midisim β comparison table
| Feature / Change | midisimx | midisim |
|---|---|---|
| Model Architecture | β One unified larger model | Two smaller models |
| Model Dimension | π₯ 768 | 512 |
| Model Depth | π₯ 16 layers | 16 + 8 layers |
| Attention Heads | π₯ 12 heads | 8 heads |
| Training Corpus Size | π 3M+ filtered & processed MIDIs | 1M+ raw MIDIs |
| MIDI Event Representation | πΌ start-time Β· note/chord Β· pitch Β· duration | start-time Β· duration Β· pitch |
| Codebase Quality | π Improved, extended, modernized | Older original codebase |
| Overall Quality | β Major upgrade | Baseline |
Main features
- Ultra-fast and flexible GPU/CPU MIDI-to-MIDI similarity calculation, search and analysis
- Quality pre-trained model and pre-computed embeddings sets
- Stand-alone, versatile, and extensive codebase for general or custom MIDI-to-MIDI similarity tasks
- Full cross-platform compatibility and support
Pre-trained model
midisimx_trained_model_14391_steps_0.255_loss_0.9036_acc.pth- Unified and fast large model for a nuanced embeddings generation. Download checkpoint from Hugging Face
This model was trained on full Discover Piano dataset for 2 complete epochs
Pre-computed embeddings sets
Weighted Mean Pool Embeddings (1-2-1-2)
- These embeddings put more emphasis on pitches and chords (weights == 2) with start-times and durations left as is (weights == 1)
discover_midi_dataset_3267574_clean_midis_embeddings_1_2_1_2_weighted_cc_by_nc_sa.npy - 3267574 clean MIDIs weighted embeddings from Discover MIDI Dataset for large scale similarity search and analysis tasks
lakh_midi_dataset_17203_clean_midis_embeddings_1_2_1_2_weighted_cc_by_nc_sa.npy - 17203 LAKH clean_midi subset weighted embeddings tailored primarily for artist/song identification tasks
Source MIDI datasets: Discover MIDI Dataset and LAKH MIDI Dataset
Similarity search output samples
midisimx-similarity-search-output-samples-1-2-1-2-weighted-CC-BY-NC-SA.zip - ~182k+ MIDIs filtered by weighted midisimx music discovery pipeline
Source MIDI dataset: Discover MIDI Dataset
Installation
midisimx PyPI package (for general use)
!pip install -U midisimx
x-transformers 2.3.1 (for raw/custom tasks)
!pip install x-transformers==2.3.1
Basic use guide
General use example
# ================================================================================================
# Initalize midisimx
# ================================================================================================
# Import main midisimx module
import midisimx
# ================================================================================================
# Prepare midisimx embeddings
# ================================================================================================
# Option 1: Download sample pre-computed embeddings corpus from Hugging Face
emb_path = midisimx.download_embeddings()
# Option 2: use custom pre-computed embeddings corpus
# See custom embeddings generation section of this README for details
# emb_path = './custom_midis_embeddings_corpus.npy'
# Load downloaded embeddings corpus
corpus_midi_names, corpus_emb = midisimx.load_embeddings(emb_path)
# ================================================================================================
# Prepare midisimx model
# ================================================================================================
# Option 1: Download main pre-trained midisimx model from Hugging Face
model_path = midisimx.download_model()
# Option 2: Use main pre-trained midisimx model included in midisimx PyPI package
# model_path = midisimx.get_package_models()[0]['path']
# Load midisimx model
model, ctx, dtype = midisimx.load_model(model_path)
# ================================================================================================
# Prepare source MIDI
# ================================================================================================
# Load source MIDI
input_toks_seqs = midisimx.midi_to_tokens('Come To My Window.mid')
# ================================================================================================
# Calculate and analyze embeddings
# ================================================================================================
# Compute source/query embeddings
query_emb = midisimx.get_embeddings_bf16(model,
input_toks_seqs,
device=torch.device('cuda'),
pooling='weighted_mean',
# The following arg is optional but recommended if
# you want to make an emphasis on music
# Remove it for overall/general similarity searches
# PLEAE NOTE: You must enable it if you are using
# included pre-computed weighted embeddings
token_type_weights={(128, 256): 2, # Pitches weight
(384, 718): 2 # Chords weight
},
)
# Calculate cosine similarity between source/query MIDI embeddings and embeddings corpus
idxs, sims = midisimx.cosine_similarity_topk(query_emb, corpus_emb)
# ================================================================================================
# Processs, print and save results
# ================================================================================================
# Convert the results to sorted list with transpose values
idxs_sims_tvs_list = midisimx.idxs_sims_to_sorted_list(idxs, sims)
# Print corpus matches (and optionally) convert the final result to a handy list for further processing
corpus_matches_list = midisimx.print_sorted_idxs_sims_list(idxs_sims_tvs_list, corpus_midi_names, return_as_list=True)
# ================================================================================================
# Copy matched MIDIs from the MIDI corpus for listening and further evaluation and analysis
# ================================================================================================
# Copy matched corpus MIDI to a desired directory for easy evaluation and analysis
out_dir_path = midisimx.copy_corpus_files(corpus_matches_list)
# ================================================================================================
Raw/custom use example
import torch
from x_transformers import TransformerWrapper, Encoder
# Original model hyperparameters
SEQ_LEN = 3072
MASK_IDX = 718 # Use this value for masked modelling
PAD_IDX = 719 # Model pad index
VOCAB_SIZE = 720 # Total vocab size
MASK_PROB = 0.15 # Original training mask probability value (use for masked modelling)
DEVICE = 'cuda' # You can use any compatible device or CPU
DTYPE = torch.bfloat16 # Original training dtype
# Official main midisimx model checkpoint name
MODEL_CKPT = 'midisimx_trained_model_14391_steps_0.255_loss_0.9036_acc.pth'
# Model architecture using x-transformers
model = TransformerWrapper(
num_tokens = VOCAB_SIZE,
max_seq_len = SEQ_LEN,
attn_layers = Encoder(
dim = 768,
depth = 16,
heads = 12,
rotary_pos_emb = True,
attn_flash = True,
),
)
model.load_state_dict(torch.load(MODEL_CKPT, map_location=DEVICE))
model.to(DEVICE)
model.eval()
# Original training autoxast setup
autocast_ctx = torch.amp.autocast(device_type=DEVICE, dtype=DTYPE)
Creating custom MIDI corpus embeddings
# ================================================================================================
# Load main midisimx module
import midisimx
# Import helper modules
import os
import tqdm
# ================================================================================================
# Call included TMIDIX module through midisimx to create MIDI files list
custom_midi_corpus_file_names = midisimx.TMIDIX.create_files_list(['./custom_midi_corpus_dir/'])
# ================================================================================================
# Create two lists: one with MIDI corpus file names
# and another with MIDI corpus tokens representations suitable for embeddings generation
midi_corpus_file_names = []
midi_corpus_tokens = []
for midi_file in tqdm.tqdm(custom_midi_corpus_file_names):
midi_corpus_file_names.append(os.path.splitext(os.path.basename(midi_file))[0])
midi_tokens = midisimx.midi_to_tokens(midi_file, transpose_factor=0, verbose=False)[0]
midi_corpus_tokens.append(midi_tokens)
# It is highly recommended to sort the resulting corpus by tokens sequence length
# This greatly speeds up embeddings calculations
sorted_midi_corpus = sorted(zip(midi_corpus_file_names, midi_corpus_tokens), key=lambda x: len(x[1]))
midi_corpus_file_names, midi_corpus_tokens = map(list, zip(*sorted_midi_corpus))
# ================================================================================================
# Now you are ready to generate embeddings as follows:
# ================================================================================================
# Load main midisimx model
model, ctx, dtype = midisimx.load_model(verbose=False)
# Generate MIDI corpus embeddings
midi_corpus_embeddings = midisimx.get_embeddings_bf16(model, midi_corpus_tokens, verbose=False)
# ================================================================================================
# Save generated MIDI corpus embeddings and MIDI corpus file names in one handy NumPy file
midisimx.save_embeddings(midi_corpus_file_names,
midi_corpus_embeddings,
verbose=False
)
# ================================================================================================
# You now can use this saved custom MIDI corpus NumPy file with midisimx.load_embeddings()
# and the rest of the pipeline outlined in the general use section above
Music discovery pipeline
Here is a complete MIDI music discovery pipeline example using midisimx and Discover MIDI Dataset
Install midisimx and discovermidi PyPI packages
!pip install -U midisimx
!pip install -U discovermidi
Download and unzip Discover MIDI Dataset
import discovermidi
from discovermidi import fast_parallel_extract
discovermidi.download_dataset()
fast_parallel_extract.fast_parallel_extract()
Prepare midisimx model and desired corresponding embeddings set
model_ckpt = 'midisimx_trained_model_14391_steps_0.255_loss_0.9036_acc.pth'
model_depth = 16
embeddings_file = 'discover_midi_dataset_3267574_clean_midis_embeddings_1_2_1_2_weighted_cc_by_nc_sa.npy'
Create Master MIDI dataset directory and upload your source/master MIDIs in it
import os
os.makedirs('./Master-MIDI-Dataset/', exist_ok=True)
Initialize midisimx, download and load midisimx model and embeddings set
# Import main midisimx module
import midisimx
# Download embeddings from Hugging Face
emb_path = midisimx.download_embeddings(filename=embeddings_file)
# Load downloaded embeddings corpus
corpus_midi_names, corpus_emb = midisimx.load_embeddings(embeddings_path=emb_path)
# Download midisimx model from Hugging Face
model_path = midisimx.download_model(filename=model_ckpt)
# Load midisimx model
model, ctx, dtype = midisimx.load_model(model_path,
depth=model_depth
)
Create Master MIDI dataset files list
filez = midisimx.TMIDIX.create_files_list(['./Master-MIDI-Dataset/'])
Launch the search
import os
import tqdm
for fa in tqdm.tqdm(filez):
# Load source MIDI
input_toks_seqs = midisimx.midi_to_tokens(fa, verbose=False)
if input_toks_seqs:
# ================================================================================================
# Calculate and analyze embeddings
# ================================================================================================
# Compute source/query embeddings
query_emb = midisimx.get_embeddings_bf16(model,
input_toks_seqs,
device=torch.device('cuda'),
pooling='weighted_mean',
# The following arg is optional but recommended if
# you want to make an emphasis on music
# Remove it for overall/general similarity searches
# PLEAE NOTE: You must enable it if you are using
# included pre-computed weighted embeddings
token_type_weights={(128, 256): 2, # Pitches weight
(384, 718): 2 # Chords weight
},
verbose=False,
show_progress_bar=False
)
# Calculate cosine similarity between source/query MIDI embeddings and embeddings corpus
idxs, sims = midisimx.cosine_similarity_topk(query_emb,
corpus_emb,
verbose=False
)
# ================================================================================================
# Processs, print and save results
# ================================================================================================
# Convert the results to sorted list with transpose values
idxs_sims_tvs_list = midisimx.idxs_sims_to_sorted_list(idxs, sims)
# Print corpus matches (and optionally) convert the final result to a handy list for further processing
corpus_matches_list = midisimx.print_sorted_idxs_sims_list(idxs_sims_tvs_list,
corpus_midi_names,
return_as_list=True
)
# ================================================================================================
# Copy matched MIDIs from the MIDI corpus for listening and further evaluation and analysis
# ================================================================================================
# Copy matched corpus MIDI to a desired directory for easy evaluation and analysis
out_dir_path = midisimx.copy_corpus_files(corpus_matches_list,
corpus_midis_dirs=['./Discover-MIDI-Dataset/MIDIs/'],
main_output_dir='Output-MIDI-Dataset',
sub_output_dir=os.path.splitext(os.path.basename(fa))[0],
verbose=False
)
# ================================================================================================
MIDI Representation Encoding
midisimx uses a compact, eventβstructured token format that lets the model understand timing, harmony, melody, and rhythm with minimal overhead.
Each event is encoded in a strict order, and notes and chords share the same structureβchords simply contain multiple pitchβduration pairs.
| Token Type | Range | Meaning | Notes |
|---|---|---|---|
| Delta StartβTime | 0β127 | Time since previous event | Encodes rhythmic spacing |
| Note/Chord Token | 384β717 | Semitone or chord class | 384β395 β 12 semitones; 396β716 β 321 chords |
| Pitch | 128β255 | MIDI pitch (0β127) | One per note; multiple for chords |
| Duration | 256β383 | Note length | One per pitch |
Event Structure
Notes
A note event always has four tokens:
[deltaβstart, note-token, pitch, duration]
Chords
A chord event starts with the same two tokens, but then includes multiple (pitch, duration) pairs:
[deltaβstart, chord-token, pitch, duration, pitch, duration, pitch, duration, ...]
This allows encoding triads, extended chords, clusters, or any multiβnote harmony.
Sample Encoded Sequence
Below is a real midisimx token sequence excerpt, formatted for readability.
Events are grouped to show how notes and chords appear:
[0, 643, 193, 321]
[186, 321, 179, 325]
[16, 391, 195, 265]
[9, 391, 195, 298]
[16, 387, 191, 272]
[16, 387, 191, 266]
[8, 689, 193, 323, 186, 321] β chord (two pitchβduration pairs)
[1, 386, 178, 323]
[15, 391, 195, 265]
[9, 391, 195, 298]
[16, 387, 191, 283]
[24, 711, 196, 321, 186, 321] β chord
[1, 384, 176, 320]
[15, 391, 195, 296]
[9, 389, 193, 297]
[23, 387, 191, 273]
[8, 391, 195, 315]
[9, 707, 183, 323, 174, 324] β chord
[50, 391, 195, 274]
...
You can clearly see:
- Deltaβtimes drive the rhythm
- Chord tokens (β₯396) introduce multiβpitch structures
- Single notes β 4 tokens
- Chords β 2 + (pitch, duration) Γ N tokens
Documentation
midisimx API Reference
midisimx API Functions Index
Legacy midisimx API Reference
Project Structure
midisimx/ # Project root
βββ LICENSE # Apache-2.0 license text
βββ MANIFEST.in # Setuptools manifest β package-data inclusion rules for sdist/wheel
βββ README.md # Main project README β features, usage guides, links, citations
βββ midisimx/ # The installable Python package
β βββ API_REFERENCE.md # This document β complete public API reference
β βββ MIDI.py # LEGACY β original parent of TMIDIX; unused, kept for reference/posterity
β βββ README.md # Package README (PyPI landing page)
β βββ TMIDIX.py # TMIDIX MIDI parsing/processing suite; re-exported as midisimx.TMIDIX
β βββ artwork/ # Project images
β β βββ Project-Los-Angeles.png # Project Los Angeles logo
β β βββ README.md # Artwork notes and credits
β β βββ Tegridy-Code-2026.png # Tegridy Code 2026 branding image
β β βββ midisimx.png # Project banner (embedded in READMEs)
β βββ crossmodal_mapper.py # Non-ML closed-form bi-directional cross-modal embedding mapper (procrustes/ridge/cca)
β βββ embeddings/ # Bundled pre-computed embeddings
β β βββ README.md # Notes on bundled embeddings sets
β β βββ lakh_midi_dataset_17209...... # Tiny 128-dim weighted (1-2-1-2) embeddings for 17 209 clean LAKH MIDIs β pairs with the bundled tiny model
β βββ helpers.py # Utilities β bundled assets listing, MIDI normalization, file hashing, apt install
β βββ instrumentation_similarity.py # Deterministic timbre-aware GM instrumentation similarity scoring
β βββ ldmb.py # LDMB β mmap-backed binary storage for large lists of dicts (lazy reads, byte-level merges)
β βββ memmap.py # Single-file memmap storage for paired names + float32 embeddings
β βββ midi_to_colab_audio.py # AUX (optional) β renders MIDIs to audio via fluidsynth + SF2 soundfont banks
β βββ midisimx.py # CORE β model/embeddings I/O, MIDIβtokens, embedding computation, similarity search
β βββ models/ # Bundled model checkpoints
β β βββ README.md # Notes on bundled models
β β βββ midisimx_tiny_trained_model... # Tiny 6.49M-param Transformer checkpoint (14 401 steps Β· 0.5146 loss Β· 0.8202 acc)
β βββ pca_reduce.py # Streaming, GPU-accelerated PCA reduction (PCAReductor, PCAReductionResult)
β βββ x_transformer_2_3_1.py # CORE (models) β vendored, stand-alone x-transformers v2.3.1 by lucidrains
βββ pyproject.toml # PEP 621 packaging metadata β version, dependencies, PyPI URLs, classifiers
Legend:
- CORE β required by the main similarity pipeline.
- AUX β optional convenience module; requires
fluidsynth(installable viamidisimx.helpers.install_apt_package('fluidsynth')) and SF2 banks; audio rendering only. - LEGACY β not used anywhere in the project; provided for reference, convenience, and posterity.
- The vendored
x_transformer_2_3_1.pymakes the core pipeline independent of the PyPIx-transformerspackage (which is only needed for raw/custom tasks).
Limitations
- Current code and models support only MIDI music elements similarity (start-times, durations, pitches and chords)
- MIDI channels and velocities are not currently supported due to practicality considerations
- Current model is limited by 3k sequence length (~1000 MIDI music notes) so long-running MIDIs can only be analyzed in chunks
Citations
@misc{project_los_angeles_2026,
author = { Project Los Angeles and Tegridy Code },
title = { midisimx (Revision cfed861) },
year = 2026,
url = { https://huggingface.co/projectlosangeles/midisimx },
doi = { 10.57967/hf/10032 },
publisher = { Hugging Face }
}
@misc{project_los_angeles_2026,
author = { Project Los Angeles and Tegridy Code },
title = { midisimx-embeddings (Revision 0af7bbc) },
year = 2026,
url = { https://huggingface.co/datasets/projectlosangeles/midisimx-embeddings },
doi = { 10.57967/hf/10082 },
publisher = { Hugging Face }
}
@misc{project_los_angeles_2026,
author = { Project Los Angeles and Tegridy Code },
title = { midisimx-samples (Revision 3c28df7) },
year = 2026,
url = { https://huggingface.co/datasets/projectlosangeles/midisimx-samples },
doi = { 10.57967/hf/10085 },
publisher = { Hugging Face }
}
@misc{project_los_angeles_2025,
author = { Project Los Angeles },
title = { Discover-MIDI-Dataset (Revision 0eaecb5) },
year = 2025,
url = { https://huggingface.co/datasets/projectlosangeles/Discover-MIDI-Dataset },
doi = { 10.57967/hf/7361 },
publisher = { Hugging Face }
}
@phdthesis{raffel2016learning,
author = { Colin Raffel },
title = { Learning-Based Methods for Comparing Sequences, with Applications to Audio-to-{MIDI} Alignment and Matching },
school = { Columbia University },
year = { 2016 },
url = { https://colinraffel.com/projects/lmd/ }
}
