DX7 timbre: encoder and catalogs
A small encoder that maps one note to a 128-dimensional unit vector, and two catalogs of Yamaha DX7 voices, one vector per voice, that name the sound by cosine. These are the files @audio/neural-timbre (npm) reads: each package version pins a commit of this repository and checks every file against its SHA-256.
| File | Size | SHA-256 |
|---|---|---|
encoder.int8.onnx |
1.87 MB (1,871,901 bytes) | aaa9fae0337325fc9ae0f15c45f711073255c09264a31a9779da635ec21b8484 |
catalogs/dx7-factory.json |
0.21 MB (209,548 bytes) | 7ab9d48c1b90a4ec04af3aeb63737352d63f959db8b1459d6b238cea365a0900 |
catalogs/dx7-all.json |
7.9 MB (7,894,349 bytes) | 0dd0fd646e8aef7e417151e8ae76b5c2bec7fde3cb531e01a3bc4e5b964e3671 |
SHA256SUMS lists the same.
import { createMatcher } from '@audio/neural-timbre'
let matcher = await createMatcher() // dx7-factory; createMatcher('dx7-all') for every voice
let [best] = await matcher.match(note, { sampleRate: 44100 }) // one note
best.name // 'E.PIANO 1'
best.cartridge // 'ROM1A'
best.score // cosine, -1 to 1
matcher.free()
Encoder
encoder.int8.onnx: four convolution stages (32, 64, 128, 256 channels), a 1×1 convolution to 96 channels, the mean
and maximum over time, two linear layers to a unit vector; 1.85M parameters, weights rounded to int8 per output row.
| name | shape | ||
|---|---|---|---|
| input | features |
[n, 1, 96, 128] | constant-Q power spectrogram of one note from its onset: 96 bins 12 an octave from 41.2 Hz, 128 frames 512 samples apart at 44.1 kHz, dB below the loudest cell with an 80 dB floor, scaled to [−1, 1] (the package's features.js) |
| output | embedding |
[n, 128] | the unit vector a catalog ranks against |
| output | algorithm |
[n, 32] | the voice's DX7 algorithm, read linearly from the vector (weak: 37.1% right on held-out queries) |
Trained once with scripts/train.py of the package on 26,206 DX7 voices, eight rendered views each (any note 33–99,
velocity 20–127, held or staccato, two thirds through an augmentation chain: chorus, EQ, reverb, another voice or drums
underneath, noise, compression, MP3 or Opus): supervised contrastive loss (Khosla et al. 2020) and a margin softmax
against one proxy per voice (CosFace, Wang et al. 2018). The renders come from
@audio/synth-dx7, bit-exact to Google's Music
Synthesizer for Android engine. ONNX output matches PyTorch within 2·10⁻⁶ in Node and in Chromium on WebGPU. The recipe
is gpu-font's, applied to sound.
Catalogs
Each voice rendered at MIDI 36, 48, 60, 72 and 84, velocities 64 and 112, held 1 s; the unit mean of the ten embeddings, stored with one scale per row as 4-bit integers. Each catalog carries the encoder's and the features' SHA-256; the package refuses a catalog bound to another encoder.
| Catalog | Voices | From |
|---|---|---|
dx7-factory |
988 | Yamaha's own cartridges: ROM1A–ROM4B (223 distinct sounds) and VRC-101 to VRC-112 (765) |
dx7-all |
29,840 | every audible voice of every source below |
An entry holds a voice's name, its cartridge (ROM1A for Yamaha's, else the file's path in its archive), index
(its number in that file), source, algorithm (its DX7 algorithm number, 1–32) and id (the first 16 hex digits of a
SHA-1 of the parameters that shape its sound, so voices that render the same samples share it). Names and vectors only: no
sysex and no other voice parameter. To play a voice, load its cartridge from its source and take the voice by number.
Benchmark
From the package's README, frozen before training: 2,925 voices from 247 cartridges never seen in training or checkpoint selection, one query each at a note and velocity no catalog reference uses.
On 2,925 voices from cartridges it never saw, one clean note names the voice 75.0% of the time and puts it in the top five 92.6%; through the augmentation chain, 56.4% and 79.0%. Those figures use 8-bit catalog vectors; the 4-bit vectors hosted here:
| held-out clean top-1 | top-5 | augmented top-1 | top-5 | all voices clean top-1 | top-5 | |
|---|---|---|---|---|---|---|
| int8 encoder, int8 catalog | 75.0% | 92.6% | 56.4% | 79.0% | 45.0% | 75.8% |
| int8 encoder, int4 catalog: these files | 74.0% | 92.1% | 56.4% | 79.0% | 44.3% | 75.1% |
"All voices": the right voice among all 29,840, the training voices included. Per condition (reverb, EQ, compression, chorus, codecs, noise, another voice underneath, drums) and the method: the package's README, Benchmark.
Intended use
Use when: naming which preset of a known synth plays a recorded note (a DX7 factory sound, a cartridge voice), as the seed of a rebuild in the original synth; indexing your own presets into a catalog and searching them by sound.
Not for: chords, layered or heavily processed parts without separating them first (another voice under the note at 0 dB: 28.6% top-1); synths no catalog covers (the closest preset comes back with a lower score, not a refusal: there is no calibrated threshold); real-time use.
Sources
The voices the encoder trained on and the catalogs name come from these sources:
source |
Sounds in dx7-all |
What, from where | Terms stated |
|---|---|---|---|
yamaha-factory |
223 | Yamaha's ROM1A–ROM4B cartridges, from yamahablackboxes.com, byte-identical to rohandrape.net's copies | none; voice data © Yamaha Corporation, 1983 |
yamaha-vrc |
765 | Yamaha's VRC-101 to VRC-112 cartridges, sides A and B, from the same site | none; © Yamaha and the named programmers |
greymatter |
56 | Grey Matter Response E! card disks 2, 5 and 7, from the same site | none; © Grey Matter Response |
alltheweb |
28,571 | DX7 All The Web: 13,192 sysex files gathered from the web by Bobby Blues | "public domain material" gleaned from the Net, per the collector, who removes material at its owners' request; it includes commercial banks |
library |
225 | BlackWinny's Yamaha-DX7-patch-library v1.0 (2015), commit 6b17188 | CC0-1.0, marked by its uploader, who does not own every bank |
No source licenses all its voices for this use. The catalogs follow gpu-font's policy for DaFont and Adobe Fonts: names and vectors only, published by the maintainer's decision, and a source comes out at its owner's request.
Removal: if you own voices listed here and want them out, open a discussion on this repository or an issue at github.com/audiojs/neural, naming the source or cartridges. Their entries are removed from the catalogs and a new revision published.
Licence
The encoder weights (encoder.int8.onnx): MIT, Copyright (c) Dmitry Iv, as
@audio/neural.
The catalogs: the vectors are this encoder's output on renders of the voices above. The voice names, cartridge names and file paths are the ones their authors and collectors gave; they are listed to identify the voices and stay theirs. No rights over them are claimed or granted here.
References
J. C. Brown, "Calculation of a constant Q spectral transform," JASA 89(1), 1991. · J. C. Brown, M. S. Puckette, "An efficient algorithm for the calculation of a constant Q transform," JASA 92(5), 1992. · P. Khosla et al., "Supervised contrastive learning," NeurIPS 2020. · H. Wang et al., "CosFace: large margin cosine loss for deep face recognition," CVPR 2018. · R. Levien, Music Synthesizer for Android (the DX7 engine, Apache-2.0). · gpu-font (the recipe).