DX7 timbre: encoder and catalogs

A small encoder that maps one note to a 128-dimensional unit vector, and two catalogs of Yamaha DX7 voices, one vector per voice, that name the sound by cosine. These are the files @audio/neural-timbre (npm) reads: each package version pins a commit of this repository and checks every file against its SHA-256.

File Size SHA-256
encoder.int8.onnx 1.87 MB (1,871,901 bytes) aaa9fae0337325fc9ae0f15c45f711073255c09264a31a9779da635ec21b8484
catalogs/dx7-factory.json 0.21 MB (209,548 bytes) 7ab9d48c1b90a4ec04af3aeb63737352d63f959db8b1459d6b238cea365a0900
catalogs/dx7-all.json 7.9 MB (7,894,349 bytes) 0dd0fd646e8aef7e417151e8ae76b5c2bec7fde3cb531e01a3bc4e5b964e3671

SHA256SUMS lists the same.

import { createMatcher } from '@audio/neural-timbre'

let matcher = await createMatcher()                           // dx7-factory; createMatcher('dx7-all') for every voice
let [best] = await matcher.match(note, { sampleRate: 44100 })  // one note
best.name       // 'E.PIANO 1'
best.cartridge  // 'ROM1A'
best.score      // cosine, -1 to 1
matcher.free()

Encoder

encoder.int8.onnx: four convolution stages (32, 64, 128, 256 channels), a 1×1 convolution to 96 channels, the mean and maximum over time, two linear layers to a unit vector; 1.85M parameters, weights rounded to int8 per output row.

name shape
input features [n, 1, 96, 128] constant-Q power spectrogram of one note from its onset: 96 bins 12 an octave from 41.2 Hz, 128 frames 512 samples apart at 44.1 kHz, dB below the loudest cell with an 80 dB floor, scaled to [−1, 1] (the package's features.js)
output embedding [n, 128] the unit vector a catalog ranks against
output algorithm [n, 32] the voice's DX7 algorithm, read linearly from the vector (weak: 37.1% right on held-out queries)

Trained once with scripts/train.py of the package on 26,206 DX7 voices, eight rendered views each (any note 33–99, velocity 20–127, held or staccato, two thirds through an augmentation chain: chorus, EQ, reverb, another voice or drums underneath, noise, compression, MP3 or Opus): supervised contrastive loss (Khosla et al. 2020) and a margin softmax against one proxy per voice (CosFace, Wang et al. 2018). The renders come from @audio/synth-dx7, bit-exact to Google's Music Synthesizer for Android engine. ONNX output matches PyTorch within 2·10⁻⁶ in Node and in Chromium on WebGPU. The recipe is gpu-font's, applied to sound.

Catalogs

Each voice rendered at MIDI 36, 48, 60, 72 and 84, velocities 64 and 112, held 1 s; the unit mean of the ten embeddings, stored with one scale per row as 4-bit integers. Each catalog carries the encoder's and the features' SHA-256; the package refuses a catalog bound to another encoder.

Catalog Voices From
dx7-factory 988 Yamaha's own cartridges: ROM1A–ROM4B (223 distinct sounds) and VRC-101 to VRC-112 (765)
dx7-all 29,840 every audible voice of every source below

An entry holds a voice's name, its cartridge (ROM1A for Yamaha's, else the file's path in its archive), index (its number in that file), source, algorithm (its DX7 algorithm number, 1–32) and id (the first 16 hex digits of a SHA-1 of the parameters that shape its sound, so voices that render the same samples share it). Names and vectors only: no sysex and no other voice parameter. To play a voice, load its cartridge from its source and take the voice by number.

Benchmark

From the package's README, frozen before training: 2,925 voices from 247 cartridges never seen in training or checkpoint selection, one query each at a note and velocity no catalog reference uses.

On 2,925 voices from cartridges it never saw, one clean note names the voice 75.0% of the time and puts it in the top five 92.6%; through the augmentation chain, 56.4% and 79.0%. Those figures use 8-bit catalog vectors; the 4-bit vectors hosted here:

held-out clean top-1 top-5 augmented top-1 top-5 all voices clean top-1 top-5
int8 encoder, int8 catalog 75.0% 92.6% 56.4% 79.0% 45.0% 75.8%
int8 encoder, int4 catalog: these files 74.0% 92.1% 56.4% 79.0% 44.3% 75.1%

"All voices": the right voice among all 29,840, the training voices included. Per condition (reverb, EQ, compression, chorus, codecs, noise, another voice underneath, drums) and the method: the package's README, Benchmark.

Intended use

Use when: naming which preset of a known synth plays a recorded note (a DX7 factory sound, a cartridge voice), as the seed of a rebuild in the original synth; indexing your own presets into a catalog and searching them by sound.

Not for: chords, layered or heavily processed parts without separating them first (another voice under the note at 0 dB: 28.6% top-1); synths no catalog covers (the closest preset comes back with a lower score, not a refusal: there is no calibrated threshold); real-time use.

Sources

The voices the encoder trained on and the catalogs name come from these sources:

source Sounds in dx7-all What, from where Terms stated
yamaha-factory 223 Yamaha's ROM1A–ROM4B cartridges, from yamahablackboxes.com, byte-identical to rohandrape.net's copies none; voice data © Yamaha Corporation, 1983
yamaha-vrc 765 Yamaha's VRC-101 to VRC-112 cartridges, sides A and B, from the same site none; © Yamaha and the named programmers
greymatter 56 Grey Matter Response E! card disks 2, 5 and 7, from the same site none; © Grey Matter Response
alltheweb 28,571 DX7 All The Web: 13,192 sysex files gathered from the web by Bobby Blues "public domain material" gleaned from the Net, per the collector, who removes material at its owners' request; it includes commercial banks
library 225 BlackWinny's Yamaha-DX7-patch-library v1.0 (2015), commit 6b17188 CC0-1.0, marked by its uploader, who does not own every bank

No source licenses all its voices for this use. The catalogs follow gpu-font's policy for DaFont and Adobe Fonts: names and vectors only, published by the maintainer's decision, and a source comes out at its owner's request.

Removal: if you own voices listed here and want them out, open a discussion on this repository or an issue at github.com/audiojs/neural, naming the source or cartridges. Their entries are removed from the catalogs and a new revision published.

Licence

The encoder weights (encoder.int8.onnx): MIT, Copyright (c) Dmitry Iv, as @audio/neural.

The catalogs: the vectors are this encoder's output on renders of the voices above. The voice names, cartridge names and file paths are the ones their authors and collectors gave; they are listed to identify the voices and stay theirs. No rights over them are claimed or granted here.

References

J. C. Brown, "Calculation of a constant Q spectral transform," JASA 89(1), 1991. · J. C. Brown, M. S. Puckette, "An efficient algorithm for the calculation of a constant Q transform," JASA 92(5), 1992. · P. Khosla et al., "Supervised contrastive learning," NeurIPS 2020. · H. Wang et al., "CosFace: large margin cosine loss for deep face recognition," CVPR 2018. · R. Levien, Music Synthesizer for Android (the DX7 engine, Apache-2.0). · gpu-font (the recipe).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support