LiveSynth

Weights for LiveSynth: A Streaming Neural Synthesizer for Instrument Cloning and Text-to-Instrument (Kyungsu Kim, Yejin Kim, Kyogu Lee, Seoul National University). Use them with the livesynth Python package, which downloads this repository automatically:

from livesynth import LiveSynth
synth = LiveSynth.from_pretrained()          # repo_id="KyungsuKim/LiveSynth"
audio = synth.render("song.mid", synth.embed_audio("reference.wav"))
File Contents
generator.safetensors 190 M-parameter feedback-free causal Transformer (Linear weights bf16, embeddings and norms fp32)
decoder.safetensors Causal Vocos latent decoder of the VAE codec (fp32) and latent standardisation statistics
text_align.safetensors Optional orthogonal Procrustes map (512 x 512) toward the audio embeddings the model was trained on (embed_text(..., align="procrustes"); the default uses the raw text embedding)
presets.safetensors CLAP timbre embeddings of 53 held-out NSynth instruments (names in config.json); each is the embedding of one reference recording in references/
references/*.flac The 10-s reference recordings behind the presets (NSynth notes, 48 kHz), fetched on demand by synth.preset_audio(name)
config.json Model hyper-parameters and provenance

The timbre encoder is the public LAION-CLAP checkpoint lukewys/laion_clap/music_audioset_epoch_15_esc_90.14.pt, downloaded separately.

The weights are released under the MIT License, like the code.

Downloads last month
25
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support