Stable Audio Open — Core ML

Text-to-Music, 2024

Text-to-music. Up to 11.9s stereo 44.1 kHz. Rectified flow DiT + T5 + Oobleck VAE.

Stable Audio Open demo

Core ML conversion of stabilityai/stable-audio-open-small for on-device inference on iPhone, iPad and Mac. Converted with coremltools; the packages are stateless, so all sequencing and buffering lives in your Swift code.

Task text to audio
Upstream stabilityai/stable-audio-open-small
Packages 4
Download size 1.41 GB
Minimum iOS 17.0
Peak RAM ~1200 MB

Files

File Size Compute units SHA-256
StableAudioT5Encoder.mlpackage.zip 94 MB cpuAndGPU 319a8ba775d30924…
StableAudioNumberEmbedder.mlpackage.zip 367 KB cpuAndGPU 04bdc5de00a2cf1c…
StableAudioDiT.mlpackage.zip 1.18 GB cpuOnly b17da4fc4df85782…
StableAudioVAEDecoder.mlpackage.zip 138 MB cpuAndGPU 7207544cca9799cc…
t5_vocab.json 732 KB - 7c9ff3ac1b3dbcaa…
Total 1.41 GB

compute_units is not a suggestion -- it is the configuration the conversion was verified against. Moving a package to a different compute unit can silently change the numerics (FP16 attention overflow) or crash on the GPU.

Download

hf download mlboydaisuke/coreml-zoo --include "stableaudio/*" --local-dir ./stable_audio
unzip './stable_audio/stableaudio/*.zip' -d ./stable_audio

Use in Swift

import CoreML

let config = MLModelConfiguration()
config.computeUnits = .cpuAndGPU   // as converted — see the table above

// Unzip the .mlpackage, drop it into your Xcode target and Xcode compiles it
// at build time:
let model = try StableAudioT5Encoder(configuration: config)

// ...or compile a downloaded .mlpackage at runtime:
let compiled = try await MLModel.compileModel(at: mlpackageURL)
let model = try MLModel(contentsOf: compiled, configuration: config)

This model is split into 4 Core ML packages that are driven in sequence from Swift. Load them one at a time, copy the outputs out of the MLMultiArray buffers and release each model before loading the next — two large Core ML models resident at once will OOM on an iPhone.

Demo

Conversion

License

The conversion inherits the upstream license: Stability AI Community License. See https://huggingface.co/stabilityai/stable-audio-open-small.

Free for non-commercial and limited commercial use; see the Stability AI Community License.

Credits

Downloads last month
34
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlboydaisuke/Stable-Audio-Open-Small-CoreML

Quantized
(1)
this model

Collection including mlboydaisuke/Stable-Audio-Open-Small-CoreML