| --- |
| license: other |
| license_name: stability-ai-community |
| license_link: https://huggingface.co/stabilityai/stable-audio-open-small |
| library_name: coreml |
| pipeline_tag: text-to-audio |
| base_model: stabilityai/stable-audio-open-small |
| base_model_relation: quantized |
| tags: |
| - coreml |
| - core-ml |
| - ios |
| - macos |
| - apple |
| - on-device |
| - text-to-audio |
| - music-generation |
| - dit |
| - arxiv:2505.08175 |
| --- |
| |
| # Stable Audio Open — Core ML |
|
|
| *Text-to-Music, 2024* |
|
|
| Text-to-music. Up to 11.9s stereo 44.1 kHz. Rectified flow DiT + T5 + Oobleck VAE. |
|
|
| <p><img src="https://huggingface.co/mlboydaisuke/Stable-Audio-Open-Small-CoreML/resolve/main/media/2fe3ada05b.gif" alt="Stable Audio Open demo"></p> |
|
|
| Core ML conversion of [stabilityai/stable-audio-open-small](https://huggingface.co/stabilityai/stable-audio-open-small) for on-device inference on iPhone, iPad and Mac. Converted with `coremltools`; the packages are stateless, so all sequencing and buffering lives in your Swift code. |
|
|
| | | | |
| |---|---| |
| | Task | text to audio | |
| | Upstream | [stabilityai/stable-audio-open-small](https://huggingface.co/stabilityai/stable-audio-open-small) | |
| | Packages | 4 | |
| | Download size | 1.41 GB | |
| | Minimum iOS | 17.0 | |
| | Peak RAM | ~1200 MB | |
|
|
| ## Files |
|
|
| | File | Size | Compute units | SHA-256 | |
| |---|---:|---|---| |
| | `StableAudioT5Encoder.mlpackage.zip` | 94 MB | `cpuAndGPU` | `319a8ba775d30924…` | |
| | `StableAudioNumberEmbedder.mlpackage.zip` | 367 KB | `cpuAndGPU` | `04bdc5de00a2cf1c…` | |
| | `StableAudioDiT.mlpackage.zip` | 1.18 GB | `cpuOnly` | `b17da4fc4df85782…` | |
| | `StableAudioVAEDecoder.mlpackage.zip` | 138 MB | `cpuAndGPU` | `7207544cca9799cc…` | |
| | `t5_vocab.json` | 732 KB | `-` | `7c9ff3ac1b3dbcaa…` | |
| | **Total** | **1.41 GB** | | | |
|
|
| `compute_units` is not a suggestion -- it is the configuration the conversion was verified against. Moving a package to a different compute unit can silently change the numerics (FP16 attention overflow) or crash on the GPU. |
|
|
| ## Download |
|
|
| ```bash |
| hf download mlboydaisuke/coreml-zoo --include "stableaudio/*" --local-dir ./stable_audio |
| unzip './stable_audio/stableaudio/*.zip' -d ./stable_audio |
| ``` |
|
|
| ## Use in Swift |
|
|
| ```swift |
| import CoreML |
| |
| let config = MLModelConfiguration() |
| config.computeUnits = .cpuAndGPU // as converted — see the table above |
| |
| // Unzip the .mlpackage, drop it into your Xcode target and Xcode compiles it |
| // at build time: |
| let model = try StableAudioT5Encoder(configuration: config) |
| |
| // ...or compile a downloaded .mlpackage at runtime: |
| let compiled = try await MLModel.compileModel(at: mlpackageURL) |
| let model = try MLModel(contentsOf: compiled, configuration: config) |
| ``` |
|
|
| > This model is split into 4 Core ML packages that are driven in sequence from Swift. Load them one at a time, copy the outputs out of the `MLMultiArray` buffers and release each model before loading the next — two large Core ML models resident at once will OOM on an iPhone. |
|
|
| ## Demo |
|
|
| - **Sample app** — [`sample_apps/StableAudioDemo`](https://github.com/john-rocky/CoreML-Models/tree/master/sample_apps/StableAudioDemo), a standalone SwiftUI project. |
| - **Models Zoo** — this model is downloadable and runnable inside the [Models Zoo app](https://apps.apple.com/app/id6762083207) on the App Store, no build required. |
|
|
| ## Conversion |
|
|
| - Script: [`convert_stable_audio.py`](https://github.com/john-rocky/CoreML-Models/blob/master/conversion_scripts/convert_stable_audio.py) |
| - Pitfalls hit during conversion (FP16 overflow, ANE buffer limits, stride handling): [`docs/coreml_conversion_notes.md`](https://github.com/john-rocky/CoreML-Models/blob/master/docs/coreml_conversion_notes.md) |
| - Model index: [CoreML-Models](https://github.com/john-rocky/CoreML-Models) |
|
|
| ## License |
|
|
| The conversion inherits the upstream license: **Stability AI Community License**. |
| See [https://huggingface.co/stabilityai/stable-audio-open-small](https://huggingface.co/stabilityai/stable-audio-open-small). |
|
|
| > Free for non-commercial and limited commercial use; see the Stability AI Community License. |
|
|
| ## Credits |
|
|
| - Upstream authors: [stabilityai/stable-audio-open-small](https://huggingface.co/stabilityai/stable-audio-open-small), 2024 |
| - Core ML conversion: john-rocky (Daisuke Majima) |
|
|