Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| DurationPredictor.mlpackage | 3 items | ||
| TextEncoder.mlpackage | 3 items | ||
| VectorEstimator.mlpackage | 3 items | ||
| Vocoder.mlpackage | 3 items | ||
| voice_styles | 10 items | ||
| .gitattributes | 1.52 kB xet | 818ba6de | |
| DurationPredictor.mlpackage.zip | 3.22 MB xet | 15bc3109 | |
| README.md | 3.32 kB xet | 3cdb2617 | |
| TextEncoder.mlpackage.zip | 33.1 MB xet | 4cb3ad9f | |
| VectorEstimator.mlpackage.zip | 239 MB xet | 8b36f775 | |
| Vocoder.mlpackage.zip | 94.5 MB xet | 8e267676 | |
| config.json | 34 Bytes xet | 504f4604 | |
| tts.json | 8.25 kB xet | f9f956cf | |
| unicode_indexer.json | 278 kB xet | 335f0ab4 |
SupertonicTTS-3 — CoreML (.mlpackage, iOS)
First-party CoreML export of Supertonic-3's four non-autoregressive flow-matching graphs, for
on-device iOS / Apple Neural Engine. Built by our own pipeline (speech-models/stmodels): weights
lifted from the Supertone/supertonic-3 ONNX
initializers → PyTorch nn.Module → coremltools (mlprogram, FP32, iOS18+).
Graphs & parity (FP32, vs ONNX Runtime)
| Module | mlpackage | parity max|Δ| |
|---|---|---|
| Duration predictor | DurationPredictor.mlpackage |
7.2e-06 ✓ |
| Vector estimator (ODE denoiser) | VectorEstimator.mlpackage |
2.5e-03 ✓ |
| Vocoder | Vocoder.mlpackage |
3.0e-04 ✓ |
| Text encoder | TextEncoder.mlpackage |
mean 2.5e-04 (max 2.5e-2 at isolated positions) |
Text/duration use fixed T=128 (relpos attention has T-dependent pad widths — pad/segment text to
128); vocoder + vector-estimator use a dynamic latent-length RangeDim. The host runs the flow-matching
ODE loop (vector_estimator ×total_steps) — the graphs contain no control flow. Assets to drive them:
tts.json, unicode_indexer.json (G2P-free tokenizer table), voice_styles/*.json.
FP32 = parity reference. For ANE residency, use the mixed-precision
Supertonic-3-CoreML-FP16— vocoder + duration FP16, text-encoder + vector-estimator FP32; measured transparent at 47–51 dB mag-STFT SNR.
Attribution & license
- Weights: derivative of
Supertone/supertonic-3(commit3cadd1ee6394adea1bd021217a0e650ede09a323), Supertone Inc., arXiv:2503.23108 — OpenRAIL-M (use-based restrictions carry over: no non-consensual impersonation/deepfakes, etc.).
Other Supertonic-3 formats
- Supertonic-3 — CoreML (FP16) — mixed-precision ANE variant (47–51 dB).
- Supertonic-3 — ONNX (INT8) — server / desktop (ONNX Runtime).
- Supertonic-3 — LiteRT — Android / Qualcomm NPU (.tflite).
Ecosystem
- soniqo.audio — website / use-case explorer (transcription, voice cloning, live ASR, voice agents).
- speech-core — C++ orchestration library; Supertonic plugs in as a
TTSInterfaceCoreML model. - speech-swift — Apple Silicon MLX + CoreML runtime.
- speech-android — Android SDK consuming on-device LiteRT bundles.
Other CoreML models
- Total size
- 770 MB
- Files
- 31
- Last updated
- Jul 21
- Pre-warmed CDN
- US EU US EU