Buckets:
| language: | |
| - en | |
| license: apache-2.0 | |
| tags: | |
| - text-to-speech | |
| - tts | |
| - matcha-tts | |
| - flow-matching | |
| - tflite | |
| - onnx | |
| - audio | |
| - on-device | |
| datasets: | |
| - ljspeech | |
| metrics: | |
| - loss | |
| pipeline_tag: text-to-speech | |
| # Sonora (Phase 0) | |
| **Sonora** is a lightweight, non-autoregressive text-to-speech (TTS) model designed for real-time, on-device, and resource-constrained environments. | |
| This model card covers the **Phase 0 baseline** fine-tune, which is built on the **Matcha-TTS** architecture and fine-tuned on the single-speaker LJSpeech dataset. The final model graph embeds the **HiFi-GAN vocoder** (`hifigan_T2_v1`) to execute full end-to-end waveform synthesis directly in a single forward pass (Text IDs $\rightarrow$ PCM Audio). | |
| * **Voice Engine Integration:** Deployed inside **Project Prosodia** (`ProsodiaActor`) | |
| * **Underlying Architecture:** Optimal Transport Conditional Flow Matching (OT-CFM) | |
| * **Parameter Count:** ~18.2 Million | |
| * **Sample Rate:** 22,050 Hz (mono) | |
| --- | |
| ## Model Details | |
| * **Developed by:** Artificial Humanity | |
| * **Model Type:** Acoustic Flow-Matching Model + HiFi-GAN Vocoder (end-to-end graph) | |
| * **Language(s):** English (en) | |
| * **License:** Apache 2.0 | |
| * **Base Checkpoint:** `matcha_ljspeech.ckpt` (v1.0) | |
| * **Parent Model Paper:** [Matcha-TTS: A Fast CFM Text-to-Speech Model (Mehta et al., 2024)](https://arxiv.org/abs/2309.03199) | |
| --- | |
| ## Active Formats | |
| The model is distributed in multiple formats for execution on different hardware backends: | |
| 1. **`checkpoint_epoch=199_e2e_float16.tflite`** (89.8 MB) — **Recommended for on-device deployment.** Quantized to Float16 for accelerated execution on mobile/edge NPUs and GPUs via LiteRT (TensorFlow Lite). | |
| 2. **`checkpoint_epoch=199_e2e_float32.tflite`** (178.2 MB) — High-precision Float32 model for desktop CPU/GPU evaluation. | |
| 3. **`checkpoint_epoch=199_e2e.onnx`** (178.2 MB) — Standalone ONNX graph for server-side inference or platform debugging. | |
| --- | |
| ## Training Environment & Hardware | |
| The model was trained locally on the `ai-lab-0` workstation using a containerized AMD ROCm configuration: | |
| * **Host CPU:** AMD Ryzen AI Max+ 395 (16 Cores, 32 Threads, Zen 5) | |
| * **Host GPU:** AMD Radeon 8060S Graphics (RDNA 3.5, 40 Compute Units, Target: `gfx1151`) | |
| * **Allocatable VRAM:** 112 GB Unified LPDDR5X (shared memory pool) | |
| * **Operating System:** Ubuntu 26.04 LTS (Kernel `7.0.0-27-generic`) | |
| * **Container Environment:** `rocm/pytorch:latest` running PyTorch 2.x + ROCm 7.2.4 | |
| * **Optimization:** 16-bit Automatic Mixed Precision (AMP), batch size 32, shared memory (`--ipc=host`, `--shm-size 16G`). | |
| --- | |
| ## Training History & Hyperparameters | |
| * **Total Epochs Trained:** 260 (over 101,000 steps) | |
| * **Selected Checkpoint:** Epoch 199 (step ~82,000) where validation loss converged and stabilized. | |
| * **Loss Function:** Joint Flow-Matching loss, duration prediction loss, and monotonic alignment prior loss. | |
| * **Troubleshooting Fixes Applied:** | |
| * **Audio Decoding:** Replaced `torchaudio.load()` with `soundfile.read()` in the data pipeline to bypass CUDA-exclusive shared library dependencies (`torchcodec` decoding failures) inside the ROCm container. | |
| * **Matplotlib Compatibility:** Patched validation figure logging to use `np.asarray(fig.canvas.buffer_rgba())[:, :, :3]`, avoiding the deprecated and removed `tostring_rgb()` method under Matplotlib 3.9+. | |
| --- | |
| ## Intended Uses & Limitations | |
| ### Intended Use | |
| * **On-Device Narration:** Real-time text-to-speech synthesis directly on mobile devices (iOS/Android) and desktop clients without sending telemetry data to external cloud APIs. | |
| * **Emotion & Directability Scaffolding:** Serves as the foundation for the Phase 1 directable voice actor, which will add Valence/Arousal/Tension (VAT) conditioning controls. | |
| ### Limitations | |
| * **LJSpeech Speaker Profile:** The model is currently trained on a single female narrator and lacks voice variety or pitch modulation. Multi-speaker voice blending and casting grids are planned for the next training iteration (using Expresso and LibriTTS-R). | |
| * **Phoneme Representation:** Requires an external G2P (Grapheme-to-Phoneme) converter to translate raw text into phoneme ID sequences before model execution. Project Prosodia handles this natively via its built-in CMUdict lexicon Rust crate (`crates/actor/src/g2p.rs`). | |
| --- | |
| ## How to Get Started | |
| ### Loading the Model in Project Prosodia (Rust) | |
| The compiled model weights are loaded dynamically by the `StageCoordinator` through the TFLite runtime C-API. | |
| To use this model locally: | |
| 1. Download `checkpoint_epoch=199_e2e_float16.tflite`. | |
| 2. Rename the file to `styletts2_lite.tflite` (or update your environment configs to load the custom filename). | |
| 3. Place it in your shared `Models/` directory. | |
| 4. Run your application: | |
| ```bash | |
| cargo run --package tuner-app | |
| ``` | |
| ### Input/Output Tensor Node Contract | |
| The model graph matches the following tensor specifications: | |
| * **Inputs:** | |
| * `x`: token IDs representing phonemes (Shape: `[1, token_count]`, `i32` / `int64`). | |
| * `x_lengths`: length of the token sequence (Shape: `[1]`, `i32` / `int64`). | |
| * `scales`: scale parameters (Shape: `[2]`, `f32` containing `[temperature, length_scale]`). | |
| * **Outputs:** | |
| * `wav`: Raw PCM audio waveform samples (Shape: `[1, audio_length]`, `f32` mono 22.05 kHz). | |
| * `wav_lengths`: The total count of generated audio samples (Shape: `[1]`, `i32` / `int64`). | |
Xet Storage Details
- Size:
- 5.4 kB
- Xet hash:
- 328b30dbecfe3c3baa12d10fa94fcf02f06cff349f1fffc720beffba1d2110f7
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.