Sonora / README.md
lmcfarlin
vat3-24k: publish milestone-3 full-VAT checkpoint (mixed verdict) + registry row
e00ca4c
|
Raw
History Blame Contribute Delete
4.26 kB
---
language:
- en
license: apache-2.0
tags:
- text-to-speech
- tts
- matcha-tts
- flow-matching
- tflite
- onnx
- audio
- on-device
- expressive-speech
datasets:
- ljspeech
- libritts_r
metrics:
- loss
- wer
pipeline_tag: text-to-speech
---
# Sonora
**Sonora** is a lightweight, non-autoregressive text-to-speech (TTS) model line built for
real-time, on-device, and resource-constrained environments — the voice engine of
**Project Prosodia**. Architecture: Matcha-TTS-family Optimal-Transport Conditional Flow
Matching (OT-CFM) acoustic model + HiFi-GAN vocoder, with a growing set of **expressive
conditioning channels** (valence / arousal-energy / tension via FiLM) aimed at directable,
story-driven narration.
This repository is the **model registry**: each top-level directory is one promoted,
audited milestone artifact with its own README, eval report, and provenance.
## Registry index
> **Current trunk:** `derisk-energy-24k` (acoustic) + `vocoder-24k-hifigan` (vocoder) —
> the 24 kHz multi-speaker line. Next up: the v1.1 VAT continuation (story-driven valence
> corpus), warm-starting from `vat3-24k`.
> Machine-readable lineage: [`registry.json`](./registry.json).
| Directory | Training pass | What it is | Status |
|---|---|---|---|
| [`baseline-ljspeech-22k/`](./baseline-ljspeech-22k) | **Phase 0** | Single-speaker LJSpeech fine-tune, 22.05 kHz, end-to-end graphs (ONNX + TFLite) and the LiteRT split-graph mobile lane. *(Renamed from `v1-ljspeech` on 2026-07-22 — update old deep links.)* | Published 2026-07-12 |
| [`vocoder-24k-hifigan/`](./vocoder-24k-hifigan) | 24 kHz trunk — vocoder | HiFi-GAN vocoder fine-tune, 24 kHz / 80-band — the vocoder pairing for the multi-speaker line. Converged + human-audited (copy-synthesis A/B indistinguishable). | Published 2026-07-16 |
| [`derisk-energy-24k/`](./derisk-energy-24k) | **Expressive de-risk** (north-star §7) | First trained expressive channel: energy/arousal FiLM conditioning, multi-speaker (247 spk, LibriTTS-R), 24 kHz. Controllability ρ≈1.0, identity-preserving, WER-safe; includes the VAT-ready LiteRT split-graph export. | Published 2026-07-16 |
| [`vat3-24k/`](./vat3-24k) | **Full 3-channel VAT** (milestone 3) | First all-three-channels training (valence/energy/tension FiLM), warm-started from `derisk-energy-24k`. **Mixed verdict**: energy PASS, tension near-pass, valence FAIL — a corpus-label limitation, not a training one (see its README). Warm-start seed for the v1.1 continuation. | Published 2026-07-22 |
Directory names are deliberately **descriptive, not phase-numbered** — artifacts keep their
identity even as the roadmap's phase labels evolve; the "Training pass" column carries the
roadmap mapping.
**Developed by:** Artificial Humanity · **License:** Apache-2.0 · **Languages:** English
## Model line at a glance
* **Phase 0 (`baseline-ljspeech-22k`)** proved the on-device lane: ~18.2M-parameter acoustic
model, single forward pass text→PCM, running on phones via TFLite/LiteRT.
* **The 24 kHz multi-speaker line** (`vocoder-24k-hifigan` + `derisk-energy-24k`) is the
current trunk: native 24 kHz, 247 speakers, and validated FiLM conditioning — the
de-risk experiment demonstrated a continuous, monotonic, speaker-preserving loudness/
energy control learned from weak labels in ~35 GPU-hours on consumer hardware.
* **Full 3-channel VAT conditioning** (valence/arousal/tension) and story-driven expressive
training data (directed teacher synthesis + aligned public-domain audiobooks) are in
active development; artifacts land here as they pass their gates.
## Provenance & audit culture
Every artifact directory documents: training corpus + license wall, warm-start lineage,
pre-registered convergence criteria, automated eval-harness results (controllability /
identity leakage / intelligibility), and human audition verdicts. Checkpoints are
distributed as **safetensors** (weights-only); full resumable training checkpoints are
retained privately.
## Companion dataset
The **`sonora-expressive-registers`** dataset (CC-BY-4.0) — certified keeps from our
directed teacher-synthesis pipeline, with exact intended-VAT labels and full render
provenance — is published separately under this organization.