Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,106 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: fish-audio-research-license
|
| 4 |
+
license_link: https://huggingface.co/fishaudio/s2-pro/blob/main/LICENSE.md
|
| 5 |
+
language: [ar]
|
| 6 |
+
base_model: fishaudio/s2-pro
|
| 7 |
+
tags: [text-to-speech, egyptian-arabic, arz, fish-speech, lora]
|
| 8 |
+
pipeline_tag: text-to-speech
|
| 9 |
+
---
|
| 10 |
+
|
| 11 |
+
# s2pro-egy — Egyptian Arabic fine-tune of Fish Audio S2-Pro
|
| 12 |
+
|
| 13 |
+
Merged production model + full training/serving toolkit. **Phase 1 complete (2026-07-18/19); Phase 2 planned — see below.**
|
| 14 |
+
|
| 15 |
+
> ⚠️ **License**: derivative of `fishaudio/s2-pro` (Fish Audio Research License — **non-commercial only**).
|
| 16 |
+
> Production/commercial use requires a written license from Fish Audio (business@fish.audio). Keep this repo private.
|
| 17 |
+
|
| 18 |
+
## What's in this repo
|
| 19 |
+
|
| 20 |
+
| Path | Contents |
|
| 21 |
+
|---|---|
|
| 22 |
+
| `/` (root) | Merged model, HF `fish_qwen3_omni` layout: safetensors shards + index, `config.json`, tokenizer, `chat_template.jinja`, `codec.pth` — loads in fish-speech `main` AND sglang-omni |
|
| 23 |
+
| `checkpoints/` | Phase-1 LoRA checkpoints (fast-AR only, steps 100–1200; `step_000001200.ckpt` is the one merged) |
|
| 24 |
+
| `scripts/` | Full pipeline: data prep, relabeling, eval (synth + Soniox judge), serving (PyTorch + sglang shim), weight conversion, patches |
|
| 25 |
+
| `configs/` | Training configs (`text2semantic_finetune_egy2.yaml` = the recipe that worked) + LoRA configs |
|
| 26 |
+
| `eval/` | Mega-paragraph A/B wavs + Soniox WER reports (baseline vs checkpoints) |
|
| 27 |
+
| `logs/` | Training log + tensorboard events |
|
| 28 |
+
|
| 29 |
+
## Phase 1 — what was done
|
| 30 |
+
|
| 31 |
+
### Data (46.9 h / 15,030 clips, all 24 kHz mono, loudness-normalized)
|
| 32 |
+
|
| 33 |
+
| Dataset (HF, private) | Clips | Hours | Transcript source |
|
| 34 |
+
|---|---|---|---|
|
| 35 |
+
| `ehabnegm/noselleel-egyptian-tts` | 8,766 | 26.2 | **Soniox stt-async-v5** (6,344 clips, `transcripts_soniox/train.jsonl`) > `text_raw` (pre-CATT Whisper) |
|
| 36 |
+
| `ehabnegm/eqkawkab-egyptian-tts` | 2,062 | 8.0 | `text_raw` (pre-CATT Whisper) |
|
| 37 |
+
| `ehabnegm/moustafa-sadek-egyptian-tts` | 2,883 | 8.0 | Deepgram nova-3 |
|
| 38 |
+
| `ehabnegm/mosaifside-egyptian-tts` | 1,319 | 4.7 | Deepgram nova-3 |
|
| 39 |
+
|
| 40 |
+
Processing rules (see `scripts/prepare_data.py`, `scripts/relab3.py`):
|
| 41 |
+
- **NO tashkeel anywhere** (production input is plain text; CATT diacritics were MSA-flavored noise).
|
| 42 |
+
- Speaker grouping = **per YouTube video** (`nos_<VID>`…, 121 groups) because the trainer packs same-folder clips into one sequence.
|
| 43 |
+
- VQ tokens via s2-pro `codec.pth` (`modded_dac_vq`), protobuf shards via `build_dataset.py`.
|
| 44 |
+
- ~231 clips contain subscribe-outro phrases (1.5%) — left in for phase 1; candidate for dedupe in phase 2.
|
| 45 |
+
|
| 46 |
+
### Training recipe (the one that works)
|
| 47 |
+
|
| 48 |
+
Community-validated recipe (credit: [Enucatl/fish-speech barbero](https://github.com/Enucatl/fish-speech/blob/barbero/lora-finetune.md)):
|
| 49 |
+
|
| 50 |
+
- LoRA **fast-AR (audio decoder) ONLY** — stock `r_32_alpha_16_fast` (r=32, α=16, α/r=0.5)
|
| 51 |
+
- batch 1 × grad-accum 4, lr **1e-5**, CosineAnnealingLR → 1e-6, weight_decay 0.01
|
| 52 |
+
- `causal: false` (random sampling), max_steps **1200**, ckpt every 100, bf16-true, `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`
|
| 53 |
+
- Run: `python fish_speech/train.py --config-name text2semantic_finetune_egy2 +lora@model.model.lora_config=r_32_alpha_16_fast`
|
| 54 |
+
|
| 55 |
+
**⚠️ NEVER LoRA the slow AR with α/r ≥ 1.0**: phase-1's first attempt (r32 α32 attention+mlp on both transformers, lr 5e-5) produced pure noise by step 200 — the slow/text transformer collapsed (confirmed by slow/fast ablation, and it is literally a row in the community doc's "didn't work" table).
|
| 56 |
+
|
| 57 |
+
Required upstream patches (in `scripts/patches_notes.md`):
|
| 58 |
+
- `llama.py`: `use_reentrant=False` in both `checkpoint()` calls (LoRA + frozen embeddings otherwise breaks the grad graph)
|
| 59 |
+
- `lit_module.py`: `strict_loading = False` on `TextToSemantic` (checkpoints are LoRA-only; Lightning resume fails otherwise)
|
| 60 |
+
|
| 61 |
+
### Results (Soniox stt-async-v5 judge — Whisper hallucinates on Arabic noise, do not use it)
|
| 62 |
+
|
| 63 |
+
325-word Egyptian mega-paragraph, `masry` reference voice:
|
| 64 |
+
|
| 65 |
+
| System | WER | CER |
|
| 66 |
+
|---|---|---|
|
| 67 |
+
| s2-pro zero-shot baseline | 0.071 | 0.05 |
|
| 68 |
+
| **fine-tuned step-1200 (this model)** | **0.036–0.06** | 0.007–0.05 |
|
| 69 |
+
|
| 70 |
+
Fine-tune preserves Egyptian forms the baseline drifts on (e.g. `معايا` vs `معي`). Fast-AR tuning = voice/timbre adaptation; text-following stays RL-aligned.
|
| 71 |
+
|
| 72 |
+
## Serving (production, RTX 5090)
|
| 73 |
+
|
| 74 |
+
**Engine: [sglang-omni](https://github.com/sgl-project/sglang-omni)** (S2-Pro supported natively) + OpenAI shim.
|
| 75 |
+
Measured: **TTFA ~0.5 s** (streaming, warm radix cache), RTF ~0.85, quality WER 0.00 on server output, output resampled to **24 kHz**.
|
| 76 |
+
|
| 77 |
+
Blackwell/sm_120 porting patches (all required, `scripts/patches_notes.md`):
|
| 78 |
+
1. delete/rename `site-packages/deep_gemm` (asserts on missing CUDA_HOME at import)
|
| 79 |
+
2. `apt install cuda-nvcc-13-0 libcublas-dev-13-0 libcusparse-dev-13-0 libcusolver-dev-13-0 libcurand-dev-13-0` (torch cu130) + `ninja`
|
| 80 |
+
3. `engine_builder.py`: attention backend `fa3` → `triton` (FA3 = Hopper-only)
|
| 81 |
+
4. `engine_builder.py`: `disable_cuda_graph: True` (graph capture calls FA3 kernels), `mem_fraction_static: 0.55`
|
| 82 |
+
5. `audio_decoder.py`: `FISH_FORCE_SDPA=1` env-gated pure-SDPA replacement for `sgl_kernel.flash_attn_with_kvcache` (`scripts/patch_audio_decoder.py`) — Fast-AR attends ≤11 positions, SDPA is exact and fast
|
| 83 |
+
|
| 84 |
+
Launch:
|
| 85 |
+
```bash
|
| 86 |
+
export CUDA_HOME=/usr/local/cuda-13.0 SGLANG_ENABLE_JIT_DEEPGEMM=0 FISH_FORCE_SDPA=1
|
| 87 |
+
sgl-omni serve --model-path <this-repo-dir> --config examples/configs/s2pro_tts.yaml --port 8001
|
| 88 |
+
python scripts/serve_shim.py --port 8000 # OpenAI contract: model=s2pro-egy, voices, 24kHz
|
| 89 |
+
```
|
| 90 |
+
|
| 91 |
+
API (OpenAI-compatible): `POST /v1/audio/speech` `{"model":"s2pro-egy","input":"...","voice":"masry","stream":true,"response_format":"pcm"}`.
|
| 92 |
+
Voices = `voices/<name>.wav+.txt` reference pairs (`masry` = eqkawkab narrator; `noselleel`/`noselleel2`/`noselleel3` = noselleel narrator candidates).
|
| 93 |
+
|
| 94 |
+
## Phase 2 — plan (next session)
|
| 95 |
+
|
| 96 |
+
**Goal:** deeper Egyptian pronunciation/prosody (slow-AR territory) without breaking RL alignment.
|
| 97 |
+
|
| 98 |
+
1. **Config:** `r32 α16 ALL modules` (α/r = 0.5, attention+mlp on slow+fast) — the community doc's "minor slow-AR degradation, usable — worth investigating" row. lr 1e-5 cosine, wd 0.01, causal false, max 1200 steps, ckpt every 100.
|
| 99 |
+
2. **Start from:** this merged model as the base (`pretrained_ckpt_path` → merged dir), so phase-1 voice gains are kept.
|
| 100 |
+
3. **Hard gates:** eval at step 200/400 with Soniox WER + human listening vs phase-1 (comparison page: `scripts/serve_eval.py`). The failure mode is audio ending early / going quiet / noise — **stop immediately if WER degrades**, keep phase-1.
|
| 101 |
+
4. **Data option:** drop/downweight the 231 subscribe-outro clips; optionally add more clean hours.
|
| 102 |
+
5. **Success metric:** phase-2 ≥ phase-1 WER AND user prefers pronunciation blind.
|
| 103 |
+
|
| 104 |
+
## Provenance
|
| 105 |
+
- Base: `fishaudio/s2-pro` · Training/eval infra: fish-speech `main` (e5e2926) · Serving: sglang-omni
|
| 106 |
+
- Fine-tuned 2026-07-18/19 on 1× RTX 5090 (32 GB), total ~1.5 h training for phase 1.
|