Instructions to use Reza2kn/Gooya-RizehPizeh-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Piper
How to use Reza2kn/Gooya-RizehPizeh-v2 with Piper:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
File size: 4,190 Bytes
17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b 4470132 a1e56c2 4470132 a1e56c2 4470132 17efd9b 4470132 17efd9b 4470132 17efd9b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 | ---
language:
- fa
license: mit
pipeline_tag: text-to-speech
base_model: Reza2kn/Gooya-RizehPizeh-v1.5
library_name: piper
model_type: text-to-speech
tags:
- audio
- text-to-speech
- piper
- vits
- persian
- farsi
- negara
pretty_name: Gooya RizehPizeh v2
---
# Gooya RizehPizeh v2
Persian single-speaker Piper/VITS2 voice trained on full, unchunked Wikipedia recordings whose audio/transcript agreement passed AvaSanj. It uses a case-sensitive Negara-style phonetic/Finglish frontend and produces 22,050 Hz mono audio.
## Current release
This revision replaces the earlier segmented-data v2 model with the user-selected unchunked/AvaSanj-approved 50K run.
- Full checkpoint: `checkpoint/gooya-rizehpizeh-v2-unchunked-avasanj-step50000-val_mel0.4138.ckpt`
- Checkpoint SHA-256: `8b52b3b6060b4698aba0c8bbd2f9b338aa6ba811d28a620e5f492c0fa1ae4d2b`
- Piper/browser model: `gooya-rizehpizeh-v2.onnx`
- ONNX SHA-256: `3f85f450fc21415e1eddaa0aa894c667e76d14aa9a342c4e77f1f9009b96f170`
- ONNX config SHA-256: `981770210ac1aea618cb8134597127911051deabf9900824df3388f209a884fa`
- Generator parameters: approximately 23.7M
- Sample rate: 22,050 Hz
- Speakers: 1
- Phoneme type: `text`
- Symbol inventory: 256 IDs
## Training data
The parent corpus comes from the approved Persian rows in [`Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731`](https://huggingface.co/datasets/Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731).
- Train: 1,290 full recordings, 21.1455 hours
- Parent-disjoint validation: 100 full recordings, 1.6445 hours
- Untouched sealed-eval agreement slice: 253 recordings, 4.1080 hours
- Recording duration: 38.68–95.96 seconds; mean 58.94 seconds
- Phoneme IDs per recording: 1,711–3,423; mean 2,174.65
- AvaSanj compact CER threshold: `<= 0.12`
- Train/validation/sealed-eval parent overlap: zero
- Audio transform: none
- Silence trimming: disabled
- Segmentation/chunking: none
The 253-row sealed-eval agreement slice was not consumed by training or checkpoint selection. The complete selection and hash receipt is in `training/unchunked-avasanj-agree-report.json`.
## Lineage and training
The original v1.5 production ONNX had no matching resumable checkpoint. Its 350 inference-time generator tensors were reconstructed into a compatible trainable Piper checkpoint; training-only branches absent from the ONNX came from the compatible Mana template. Recovery receipts remain under `training/`.
The promoted run resumed from that recovered v1.5 state and used:
- BF16 mixed precision
- Batch size 1
- Generator learning rate `5e-6`
- Discriminator learning rate `2.5e-6`
- Configured limit: 50,000 optimizer steps
- Final/best validation mel: `0.4138`
The exact continuation script is `training/train_unchunked_avasanj_agree_full50k.sh`.
## Validation
The published ONNX is self-contained and has no external tensor file. Direct ONNX Runtime CPU inference passed on both the canonical and long colloquial probes with finite, positive-length output. The same WAVs used for human listening are included under `samples/`.
- Inputs: `input`, `input_lengths`, `scales`
- Output: `output`
- File size: 63,516,051 bytes
- Canonical sample: 4.7137 seconds
- Long colloquial sample: 12.6084 seconds
Machine-readable receipts are under `browser/` and `training/`.
## Inference
Input must already be case-sensitive Negara-compatible phonetic text:
```bash
echo "zurxAneye vahdat yek makAne varzeSi dar Sahre borujerd ast." | \
piper --model gooya-rizehpizeh-v2.onnx \
--config gooya-rizehpizeh-v2.onnx.json \
--output_file output.wav
```
For long-form narration, split at semantic sentence or clause boundaries into substantial spans and join them with boundary-aware pauses. Tiny fixed chunks tend to sound breathy and robotic.
## Limitations
- G2P/Finglish quality remains the dominant source of pronunciation errors.
- The approved training labels are flat Negara phoneme sequences; existing internal punctuation was not available in those labels.
- Very long single-pass phoneme sequences can still collapse or omit content; use substantial semantic spans.
- Proper nouns may require corrected phonetic input.
## License
MIT. See `LICENSE`.
|