BlueV3 / README.md
notmax123's picture
Add model card README for PyTorch weights
8a03141 verified
|
Raw
History Blame Contribute Delete
2.45 kB
---
license: mit
library_name: pytorch
tags:
- text-to-speech
- tts
- hebrew
- multilingual
- flow-matching
- pytorch
pipeline_tag: text-to-speech
---
# BlueV3 — PyTorch Weights
Multilingual neural TTS (Hebrew-first) based on flow matching in a compressed latent space.
**TTS version:** `v1.7.3` · **Sample rate:** 44.1 kHz · **Checkpoint:** `ckpt_step_767000`
## Contents
| Path | Description |
|------|-------------|
| `checkpoints/text2latent/ckpt_step_767000.pt` | Text encoder + vector-field estimator + reference encoder |
| `checkpoints/duration_predictor/duration_predictor_final.pt` | Utterance duration predictor |
| `stats_multilingual.pt` | Latent mean / std for normalization |
| `configs/tts.json` | Model / AE / DP config |
> **Not included:** the autoencoder / vocoder codec (`ae_*.pt`). Use your local AE checkpoint (e.g. `checkpoints/44k_decoder/ae_541000.pt`) or the ONNX vocoder from [`notmax123/BlueV3-onnx`](https://huggingface.co/notmax123/BlueV3-onnx).
## Companion repos
- ONNX runtime bundle (incl. vocoder): [`notmax123/BlueV3-onnx`](https://huggingface.co/notmax123/BlueV3-onnx)
## Quick start (PyTorch)
```bash
# from the BlueV3 code repo
hf download notmax123/BlueV3 --local-dir ./hf_bluev3
# place weights where run_pt_inference / synth scripts expect them, e.g.:
mkdir -p checkpoints/text2latent checkpoints/duration_predictor
cp hf_bluev3/checkpoints/text2latent/ckpt_step_767000.pt checkpoints/text2latent/
cp hf_bluev3/checkpoints/duration_predictor/duration_predictor_final.pt checkpoints/duration_predictor/
cp hf_bluev3/stats_multilingual.pt .
cp hf_bluev3/configs/tts.json configs/tts.json
```
Example synthesis (IPA / phonemes, Hebrew):
```bash
uv run python run_pt_inference.py \
--text "metsujˈan. vetatˈus levˈad?" \
--lang he \
--style_json voice_styles/Rotem.json \
--steps 8 \
--cfg 3.0 \
--out out.wav
```
You still need:
1. Code from this project’s Git repo
2. An AE codec checkpoint for waveform decode (or use ONNX vocoder)
3. A voice style JSON (`style_ttl` + `style_dp`), e.g. export via `export_ref_latent.py` / reference WAV encoding
## Model notes
- Latent: 24-dim AE → compressed **144** channels (`chunk_compress_factor=6`)
- Flow matching with classifier-free guidance (`cfg_scale` typically `3.0`)
- Universal IPA vocab (size 256)
- Duration is predicted in **seconds**, then mapped to latent frames
## License
MIT (see frontmatter).