| --- |
| license: apache-2.0 |
| base_model: Qwen/Qwen3-TTS-12Hz-0.6B-Base |
| pipeline_tag: text-to-speech |
| language: |
| - en |
| tags: |
| - audio |
| - tts |
| - qwen3-tts |
| - custom-voice |
| - fine-tuned |
| --- |
| |
| # E3 Voice4 / Jaiden — Qwen3-TTS 0.6B direct fine-tune |
|
|
| This is the ear-selected **epoch-0 full fine-tune** of |
| [`Qwen/Qwen3-TTS-12Hz-0.6B-Base`](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base) |
| for the E3 Voice4 single-speaker voice. The embedded speaker is exposed as |
| `speaker_test`; inference does not require reference audio. |
|
|
| The checkpoint was selected from five fixed-seed draws. The preferred output |
| was draw 1 (seed 43), chosen for its balance of speaker identity and expressive |
| range. This replaces the older Vivian/Fish-distilled weights previously on the |
| repository's `main` branch. Historical models remain available through the |
| repository's `v1`–`v5` tags and commit history. |
|
|
| ## Quickstart |
|
|
| ```bash |
| pip install -U qwen-tts soundfile |
| ``` |
|
|
| ```python |
| import soundfile as sf |
| import torch |
| from qwen_tts import Qwen3TTSModel |
| |
| model = Qwen3TTSModel.from_pretrained( |
| "e3group-research/enderv1.0", |
| device_map="cuda:0", |
| dtype=torch.bfloat16, |
| ) |
| |
| torch.manual_seed(43) |
| wavs, sample_rate = model.generate_custom_voice( |
| text="This is a test of the E3 Voice4 checkpoint.", |
| speaker="speaker_test", |
| language="english", |
| ) |
| sf.write("voice4.wav", wavs[0], sample_rate) |
| ``` |
|
|
| The only configured speaker key is `speaker_test` (speaker slot 3000). |
|
|
| ## Training recipe |
|
|
| - Base model: Qwen3-TTS 12 Hz 0.6B Base |
| - Method: single-stage full fine-tune, not LoRA or teacher distillation |
| - Data: 130 single-speaker speech clips, 16.56 minutes total |
| - Retained checkpoint: epoch 0 (65 optimizer steps) |
| - Learning rate: `2e-5`, with 50 warmup steps |
| - Objective: codec-0 cross entropy + `0.3 ×` sub-talker loss |
| - Embedded speaker: slot 3000, populated from the training reference voice |
|
|
| `history.json` records the complete eight-epoch training run. Only epoch 0 was |
| retained for this release; later epochs were rejected because they overfit and |
| could produce severe duration runaway. |
|
|
| ## Selection measurements |
|
|
| For the selected seed-43 evaluation draw: |
|
|
| | Measurement | Value | |
| |---|---:| |
| | Speaker similarity vs. internal reference | 0.9261 | |
| | F0 p10–p90 range | 12.45 semitones | |
| | F0 median | 246.83 Hz | |
| | Duration | 22.78 seconds | |
| | Speaking rate | 187 WPM | |
|
|
| These are measurements for the selected generation, not aggregate benchmark |
| results. Generation is stochastic and results vary with text and seed. |
|
|
| ## Integrity |
|
|
| - `model.safetensors` SHA-256: `54f399102bdfc0b55e5c2ffa7fd1d3eb3d68cae4a516e151cbe6df8d01e96b7a` |
| - `speech_tokenizer/model.safetensors` SHA-256: `836b7b357f5ea43e889936a3709af68dfe3751881acefe4ecf0dbd30ba571258` |
|
|
| ## Limitations and responsible use |
|
|
| The checkpoint was selected primarily on English narration and has not been |
| validated equally across all languages or production workloads. It may inherit |
| pronunciation, hallucination, pacing, and robustness limitations from the base |
| model. |
|
|
| This is a research voice checkpoint. It is not affiliated with or endorsed by |
| the source speaker. Do not use it for deceptive impersonation, fraud, |
| harassment, or to imply endorsement. Users are responsible for obtaining any |
| required consent and complying with applicable privacy, publicity, copyright, |
| and disclosure requirements. |
|
|
| ## Base model |
|
|
| - [Qwen3-TTS technical report](https://huggingface.co/papers/2601.15621) |
| - [Qwen3-TTS source repository](https://github.com/QwenLM/Qwen3-TTS) |
|
|
|
|