Text-to-Speech
ONNX
Safetensors
Qwen3-TTS
Text-to-Speech
TTS
speech-edit
aivoice / README.md
Athspi's picture
Update README.md
2d20f26 verified
|
Raw
History Blame Contribute Delete
2.72 kB
---
license: apache-2.0
datasets:
- amphion/Emilia-Dataset
- amphion/Emilia-NV
- simon3000/genshin-voice
language:
- zh
- en
- ja
- ko
base_model:
- Qwen/Qwen3-0.6B
pipeline_tag: text-to-speech
tags:
- Text-to-Speech
- TTS
- speech-edit
---
notebook https://cuty.io/pQGWF3c7
# ViiTorVoice-NAR Local Models
[![GitHub](https://img.shields.io/badge/GitHub-viitor--voice--nar-181717?logo=github)](https://github.com/viitor-ai/viitor-voice-nar)
[![Hugging Face Demo](https://img.shields.io/badge/Hugging%20Face-Demo-FFD21E?logo=huggingface&logoColor=000)](https://huggingface.co/spaces/ZzWater/ViiTorVoice)
This directory contains the local model files used by
[viitor-ai/viitor-voice-nar](https://github.com/viitor-ai/viitor-voice-nar).
ViiTorVoice-NAR is a non-autoregressive speech generation model for voice
cloning, local speech editing, and emotion / paralinguistic speech control.
The files in this directory are split by function so each model component can
be loaded independently.
## Directory
```text
local_models/
β”œβ”€β”€ aligner/
β”‚ └── Qwen3-ForcedAligner-0.6B/
β”œβ”€β”€ assets/
β”‚ └── dualcodec_silence_2s.pt
β”œβ”€β”€ dualcodec/
β”‚ β”œβ”€β”€ dualcodec_ckpts/
β”‚ └── w2v-bert-2.0/
└── llm/
└── 0p6_emotion/
```
## Model Components
| Component | Path | Purpose |
| --- | --- | --- |
| ViiTorVoice-NAR LLM | `llm/0p6_emotion/` | Generates target speech tokens from text, prompt speech tokens, edit masks, duration conditions, and emotion or non-verbal tags. |
| DualCodec | `dualcodec/dualcodec_ckpts/` | Converts waveform audio into discrete speech codebook tokens and decodes generated tokens back into waveform audio. |
| W2V-BERT 2.0 | `dualcodec/w2v-bert-2.0/` | Extracts semantic speech features used by the DualCodec encoder. |
| Qwen3 Forced Aligner | `aligner/Qwen3-ForcedAligner-0.6B/` | Aligns speech audio with text and provides timestamps for local speech editing. |
| Runtime Assets | `assets/` | Stores small auxiliary files, such as precomputed silence tokens used during generation or padding. |
## Main Uses
- Voice cloning: synthesize new speech from target text while preserving the
speaker characteristics of prompt audio.
- Local speech editing: replace only the changed region of an utterance while
keeping the rest of the audio stable.
- Emotion and paralinguistic control: condition generation with tags such as
emotion labels or non-verbal vocal events.
## Notes
- Keep the directory structure unchanged unless the loading code is updated as
well.
- Model weights are large binary files and are usually stored outside normal
git tracking.
- Check the upstream project and each submodel for license and usage terms.