--- license: apache-2.0 datasets: - amphion/Emilia-Dataset - amphion/Emilia-NV - simon3000/genshin-voice language: - zh - en - ja - ko base_model: - Qwen/Qwen3-0.6B pipeline_tag: text-to-speech tags: - Text-to-Speech - TTS - speech-edit --- notebook https://cuty.io/pQGWF3c7 # ViiTorVoice-NAR Local Models [![GitHub](https://img.shields.io/badge/GitHub-viitor--voice--nar-181717?logo=github)](https://github.com/viitor-ai/viitor-voice-nar) [![Hugging Face Demo](https://img.shields.io/badge/Hugging%20Face-Demo-FFD21E?logo=huggingface&logoColor=000)](https://huggingface.co/spaces/ZzWater/ViiTorVoice) This directory contains the local model files used by [viitor-ai/viitor-voice-nar](https://github.com/viitor-ai/viitor-voice-nar). ViiTorVoice-NAR is a non-autoregressive speech generation model for voice cloning, local speech editing, and emotion / paralinguistic speech control. The files in this directory are split by function so each model component can be loaded independently. ## Directory ```text local_models/ ├── aligner/ │ └── Qwen3-ForcedAligner-0.6B/ ├── assets/ │ └── dualcodec_silence_2s.pt ├── dualcodec/ │ ├── dualcodec_ckpts/ │ └── w2v-bert-2.0/ └── llm/ └── 0p6_emotion/ ``` ## Model Components | Component | Path | Purpose | | --- | --- | --- | | ViiTorVoice-NAR LLM | `llm/0p6_emotion/` | Generates target speech tokens from text, prompt speech tokens, edit masks, duration conditions, and emotion or non-verbal tags. | | DualCodec | `dualcodec/dualcodec_ckpts/` | Converts waveform audio into discrete speech codebook tokens and decodes generated tokens back into waveform audio. | | W2V-BERT 2.0 | `dualcodec/w2v-bert-2.0/` | Extracts semantic speech features used by the DualCodec encoder. | | Qwen3 Forced Aligner | `aligner/Qwen3-ForcedAligner-0.6B/` | Aligns speech audio with text and provides timestamps for local speech editing. | | Runtime Assets | `assets/` | Stores small auxiliary files, such as precomputed silence tokens used during generation or padding. | ## Main Uses - Voice cloning: synthesize new speech from target text while preserving the speaker characteristics of prompt audio. - Local speech editing: replace only the changed region of an utterance while keeping the rest of the audio stable. - Emotion and paralinguistic control: condition generation with tags such as emotion labels or non-verbal vocal events. ## Notes - Keep the directory structure unchanged unless the loading code is updated as well. - Model weights are large binary files and are usually stored outside normal git tracking. - Check the upstream project and each submodel for license and usage terms.