Text-to-Speech
Transformers
ONNX
teratts_onnx
feature-extraction
onnxruntime
russian
english
custom-code
custom_code
Instructions to use TeraSpace/TeraTTSv2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TeraSpace/TeraTTSv2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="TeraSpace/TeraTTSv2", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("TeraSpace/TeraTTSv2", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
| library_name: transformers | |
| pipeline_tag: text-to-speech | |
| tags: | |
| - onnxruntime | |
| - text-to-speech | |
| - russian | |
| - english | |
| - custom-code | |
| # TeraTTSv2 ONNX | |
| TeraTTSv2 is a self-contained ONNX Runtime text-to-speech release with | |
| selectable diffusion samplers, ten voice styles, Russian stress marking, and | |
| streamed audio output. | |
| This release uses the clean English/Russian 25-second teacher and its matching | |
| eight-step CFG-3 distilled student. | |
| > **Important — Russian stress is automatic.** Text inside `<ru>…</ru>` receives | |
| > stress markers automatically by default. Explicit `+` markers always win. | |
| > | |
| > **Important — cross-language prompts.** When an English reference voice is | |
| > speaking Russian, experiment with `duration_scale` below `1` (for example | |
| > `0.8`). It is usually a better starting point than the default `1`. | |
| > | |
| > **Recommended voices:** `ru_f1` and `ru_m5` are the preferred Russian voice | |
| > prompts. | |
| ## Installation | |
| ```bash | |
| pip install -r requirements.txt | |
| ``` | |
| `sounddevice` is only required for direct speaker playback. On Linux, install | |
| the system PortAudio library if it is not already present. | |
| ## Load with Transformers | |
| ```python | |
| from transformers import AutoModel | |
| tts = AutoModel.from_pretrained( | |
| "TeraSpace/TeraTTSv2", | |
| trust_remote_code=True, | |
| provider="CPUExecutionProvider", | |
| threads=6 | |
| ) | |
| waveform = tts.generate_speech( | |
| "<ru>Привет от TeraTTS.</ru>", | |
| voice="ru_f1", | |
| duration_scale=1, | |
| ) | |
| tts.save_wav("teratts.wav", waveform) | |
| ``` | |
| `waveform` is a mono `float32` NumPy array at 44,100 Hz. `save_wav` writes | |
| standard signed-16-bit PCM WAV without an extra audio package. | |
| To inspect the exact text passed to the encoder after number expansion, stress | |
| marking, and Unicode normalization, call `tts.normalize_text(text)`. | |
| ## Controls | |
| | Control | Values | Effect | | |
| | --- | --- | --- | | |
| | `voice` | `ru_f1` ★, `ru_m5` ★, `ru_f2`, `ru_m1`, `eng_f3`, `eng_f4_whisper`, `eng_f5`, `eng_m2_whisper`, `eng_m3`, `eng_m4` | Selects a bundled precomputed voice style named after its reference audio. ★ marks the recommended Russian prompts. | | |
| | `duration_scale` | Positive float, default `1` | Higher values produce slower, longer speech. | | |
| | `diffusion_model` | `distilled` (default), `teacher` | Distilled is faster; teacher supports adjustable CFG. | | |
| | `ruaccent_mode` | `full` (default), `dictionary` | Full uses RUAccent neural ONNX graphs plus dictionaries; dictionary mode loads dictionaries only. | | |
| The default `diffusion_model="distilled"` is the fast eight-step sampler. To | |
| use the teacher sampler, choose it while loading: | |
| ```python | |
| teacher_tts = AutoModel.from_pretrained( | |
| "TeraSpace/TeraTTSv2", | |
| trust_remote_code=True, | |
| provider="CPUExecutionProvider", | |
| threads=6, | |
| diffusion_model="teacher", | |
| ) | |
| ``` | |
| `guidance` can be adjusted when generating with the teacher sampler. The | |
| distilled sampler has CFG 3 baked into its graph. | |
| ## Language tags, numbers, and Russian stress | |
| Language tags are required: wrap text in `<en>…</en>` or `<ru>…</ru>`. The | |
| runtime rejects untagged or unbalanced input with a tag-specific error. Before | |
| number expansion and stress marking, it inserts spaces after punctuation and | |
| between a number and a following word. Characters outside the model vocabulary | |
| are skipped with a runtime warning. Numbers inside language tags are expanded to | |
| words in the matching language before synthesis: | |
| ```python | |
| waveform = tts.generate_speech( | |
| "<ru>У меня 21 яблоко.</ru> <en>I have 42 apples.</en>", | |
| voice="ru_f1", | |
| duration_scale=1, | |
| ) | |
| ``` | |
| Russian text is automatically stress-marked by the bundled RUAccent-derived | |
| runtime. Manual `+` markers remain authoritative. For a lower-memory, | |
| deterministic dictionary-only path, choose the mode while loading: | |
| ```python | |
| dictionary_tts = AutoModel.from_pretrained( | |
| "TeraSpace/TeraTTSv2", | |
| trust_remote_code=True, | |
| provider="CPUExecutionProvider", | |
| threads=6, | |
| ruaccent_mode="dictionary", | |
| ) | |
| ``` | |
| Dictionary mode does not load RUAccent neural ONNX graphs. It marks known | |
| words and applies deterministic `ё` replacements, while unknown words and | |
| ambiguous homographs are left unchanged. Set `russian_stress=False` to disable | |
| automatic Russian stress processing entirely. | |
| When using an English voice such as `eng_f3` for Russian text, start by trying | |
| `duration_scale=0.8` and adjust by ear: | |
| ```python | |
| waveform = tts.generate_speech( | |
| "<ru>Это русский текст английским голосом.</ru>", | |
| voice="eng_f3", | |
| duration_scale=0.8, | |
| ) | |
| ``` | |
| ## Stream audio | |
| ```python | |
| for chunk in tts.generate_speech_stream( | |
| "<en>Streaming speech is ready.</en>", | |
| voice="eng_f3", | |
| duration_scale=1, | |
| ): | |
| # Send float32 mono chunks (44,100 Hz) to a player or network client. | |
| consume(chunk) | |
| ``` | |
| The remote code loads only the selected sampler graph plus shared ONNX graphs. | |
| For security, pin a specific Hub commit when using `trust_remote_code=True`. | |
| ## Attribution | |
| The local Russian stress annotator and its assets are adapted from | |
| [RUAccent](https://github.com/Den4ikAI/ruaccent), Copyright 2026 Denis Petrov, | |
| under the MIT License. See `RUACCENT_NOTICE.txt`. | |