--- library_name: transformers pipeline_tag: text-to-speech tags: - onnxruntime - text-to-speech - russian - english - custom-code --- # TeraTTSv2 ONNX TeraTTSv2 is a self-contained ONNX Runtime text-to-speech release with selectable diffusion samplers, ten voice styles, Russian stress marking, and streamed audio output. This release uses the clean English/Russian 25-second teacher and its matching eight-step CFG-3 distilled student. > **Important — Russian stress is automatic.** Text inside `` receives > stress markers automatically by default. Explicit `+` markers always win. > > **Important — cross-language prompts.** When an English reference voice is > speaking Russian, experiment with `duration_scale` below `1` (for example > `0.8`). It is usually a better starting point than the default `1`. > > **Recommended voices:** `ru_f1` and `ru_m5` are the preferred Russian voice > prompts. ## Installation ```bash pip install -r requirements.txt ``` `sounddevice` is only required for direct speaker playback. On Linux, install the system PortAudio library if it is not already present. ## Load with Transformers ```python from transformers import AutoModel tts = AutoModel.from_pretrained( "TeraSpace/TeraTTSv2", trust_remote_code=True, provider="CPUExecutionProvider", threads=6 ) waveform = tts.generate_speech( "Привет от TeraTTS.", voice="ru_f1", duration_scale=1, ) tts.save_wav("teratts.wav", waveform) ``` `waveform` is a mono `float32` NumPy array at 44,100 Hz. `save_wav` writes standard signed-16-bit PCM WAV without an extra audio package. To inspect the exact text passed to the encoder after number expansion, stress marking, and Unicode normalization, call `tts.normalize_text(text)`. ## Controls | Control | Values | Effect | | --- | --- | --- | | `voice` | `ru_f1` ★, `ru_m5` ★, `ru_f2`, `ru_m1`, `eng_f3`, `eng_f4_whisper`, `eng_f5`, `eng_m2_whisper`, `eng_m3`, `eng_m4` | Selects a bundled precomputed voice style named after its reference audio. ★ marks the recommended Russian prompts. | | `duration_scale` | Positive float, default `1` | Higher values produce slower, longer speech. | | `diffusion_model` | `distilled` (default), `teacher` | Distilled is faster; teacher supports adjustable CFG. | | `ruaccent_mode` | `full` (default), `dictionary` | Full uses RUAccent neural ONNX graphs plus dictionaries; dictionary mode loads dictionaries only. | The default `diffusion_model="distilled"` is the fast eight-step sampler. To use the teacher sampler, choose it while loading: ```python teacher_tts = AutoModel.from_pretrained( "TeraSpace/TeraTTSv2", trust_remote_code=True, provider="CPUExecutionProvider", threads=6, diffusion_model="teacher", ) ``` `guidance` can be adjusted when generating with the teacher sampler. The distilled sampler has CFG 3 baked into its graph. ## Language tags, numbers, and Russian stress Language tags are required: wrap text in `` or ``. The runtime rejects untagged or unbalanced input with a tag-specific error. Before number expansion and stress marking, it inserts spaces after punctuation and between a number and a following word. Characters outside the model vocabulary are skipped with a runtime warning. Numbers inside language tags are expanded to words in the matching language before synthesis: ```python waveform = tts.generate_speech( "У меня 21 яблоко. I have 42 apples.", voice="ru_f1", duration_scale=1, ) ``` Russian text is automatically stress-marked by the bundled RUAccent-derived runtime. Manual `+` markers remain authoritative. For a lower-memory, deterministic dictionary-only path, choose the mode while loading: ```python dictionary_tts = AutoModel.from_pretrained( "TeraSpace/TeraTTSv2", trust_remote_code=True, provider="CPUExecutionProvider", threads=6, ruaccent_mode="dictionary", ) ``` Dictionary mode does not load RUAccent neural ONNX graphs. It marks known words and applies deterministic `ё` replacements, while unknown words and ambiguous homographs are left unchanged. Set `russian_stress=False` to disable automatic Russian stress processing entirely. When using an English voice such as `eng_f3` for Russian text, start by trying `duration_scale=0.8` and adjust by ear: ```python waveform = tts.generate_speech( "Это русский текст английским голосом.", voice="eng_f3", duration_scale=0.8, ) ``` ## Stream audio ```python for chunk in tts.generate_speech_stream( "Streaming speech is ready.", voice="eng_f3", duration_scale=1, ): # Send float32 mono chunks (44,100 Hz) to a player or network client. consume(chunk) ``` The remote code loads only the selected sampler graph plus shared ONNX graphs. For security, pin a specific Hub commit when using `trust_remote_code=True`. ## Attribution The local Russian stress annotator and its assets are adapted from [RUAccent](https://github.com/Den4ikAI/ruaccent), Copyright 2026 Denis Petrov, under the MIT License. See `RUACCENT_NOTICE.txt`.