File size: 3,298 Bytes
8d40de6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
319613b
8d40de6
319613b
8d40de6
 
 
 
 
319613b
 
d124e07
319613b
d124e07
319613b
 
 
8d40de6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b880ca8
8d40de6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
319613b
82ef16b
 
319613b
 
82ef16b
 
 
 
 
 
8d40de6
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
---
language:
  - tr
license: apache-2.0
pipeline_tag: text-to-speech
tags:
  - text-to-speech
  - tts
  - turkish
  - flow-matching
  - diffusion-transformer
  - speech-synthesis
library_name: freyatts
---

# FreyaTTS-small

FreyaTTS-small is a 183M-parameter Turkish text-to-speech model, and the open-source member of the FreyaTTS family. It is tokenizer-free at the character level (92 symbols, no phonemizer or G2P) and generates speech with a non-autoregressive conditional flow-matching DiT in the frozen [AudioVAE2](https://huggingface.co/openbmb/VoxCPM2) latent space (25 Hz, 64-dim latents, 16 kHz encode / 48 kHz decode). Output is 48 kHz mono.

- **Repository:** https://github.com/freyavoiceai/FreyaTTS
- **Eval set:** https://huggingface.co/datasets/freyavoice/freya-tr-eval
- **License:** Apache-2.0

## The FreyaTTS family

**FreyaTTS-small** (this model) is our compact, Apache-2.0, self-hostable model, released in full: weights, inference code, and training pipeline.

**FreyaTTS-large** is our production model, serving Turkish voice agents at [Freya (YC S25)](https://freyavoice.ai) with higher naturalness and expressivity. It is available commercially rather than as open weights. For access, contact us at **dev@freyavoice.ai**.

The model id `freyavoice/freya-tts` predates this naming and is unchanged; it refers to FreyaTTS-small.

## Usage

```python
from freyatts import FreyaTTS

tts = FreyaTTS.from_pretrained("freyavoice/freya-tts", device="cuda")
wav = tts.synthesize("Merhaba, size nasıl yardımcı olabilirim?")   # np.float32, 48 kHz
tts.save_wav(wav, "output.wav")
```

Clone [freyavoiceai/FreyaTTS](https://github.com/freyavoiceai/FreyaTTS) and install `requirements.txt`.

## Model details

- **Architecture:** conditional flow-matching diffusion transformer, non-autoregressive, 32-step Euler ODE, no CFG
- **Parameters:** 183.2M
- **Input:** character-level Turkish, 92-symbol vocabulary
- **Latent space:** frozen AudioVAE2 (Apache-2.0, [openbmb/VoxCPM2](https://huggingface.co/openbmb/VoxCPM2)), 64-dim at 25 Hz, decodes to 48 kHz
- **Training:** from scratch on Turkish speech; pretraining followed by SFT stage 1/2 (voice lock, short-utterance coverage)
- **Voice:** single target speaker, no cloning

## Evaluation

On [Freya-TR-Eval](https://huggingface.co/datasets/freyavoice/freya-tr-eval): **WER 8.0% / CER 3.0%**, ranking 3rd of 7 among open sub-1B Turkish TTS models, ahead of XTTS-v2 (11.1% WER) and F5-TTS (24.3% WER).

## Speed

- RTX 4090: RTF 0.10-0.11, TTFT ~0.5 s, 1.5 GB VRAM, 9.4 audio-s/s at concurrency 4
- About 3.2x faster RTF and 3.7x less VRAM than the 2B VoxCPM2
- Apple M3 laptop CPU: RTF 0.70 (fp32); ~0.12 end to end via Core ML on Apple silicon

## Citation

FreyaTTS-small is described in our technical report, [arXiv:2607.09530](https://arxiv.org/abs/2607.09530):

```bibtex
@misc{pamuk2026freyattscompacttokenizerfreeflowmatching,
      title={FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis}, 
      author={Ahmet Erdem Pamuk and Ömer Yentür and Ahmet Tunga Bayrak and Yavuz Alp Sencer Öztürk and Mustafa Yavuz},
      year={2026},
      eprint={2607.09530},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2607.09530}, 
}
```