Antalia 1 Foundation
The speaker-agnostic Turkish acoustic model behind Antalia 1, trained from scratch on public data only.
Türkçe: README.tr.md
Development is discontinued. Released as-is so that others can fine-tune their own consented Turkish voice with the code in https://github.com/0daycloud/antalia. Contact: sezgin@patientdesk.ai or GitHub issues.
At a glance
| What it is | The base model of Antalia 1: a ~300M-parameter character-conditioned rectified-flow model over 24 kHz, 100-band log-mel spectrograms. |
| Speaker | None. There is no speaker conditioning; the voice of any output is an unspecified average of Common Voice contributors and changes with the sampling seed. |
| Size | 299.6M parameters, fp32 safetensors, 1.20 GB. |
| Training data | Common Voice 26.0 Turkish (CC0) and FLEURS Turkish (CC-BY-4.0), 61,469 filtered clips, 67.55 h. Nothing else. |
| Intelligibility | CER 0.1835 single seed on the 120-prompt turkish-v2 suite. An intelligibility baseline, not a product voice. |
| Use it for | Fine-tuning a new consented Turkish voice with the reference lineage below. |
| License | Weights: Antalia OpenRAIL-M (LICENSE.md), which requires credit to the authors in anything that uses or redistributes the model. Code: Apache-2.0. |
Listen
Both clips are AI-generated with speaker id 0, the unconditioned path. Each is paired with the same sentence from the fine-tuned Antalia 1 so you can hear what the fine-tuning stages add. More on the sample site.
| Prompt | Takes |
|---|---|
| Sabah güneşi sessiz sokağın taşlarına yavaşça vuruyordu. General |
Foundation, unconditioned Antalia 1, best of 8 · CER 0.00 · Sim 0.972 |
| Size yardımcı olabilmem için randevu tarihini paylaşır mısınız? Voice agent |
Foundation, unconditioned Antalia 1, best of 8 · CER 0.00 · Sim 0.959 |
Quick start
git clone https://github.com/0daycloud/antalia
cd antalia
uv sync # or: pip install -e .
# Vocoder: BigVGAN is not redistributed. Pinned commit + one-line hub patch.
git clone https://github.com/NVIDIA/BigVGAN /opt/bigvgan
git -C /opt/bigvgan checkout 7d2b454564a6c7d014227f635b7423881f14bdac
patch -d /opt/bigvgan -p4 < scripts/patches/bigvgan-huggingface-hub-1.patch
export PYTHONPATH=/opt/bigvgan
python scripts/synthesize-crossflow.py \
--checkpoint cloud0day3/antalia-1-foundation \
--vocoder nvidia/bigvgan_v2_24khz_100band_256x \
--sway -0.8 --steps 32 --mel-clamp 5.0 \
--text "Bugün hava çok güzel, dışarıda yürüyüş yapmak istiyorum." \
--output foundation.wav
To fine-tune a new voice, follow the reference lineage in configs/crossflow/:
foundation-speaker-v2.json → candidate-b-speaker-v1.json → cfg-foundation-v1.json →
candidate-consistency-v1.json → candidate-b-timbre-adapter-v2.json. The voice corpus that
lineage was run on is public: antalia-voice-corpus
(5.008 h, CC-BY-4.0).
Training data
| Source | Clips used | License |
|---|---|---|
| Common Voice 26.0 Turkish (validated, filtered) | 59,593 | CC0-1.0 |
| FLEURS Turkish | 1,876 | CC-BY-4.0 |
Total 61,469 clips / 67.55 h. Filtering: forced-Turkish Whisper-large-v3 transcript check (accepted median CER 0, p90 0.111), active-audio, level, clipping, and SNR gates, per-speaker caps, speaker-disjoint validation split. Rejections from 120,407 validated CV clips: 7,098 insufficient active audio, 1,449 transcript mismatch, 356 too short, 351 low level, 5 low SNR, 4 clipping. Filter manifests with clip ids and checksums ship in the code repository; the audio itself is not re-hosted.
Training
Clean initialization; 100,000 updates on one A100-80GB; AdamW, LR 1e-4 with 2,000 warmup updates decaying to 1e-5, weight decay 0.01, gradient clip 1.0, EMA 0.9999; 6,000 mel frames per batch (≤32 samples, ≤24 s per clip); duration-loss weight 0.1; per-band mel mean/std normalization measured on 256 training records (89,966 frames), clipped at ±5. Two earlier foundation runs without per-band normalization diverged and were abandoned.
Quality
On the 120-prompt turkish-v2 suite, single seed, 32 Euler steps, no guidance:
| Metric | Foundation | Antalia 1 (single seed) |
|---|---|---|
| CER mean | 0.1835 | 0.0528 |
| CER p90 | 0.4049 | 0.1348 |
| WER mean | 0.3952 | 0.1297 |
This is an intelligibility baseline for a speaker-agnostic model with a sampled voice; it is not usable as a product voice on its own.
Files
| File | Purpose |
|---|---|
model.safetensors |
EMA weights, fp32, 299,623,013 parameters. SHA-256 89ecf310a14333c6bd5a754360cee2207027dd23b2a0171984a32c423afcfcd1 |
config.json |
Release format antalia-crossflow-release-v1: architecture, mel statistics, character vocabulary, vocoder pointer, provenance |
LICENSE.md |
Antalia OpenRAIL-M |
Vocoder: nvidia/bigvgan_v2_24khz_100band_256x (MIT), downloaded by the loader.
Architecture
16 Transformer blocks (dim 768, 12 heads, SwiGLU FFN 3072) with self-attention over mel frames, cross-attention over characters, and zero-initialized adaLN modulation from the flow timestep; 4-block ConvNeXt-style character encoder; a scalar log-total-duration head. Deterministic Turkish text normalizer; grapheme input.
Limitations
- Unconditioned voice identity is arbitrary and unstable across seeds.
- Long inputs are under-budgeted by the duration head; use clause chunking.
- Numbers, foreign names, and abbreviations have the highest error rates.
- Turkish only. No memorization audit was run.
License and use
Weights: Antalia Open RAIL-M (see LICENSE.md): no impersonation, no undisclosed synthetic
speech, no fraud or robocalls. Any distribution of the model or its derivatives, and any product,
service or publication that uses them, must credit "Antalia 1" by Sezgin Saygili, Emre Kaplaner,
Oncel Ozgul and Fikri San Koktas (Patientdesk.ai) with a link to this repository or the code
repository. Code: Apache-2.0.
Citation
@misc{antalia1_2026,
title = {Antalia 1: An Open Turkish Text-to-Speech Model from a Rights-Clean Pipeline},
author = {Saygili, Sezgin and Kaplaner, Emre and Ozgul, Oncel and Koktas, Fikri San},
year = {2026},
note = {Technical report},
url = {https://huggingface.co/cloud0day3/antalia-1-foundation}
}
- Downloads last month
- 69