MasriSwitch-TTS model card
Status and intended use
Experimental Egyptian Arabic and English code-switched text-to-speech for research and voice-agent prototyping. Audio is AI-generated. Use a reference voice only with documented speaker consent; no public reference recording is included.
Base, training data, and provenance
Base checkpoint: SILMA TTS v1 (Apache-2.0 weights), using F5-TTS 1.1.7 (MIT code). Training uses only audited D1 rows labeled CC BY 4.0 and generated by the dataset author. Noncommercial D1 rows were excluded.
D1 revision: eae9a87c17e91e3f59a9696d5f4ff3eb51502e82; selected train rows: 3414; seed: 42.
Selected checkpoint SHA256: 558e2ab53e1b5450bcd1a1c30683a1b3e6be1234b7e198b74209f362693eaf92. Checkpoint was chosen on validation prompts by English EER, then code-switch WER, with an overall Arabic CER guardrail.
Training reached 5,000 of 8,000 planned optimizer updates; the selected checkpoint is from update 2,000. Training stopped after three stale validation checks.
Attribution: Abdelrahman R. Hashem, Synthetic Arabic-English Code-Switched Speech for ASR (2026). See NOTICE and the data audit.
Evaluation
The same 189 validation and 300 locked MasriSwitch-Bench prompts, private reference voice, 16 inference steps, and pinned independent ASR were used for E0/E1.
| Model | Code-switch WER | English EER | Arabic-only CER | Critical entity accuracy | Mean speaker cosine | Mean RTF |
|---|---|---|---|---|---|---|
| E0 SILMA | 61.38% | 51.60% | 38.57% | 10.19% | 0.728 | 0.308 |
| E1 selected | 61.34% | 52.53% | 38.03% | 10.19% | 0.728 | 0.305 |
These metrics are ASR proxies, not human pronunciation or naturalness scores. See EVALUATION.md for confidence intervals and errors.
The declared accuracy target (≥15% relative code-switch WER reduction or ≥25% English EER reduction) was not met. This release remains experimental.
Limitations and misuse
Training speech is synthetic and has no speaker IDs, so real Egyptian speaker, accent, and acoustic robustness are unproven. English entities, numbers, and dates can be mispronounced. Do not use for consequential decisions without human verification. Do not impersonate, deceive, or commit fraud. Disclose generated audio and obtain consent before synthesizing a custom voice.
- Downloads last month
- 13
Model tree for Tarek737/MasriSwitch-TTS
Base model
silma-ai/silma-tts