MasriSwitch-TTS model card

Status and intended use

Experimental Egyptian Arabic and English code-switched text-to-speech for research and voice-agent prototyping. Audio is AI-generated. Use a reference voice only with documented speaker consent; no public reference recording is included.

Base, training data, and provenance

Base checkpoint: SILMA TTS v1 (Apache-2.0 weights), using F5-TTS 1.1.7 (MIT code). Training uses only audited D1 rows labeled CC BY 4.0 and generated by the dataset author. Noncommercial D1 rows were excluded. D1 revision: eae9a87c17e91e3f59a9696d5f4ff3eb51502e82; selected train rows: 3414; seed: 42. Selected checkpoint SHA256: 558e2ab53e1b5450bcd1a1c30683a1b3e6be1234b7e198b74209f362693eaf92. Checkpoint was chosen on validation prompts by English EER, then code-switch WER, with an overall Arabic CER guardrail. Training reached 5,000 of 8,000 planned optimizer updates; the selected checkpoint is from update 2,000. Training stopped after three stale validation checks. Attribution: Abdelrahman R. Hashem, Synthetic Arabic-English Code-Switched Speech for ASR (2026). See NOTICE and the data audit.

Evaluation

The same 189 validation and 300 locked MasriSwitch-Bench prompts, private reference voice, 16 inference steps, and pinned independent ASR were used for E0/E1.

Model Code-switch WER English EER Arabic-only CER Critical entity accuracy Mean speaker cosine Mean RTF
E0 SILMA 61.38% 51.60% 38.57% 10.19% 0.728 0.308
E1 selected 61.34% 52.53% 38.03% 10.19% 0.728 0.305

These metrics are ASR proxies, not human pronunciation or naturalness scores. See EVALUATION.md for confidence intervals and errors. The declared accuracy target (≥15% relative code-switch WER reduction or ≥25% English EER reduction) was not met. This release remains experimental.

Limitations and misuse

Training speech is synthetic and has no speaker IDs, so real Egyptian speaker, accent, and acoustic robustness are unproven. English entities, numbers, and dates can be mispronounced. Do not use for consequential decisions without human verification. Do not impersonate, deceive, or commit fraud. Disclose generated audio and obtain consent before synthesizing a custom voice.

Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Tarek737/MasriSwitch-TTS

Finetuned
(2)
this model

Space using Tarek737/MasriSwitch-TTS 1