TTSModel / README.md
MCAA1-MSU's picture
Upload README.md with huggingface_hub
245c6f0 verified
|
Raw
History Blame Contribute Delete
2.86 kB
metadata
language:
  - sw
license: cc-by-nc-4.0
base_model: facebook/mms-tts-swh
tags:
  - text-to-speech
  - tts
  - vits
  - swahili
  - mms
datasets:
  - google/WaxalNLP
pipeline_tag: text-to-speech

TTSModel — Swahili TTS

A Swahili text-to-speech model, finetuned from Meta's MMS-TTS Swahili checkpoint on the swa_tts split of google/WaxalNLP, using the VITS finetuning recipe from ylacombe/finetune-hf-vits. Developed by the Maseno Centre for Applied AI (MCAAI).

Model details

  • Base model: facebook/mms-tts-swh (Meta's Massively Multilingual Speech TTS, Swahili)
  • Architecture: VITS (single-speaker)
  • Training data: google/WaxalNLP, swa_tts config — 1,387 train utterances (805 after duration/length filtering), single Swahili speaker, 16kHz audio, sourced via the Loud & Clear initiative.
  • Training checkpoint: step 4,500 / 20,200 planned steps
  • Sample rate: 16,000 Hz
  • Language: Swahili (swh / ISO 639-3)

Training configuration

Setting Value
Learning rate 2e-5
Batch size 8
Precision fp16
Max clip duration 20s
Min clip duration 0.5s
Loss weights mel=35, kl=1.5, disc=3, gen/fmaps/duration=1

Training used the finetune-hf-vits recipe with transformers==4.35.1, datasets==2.14.7, accelerate==0.24.1, numpy<2.0.

Usage

import numpy as np
from transformers import pipeline
from IPython.display import Audio as IPyAudio

synthesiser = pipeline("text-to-speech", model="MCAA1-MSU/TTSModel")
speech = synthesiser("Habari yako, karibu Kenya.")

audio = np.squeeze(speech["audio"])
display(IPyAudio(audio, rate=speech["sampling_rate"]))

Verified loading cleanly with transformers pipeline("text-to-speech", ...) — no missing/unexpected weight warnings on load.

Intended use

Research and experimentation with Swahili TTS.

Limitations

  • Single speaker, single dataset: may not generalize well to varied Swahili dialects, accents, or speaking styles.
  • No formal evaluation yet: no MOS, WER, or MCD scores computed. Qualitative spot-checks suggest intelligible but rough output.
  • Small dataset: ~800 training utterances after filtering.

License

Derived from facebook/mms-tts-swh (CC-BY-NC-4.0, non-commercial). This finetuned model inherits that license. swa_tts training data is CC-BY-SA-4.0.

Acknowledgements