TTSModel / README.md
Mwau's picture
Upload README.md with huggingface_hub
575586f verified
|
Raw
History Blame
2.68 kB
metadata
language:
  - sw
license: cc-by-nc-4.0
base_model: facebook/mms-tts-swh
tags:
  - text-to-speech
  - tts
  - vits
  - swahili
  - mms
datasets:
  - google/WaxalNLP
pipeline_tag: text-to-speech

TTSModel — Swahili TTS

A Swahili text-to-speech model, finetuned from Meta's MMS-TTS Swahili checkpoint on the swa_tts split of google/WaxalNLP, using the VITS finetuning recipe from ylacombe/finetune-hf-vits. Developed by the Maseno Centre for Applied AI (MCAAI).

Model details

  • Base model: facebook/mms-tts-swh (Meta's Massively Multilingual Speech TTS, Swahili)
  • Architecture: VITS (single-speaker)
  • Training data: google/WaxalNLP, swa_tts config — 1,387 train utterances (805 after duration/length filtering), single Swahili speaker, 16kHz audio, sourced via the Loud & Clear initiative.
  • Training checkpoint: step 4,500 / 20,200 planned steps (~22 epochs of ~805 examples)
  • Sample rate: 16,000 Hz
  • Language: Swahili (swh / ISO 639-3)

Training configuration

Setting Value
Learning rate 2e-5
Batch size 8
Precision fp16
Max clip duration 20s
Min clip duration 0.5s
Loss weights mel=35, kl=1.5, disc=3, gen/fmaps/duration=1

Training used the finetune-hf-vits recipe with transformers==4.35.1, datasets==2.14.7, accelerate==0.24.1, numpy<2.0.

Usage

import numpy as np
from transformers import pipeline
import scipy.io.wavfile

synthesiser = pipeline("text-to-speech", model="MCAA1-MSU/TTSModel")
speech = synthesiser("Habari yako, karibu Kenya.")

audio = np.squeeze(speech["audio"])
scipy.io.wavfile.write("output.wav", rate=speech["sampling_rate"], data=audio)

Verified loading cleanly with transformers pipeline("text-to-speech", ...) — no missing/unexpected weight warnings on load.

Intended use

Research and experimentation with Swahili TTS. Not yet suitable for production or user-facing applications given the early training stage.

License

Derived from facebook/mms-tts-swh (CC-BY-NC-4.0, non-commercial). This finetuned model inherits that license. swa_tts training data is CC-BY-SA-4.0.

Acknowledgements