ZamVoice / README.md
buumba641's picture
Update README.md
3e3b11e verified
|
Raw
History Blame Contribute Delete
3.41 kB
---
title: ZamVoice ASR Demo
emoji: πŸŽ™οΈ
colorFrom: green
colorTo: blue
sdk: gradio
sdk_version: "6.20.0"
app_file: app.py
pinned: false
---
# ZamVoice β€” Zambian Language ASR Demo
A free public demo for testing automatic speech recognition (ASR) in three
Zambian languages β€” Tonga, Bemba, and Nyanja β€” powered by fine-tuned Whisper
models. Built as part of a final-year research project comparative analysis of pretrained
ASR models on low-resource Zambian languages.
## Models used
| Language | Hugging Face repo |
|---|---|
| Tonga | `buumba641/tonga-whisper-medium-shona-proxy` |
| Bemba | `buumba641/bemba-whisper-medium-siwahili-final` |
| Nyanja | `buumba641/Nyanja-whisper-medium-shona-proxy` |
## Translation & speech (English output)
After transcribing, you can translate the result to English (via Gemini) and
listen to it spoken aloud (via ElevenLabs text-to-speech, which starts
playing automatically as soon as it's ready). Translation only goes one
direction: Tonga, Bemba, or Nyanja β†’ English. Gemini is told which of the
three languages it's translating (with enough context to tell Zambian
Chitonga apart from the unrelated Tongan language, for example), so it
doesn't have to guess from the text alone. This logic lives in its own file,
`translation.py`, kept separate from `app.py` for easier debugging.
### Required secrets (Space Settings β†’ Variables and secrets)
| Secret | Used for |
|---|---|
| `geminikey` | Gemini translation |
| `elevenlabskey` | ElevenLabs text-to-speech |
| `HF_TOKEN` | Saving transcripts and feedback to the training dataset |
Each of these degrades gracefully if its secret is missing β€” translation,
speech, and dataset saving each show a clear message (or just skip) instead
of crashing the app; transcription itself always works independently.
### Optional environment variables
| Variable | Default |
|---|---|
| `GEMINI_MODEL` | `gemini-3.6-flash` |
| `ELEVENLABS_VOICE_ID` | `JBFqnCBsd6RMkjVDRZzb` ("George") |
| `ELEVENLABS_MODEL_ID` | `eleven_v3` |
| `ELEVENLABS_OUTPUT_FORMAT` | `mp3_44100_128` |
## Feedback & the training dataset
Every transcription is uploaded to the `buumba641/Zambia-MultiLigual-ASR-Dataset`
dataset on the Hugging Face Hub. After transcribing, rate the transcript
with the star buttons β€” anything below 5 stars reveals a box where you can
type the correct transcript. Click **Submit Feedback** to save the rating
(and correction, if any) back to that sample's row, which is what actually
improves the models over time. This logic lives in `dataset_manager.py`.
## How it works
1. Pick a language from the dropdown.
2. Upload an audio file or record from your microphone.
3. Click **Transcribe** to see the model's output.
4. Rate the transcript (and correct it, if needed) to help improve the model.
5. Optionally translate it to English and listen to it spoken aloud.
## Files
- `app.py` β€” Gradio UI: language dropdown, audio input, transcription, star rating, and translation controls.
- `translation.py` β€” Gemini translation + ElevenLabs text-to-speech, kept separate for easier debugging.
- `dataset_manager.py` β€” uploads audio/transcripts and feedback (ratings + corrections) to the Hugging Face dataset.
- `requirements.txt` β€” Python dependencies.
## Running locally
```bash
pip install -r requirements.txt
python app.py
```
Then open the local URL Gradio prints in your terminal.