Spaces:
Running
title: README
emoji: 🎙️
colorFrom: indigo
colorTo: blue
sdk: static
pinned: false
Silencio Network
Consented, crowdsourced speech data. Contributors record on their own devices, in their own environments, under terms covering AI/ML training use, with a per-contributor consent record and a deletion pathway.
Everything published here is a sample. The cards state exactly what shipped — real hour counts, real speaker counts, and a Limitations section on every release. Where a sample has no transcripts, the card says so rather than implying otherwise.
Start here — human-validated transcription with word-level alignment
Our current standard. Human-checked transcript text, forced-aligned per token, speaker-level language proficiency, dialect and demographics.
| Dataset | Hours | Speakers | Notes |
|---|---|---|---|
| kenyan swahili | 1.4 | 103 | Spontaneous. One clip per speaker — 66 Nairobi Sheng-influenced, 26 Mombasa coastal |
| cebuano | 3.5 | 15 | Spontaneous long-form, mean clip over two minutes, 27,000+ timestamped tokens |
| tagalog / filipino | 2.0 | 16 | Spontaneous, 13,782 timestamped tokens |
These are evaluation sets, not training corpora. They are sized for measuring how a model performs across a population of speakers rather than for fitting one.
Earlier samples — audio and metadata
Published April 2026, rebuilt in August against what the repos actually contain. Most carry no transcripts; human transcription with word-level alignment is available on request for any subset, in the format shown above.
complete-voiceai · amharic · yoruba · hausa · igbo · global french · medical
What sits behind the samples
Two figures, because they answer different questions.
Recorded and available off the shelf — already collected, with metadata, licensable today:
| Hours recorded | 127,793 |
| Recordings | 9,392,870 |
| Contributors who recorded | 222,145 |
| Languages | 156 |
| Countries and territories of origin | 216 |
Contributor network for commissioned collection — registered, consented contributors who can be activated for a specific brief: a language, a dialect region, a demographic, a recording condition, a speech style. These are not active contributors to the corpus above; they are the pool it is drawn from and extended through:
| Registered contributors | 2,000,000+ |
| Countries | 180+ |
| Languages reachable | 250+ |
The gap between the two is the point. Several East African languages we hold — Taita, Teso, Kikuyu, Luo, Kamba, Luyia, Kalenjin — have no dedicated speech dataset on this Hub at all.
In active collection: 7,500 hours across Hiligaynon, Tagalog and Cebuano — 2,500 hours each, split 1,000 hours single-speaker and 1,500 hours multi-speaker.
Get in touch
Volume licensing, human-validated transcription over a larger subset, or commissioned collection in a language not listed here — info@silencio.network
Licence on published samples is CC BY-NC 4.0. Non-commercial covers research, evaluation and publication; benchmarking a commercial product model is a commercial use and needs a licence, which is usually granted for evaluation. Ask.