Spaces:
Running
Running
| title: README | |
| emoji: ποΈ | |
| colorFrom: indigo | |
| colorTo: blue | |
| sdk: static | |
| pinned: false | |
| # Silencio Network | |
| Consented, crowdsourced speech data. Contributors record on their own devices, in their own | |
| environments, under terms covering AI/ML training use, with a per-contributor consent record and | |
| a deletion pathway. | |
| Everything published here is a **sample**. The cards state exactly what shipped β real hour | |
| counts, real speaker counts, and a Limitations section on every release. Where a sample has no | |
| transcripts, the card says so rather than implying otherwise. | |
| ## Start here β human-validated transcription with word-level alignment | |
| Our current standard. Human-checked transcript text, forced-aligned per token, speaker-level | |
| language proficiency, dialect and demographics. | |
| | Dataset | Hours | Speakers | Notes | | |
| |---|---:|---:|---| | |
| | [**kenyan swahili**](https://huggingface.co/datasets/SilencioNetwork/swahili-speech) | 1.4 | **103** | Spontaneous. One clip per speaker β 66 Nairobi Sheng-influenced, 26 Mombasa coastal | | |
| | [**cebuano**](https://huggingface.co/datasets/SilencioNetwork/cebuano-speech) | 3.5 | 15 | Spontaneous long-form, mean clip over two minutes, 27,000+ timestamped tokens | | |
| | [**tagalog / filipino**](https://huggingface.co/datasets/SilencioNetwork/tagalog-filipino-speech) | 2.0 | 16 | Spontaneous, 13,782 timestamped tokens | | |
| These are evaluation sets, not training corpora. They are sized for measuring how a model | |
| performs across a population of speakers rather than for fitting one. | |
| ## Earlier samples β audio and metadata | |
| Published April 2026, rebuilt in August against what the repos actually contain. Most carry no | |
| transcripts; **human transcription with word-level alignment is available on request** for any | |
| subset, in the format shown above. | |
| [complete-voiceai](https://huggingface.co/datasets/SilencioNetwork/complete-voiceai-speech-dataset) Β· | |
| [amharic](https://huggingface.co/datasets/SilencioNetwork/amharic-speech) Β· | |
| [yoruba](https://huggingface.co/datasets/SilencioNetwork/yoruba-speech) Β· | |
| [hausa](https://huggingface.co/datasets/SilencioNetwork/hausa-speech) Β· | |
| [igbo](https://huggingface.co/datasets/SilencioNetwork/igbo-speech) Β· | |
| [global french](https://huggingface.co/datasets/SilencioNetwork/global-french-speech) Β· | |
| [medical](https://huggingface.co/datasets/SilencioNetwork/medical-speech-dataset) | |
| ## What sits behind the samples | |
| Two figures, because they answer different questions. | |
| **Recorded and available off the shelf** β already collected, with metadata, licensable today: | |
| | | | | |
| |---|---| | |
| | Hours recorded | **127,793** | | |
| | Recordings | **9,392,870** | | |
| | Contributors who recorded | **222,145** | | |
| | Languages | **156** | | |
| | Countries and territories of origin | **216** | | |
| **Contributor network for commissioned collection** β registered, consented contributors who can | |
| be activated for a specific brief: a language, a dialect region, a demographic, a recording | |
| condition, a speech style. These are *not* active contributors to the corpus above; they are the | |
| pool it is drawn from and extended through: | |
| | | | | |
| |---|---| | |
| | Registered contributors | **2,000,000+** | | |
| | Countries | **180+** | | |
| | Languages reachable | **250+** | | |
| The gap between the two is the point. Several East African languages we hold β Taita, Teso, | |
| Kikuyu, Luo, Kamba, Luyia, Kalenjin β have no dedicated speech dataset on this Hub at all. | |
| **In active collection:** 7,500 hours across Hiligaynon, Tagalog and Cebuano β 2,500 hours each, | |
| split 1,000 hours single-speaker and 1,500 hours multi-speaker. | |
| ## Get in touch | |
| Volume licensing, human-validated transcription over a larger subset, or commissioned collection | |
| in a language not listed here β **info@silencio.network** | |
| Licence on published samples is CC BY-NC 4.0. Non-commercial covers research, evaluation and | |
| publication; benchmarking a commercial product model is a commercial use and needs a licence, | |
| which is usually granted for evaluation. Ask. | |