--- title: README emoji: 🎙️ colorFrom: indigo colorTo: blue sdk: static pinned: false --- # Silencio Network Consented, crowdsourced speech data. Contributors record on their own devices, in their own environments, under terms covering AI/ML training use, with a per-contributor consent record and a deletion pathway. Everything published here is a **sample**. The cards state exactly what shipped — real hour counts, real speaker counts, and a Limitations section on every release. Where a sample has no transcripts, the card says so rather than implying otherwise. ## Start here — human-validated transcription with word-level alignment Our current standard. Human-checked transcript text, forced-aligned per token, speaker-level language proficiency, dialect and demographics. | Dataset | Hours | Speakers | Notes | |---|---:|---:|---| | [**kenyan swahili**](https://huggingface.co/datasets/SilencioNetwork/swahili-speech) | 1.4 | **103** | Spontaneous. One clip per speaker — 66 Nairobi Sheng-influenced, 26 Mombasa coastal | | [**cebuano**](https://huggingface.co/datasets/SilencioNetwork/cebuano-speech) | 3.5 | 15 | Spontaneous long-form, mean clip over two minutes, 27,000+ timestamped tokens | | [**tagalog / filipino**](https://huggingface.co/datasets/SilencioNetwork/tagalog-filipino-speech) | 2.0 | 16 | Spontaneous, 13,782 timestamped tokens | These are evaluation sets, not training corpora. They are sized for measuring how a model performs across a population of speakers rather than for fitting one. ## Earlier samples — audio and metadata Published April 2026, rebuilt in August against what the repos actually contain. Most carry no transcripts; **human transcription with word-level alignment is available on request** for any subset, in the format shown above. [complete-voiceai](https://huggingface.co/datasets/SilencioNetwork/complete-voiceai-speech-dataset) · [amharic](https://huggingface.co/datasets/SilencioNetwork/amharic-speech) · [yoruba](https://huggingface.co/datasets/SilencioNetwork/yoruba-speech) · [hausa](https://huggingface.co/datasets/SilencioNetwork/hausa-speech) · [igbo](https://huggingface.co/datasets/SilencioNetwork/igbo-speech) · [global french](https://huggingface.co/datasets/SilencioNetwork/global-french-speech) · [medical](https://huggingface.co/datasets/SilencioNetwork/medical-speech-dataset) ## What sits behind the samples Two figures, because they answer different questions. **Recorded and available off the shelf** — already collected, with metadata, licensable today: | | | |---|---| | Hours recorded | **127,793** | | Recordings | **9,392,870** | | Contributors who recorded | **222,145** | | Languages | **156** | | Countries and territories of origin | **216** | **Contributor network for commissioned collection** — registered, consented contributors who can be activated for a specific brief: a language, a dialect region, a demographic, a recording condition, a speech style. These are *not* active contributors to the corpus above; they are the pool it is drawn from and extended through: | | | |---|---| | Registered contributors | **2,000,000+** | | Countries | **180+** | | Languages reachable | **250+** | The gap between the two is the point. Several East African languages we hold — Taita, Teso, Kikuyu, Luo, Kamba, Luyia, Kalenjin — have no dedicated speech dataset on this Hub at all. **In active collection:** 7,500 hours across Hiligaynon, Tagalog and Cebuano — 2,500 hours each, split 1,000 hours single-speaker and 1,500 hours multi-speaker. ## Get in touch Volume licensing, human-validated transcription over a larger subset, or commissioned collection in a language not listed here — **info@silencio.network** Licence on published samples is CC BY-NC 4.0. Non-commercial covers research, evaluation and publication; benchmarking a commercial product model is a commercial use and needs a licence, which is usually granted for evaluation. Ask.