README / README.md
ASMSIlencio's picture
Organization card
de64c45 verified
|
Raw
History Blame Contribute Delete
3.98 kB
---
title: README
emoji: πŸŽ™οΈ
colorFrom: indigo
colorTo: blue
sdk: static
pinned: false
---
# Silencio Network
Consented, crowdsourced speech data. Contributors record on their own devices, in their own
environments, under terms covering AI/ML training use, with a per-contributor consent record and
a deletion pathway.
Everything published here is a **sample**. The cards state exactly what shipped β€” real hour
counts, real speaker counts, and a Limitations section on every release. Where a sample has no
transcripts, the card says so rather than implying otherwise.
## Start here β€” human-validated transcription with word-level alignment
Our current standard. Human-checked transcript text, forced-aligned per token, speaker-level
language proficiency, dialect and demographics.
| Dataset | Hours | Speakers | Notes |
|---|---:|---:|---|
| [**kenyan swahili**](https://huggingface.co/datasets/SilencioNetwork/swahili-speech) | 1.4 | **103** | Spontaneous. One clip per speaker β€” 66 Nairobi Sheng-influenced, 26 Mombasa coastal |
| [**cebuano**](https://huggingface.co/datasets/SilencioNetwork/cebuano-speech) | 3.5 | 15 | Spontaneous long-form, mean clip over two minutes, 27,000+ timestamped tokens |
| [**tagalog / filipino**](https://huggingface.co/datasets/SilencioNetwork/tagalog-filipino-speech) | 2.0 | 16 | Spontaneous, 13,782 timestamped tokens |
These are evaluation sets, not training corpora. They are sized for measuring how a model
performs across a population of speakers rather than for fitting one.
## Earlier samples β€” audio and metadata
Published April 2026, rebuilt in August against what the repos actually contain. Most carry no
transcripts; **human transcription with word-level alignment is available on request** for any
subset, in the format shown above.
[complete-voiceai](https://huggingface.co/datasets/SilencioNetwork/complete-voiceai-speech-dataset) Β·
[amharic](https://huggingface.co/datasets/SilencioNetwork/amharic-speech) Β·
[yoruba](https://huggingface.co/datasets/SilencioNetwork/yoruba-speech) Β·
[hausa](https://huggingface.co/datasets/SilencioNetwork/hausa-speech) Β·
[igbo](https://huggingface.co/datasets/SilencioNetwork/igbo-speech) Β·
[global french](https://huggingface.co/datasets/SilencioNetwork/global-french-speech) Β·
[medical](https://huggingface.co/datasets/SilencioNetwork/medical-speech-dataset)
## What sits behind the samples
Two figures, because they answer different questions.
**Recorded and available off the shelf** β€” already collected, with metadata, licensable today:
| | |
|---|---|
| Hours recorded | **127,793** |
| Recordings | **9,392,870** |
| Contributors who recorded | **222,145** |
| Languages | **156** |
| Countries and territories of origin | **216** |
**Contributor network for commissioned collection** β€” registered, consented contributors who can
be activated for a specific brief: a language, a dialect region, a demographic, a recording
condition, a speech style. These are *not* active contributors to the corpus above; they are the
pool it is drawn from and extended through:
| | |
|---|---|
| Registered contributors | **2,000,000+** |
| Countries | **180+** |
| Languages reachable | **250+** |
The gap between the two is the point. Several East African languages we hold β€” Taita, Teso,
Kikuyu, Luo, Kamba, Luyia, Kalenjin β€” have no dedicated speech dataset on this Hub at all.
**In active collection:** 7,500 hours across Hiligaynon, Tagalog and Cebuano β€” 2,500 hours each,
split 1,000 hours single-speaker and 1,500 hours multi-speaker.
## Get in touch
Volume licensing, human-validated transcription over a larger subset, or commissioned collection
in a language not listed here β€” **info@silencio.network**
Licence on published samples is CC BY-NC 4.0. Non-commercial covers research, evaluation and
publication; benchmarking a commercial product model is a commercial use and needs a licence,
which is usually granted for evaluation. Ask.