Spaces:
Running
Running
File size: 3,976 Bytes
20effd3 de64c45 20effd3 de64c45 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 | ---
title: README
emoji: 🎙️
colorFrom: indigo
colorTo: blue
sdk: static
pinned: false
---
# Silencio Network
Consented, crowdsourced speech data. Contributors record on their own devices, in their own
environments, under terms covering AI/ML training use, with a per-contributor consent record and
a deletion pathway.
Everything published here is a **sample**. The cards state exactly what shipped — real hour
counts, real speaker counts, and a Limitations section on every release. Where a sample has no
transcripts, the card says so rather than implying otherwise.
## Start here — human-validated transcription with word-level alignment
Our current standard. Human-checked transcript text, forced-aligned per token, speaker-level
language proficiency, dialect and demographics.
| Dataset | Hours | Speakers | Notes |
|---|---:|---:|---|
| [**kenyan swahili**](https://huggingface.co/datasets/SilencioNetwork/swahili-speech) | 1.4 | **103** | Spontaneous. One clip per speaker — 66 Nairobi Sheng-influenced, 26 Mombasa coastal |
| [**cebuano**](https://huggingface.co/datasets/SilencioNetwork/cebuano-speech) | 3.5 | 15 | Spontaneous long-form, mean clip over two minutes, 27,000+ timestamped tokens |
| [**tagalog / filipino**](https://huggingface.co/datasets/SilencioNetwork/tagalog-filipino-speech) | 2.0 | 16 | Spontaneous, 13,782 timestamped tokens |
These are evaluation sets, not training corpora. They are sized for measuring how a model
performs across a population of speakers rather than for fitting one.
## Earlier samples — audio and metadata
Published April 2026, rebuilt in August against what the repos actually contain. Most carry no
transcripts; **human transcription with word-level alignment is available on request** for any
subset, in the format shown above.
[complete-voiceai](https://huggingface.co/datasets/SilencioNetwork/complete-voiceai-speech-dataset) ·
[amharic](https://huggingface.co/datasets/SilencioNetwork/amharic-speech) ·
[yoruba](https://huggingface.co/datasets/SilencioNetwork/yoruba-speech) ·
[hausa](https://huggingface.co/datasets/SilencioNetwork/hausa-speech) ·
[igbo](https://huggingface.co/datasets/SilencioNetwork/igbo-speech) ·
[global french](https://huggingface.co/datasets/SilencioNetwork/global-french-speech) ·
[medical](https://huggingface.co/datasets/SilencioNetwork/medical-speech-dataset)
## What sits behind the samples
Two figures, because they answer different questions.
**Recorded and available off the shelf** — already collected, with metadata, licensable today:
| | |
|---|---|
| Hours recorded | **127,793** |
| Recordings | **9,392,870** |
| Contributors who recorded | **222,145** |
| Languages | **156** |
| Countries and territories of origin | **216** |
**Contributor network for commissioned collection** — registered, consented contributors who can
be activated for a specific brief: a language, a dialect region, a demographic, a recording
condition, a speech style. These are *not* active contributors to the corpus above; they are the
pool it is drawn from and extended through:
| | |
|---|---|
| Registered contributors | **2,000,000+** |
| Countries | **180+** |
| Languages reachable | **250+** |
The gap between the two is the point. Several East African languages we hold — Taita, Teso,
Kikuyu, Luo, Kamba, Luyia, Kalenjin — have no dedicated speech dataset on this Hub at all.
**In active collection:** 7,500 hours across Hiligaynon, Tagalog and Cebuano — 2,500 hours each,
split 1,000 hours single-speaker and 1,500 hours multi-speaker.
## Get in touch
Volume licensing, human-validated transcription over a larger subset, or commissioned collection
in a language not listed here — **info@silencio.network**
Licence on published samples is CC BY-NC 4.0. Non-commercial covers research, evaluation and
publication; benchmarking a commercial product model is a commercial use and needs a licence,
which is usually granted for evaluation. Ask.
|