File size: 3,976 Bytes
20effd3
 
de64c45
 
 
20effd3
 
 
 
de64c45
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
---
title: README
emoji: 🎙️
colorFrom: indigo
colorTo: blue
sdk: static
pinned: false
---

# Silencio Network

Consented, crowdsourced speech data. Contributors record on their own devices, in their own
environments, under terms covering AI/ML training use, with a per-contributor consent record and
a deletion pathway.

Everything published here is a **sample**. The cards state exactly what shipped — real hour
counts, real speaker counts, and a Limitations section on every release. Where a sample has no
transcripts, the card says so rather than implying otherwise.

## Start here — human-validated transcription with word-level alignment

Our current standard. Human-checked transcript text, forced-aligned per token, speaker-level
language proficiency, dialect and demographics.

| Dataset | Hours | Speakers | Notes |
|---|---:|---:|---|
| [**kenyan swahili**](https://huggingface.co/datasets/SilencioNetwork/swahili-speech) | 1.4 | **103** | Spontaneous. One clip per speaker — 66 Nairobi Sheng-influenced, 26 Mombasa coastal |
| [**cebuano**](https://huggingface.co/datasets/SilencioNetwork/cebuano-speech) | 3.5 | 15 | Spontaneous long-form, mean clip over two minutes, 27,000+ timestamped tokens |
| [**tagalog / filipino**](https://huggingface.co/datasets/SilencioNetwork/tagalog-filipino-speech) | 2.0 | 16 | Spontaneous, 13,782 timestamped tokens |

These are evaluation sets, not training corpora. They are sized for measuring how a model
performs across a population of speakers rather than for fitting one.

## Earlier samples — audio and metadata

Published April 2026, rebuilt in August against what the repos actually contain. Most carry no
transcripts; **human transcription with word-level alignment is available on request** for any
subset, in the format shown above.

[complete-voiceai](https://huggingface.co/datasets/SilencioNetwork/complete-voiceai-speech-dataset) ·
[amharic](https://huggingface.co/datasets/SilencioNetwork/amharic-speech) ·
[yoruba](https://huggingface.co/datasets/SilencioNetwork/yoruba-speech) ·
[hausa](https://huggingface.co/datasets/SilencioNetwork/hausa-speech) ·
[igbo](https://huggingface.co/datasets/SilencioNetwork/igbo-speech) ·
[global french](https://huggingface.co/datasets/SilencioNetwork/global-french-speech) ·
[medical](https://huggingface.co/datasets/SilencioNetwork/medical-speech-dataset)

## What sits behind the samples

Two figures, because they answer different questions.

**Recorded and available off the shelf** — already collected, with metadata, licensable today:

| | |
|---|---|
| Hours recorded | **127,793** |
| Recordings | **9,392,870** |
| Contributors who recorded | **222,145** |
| Languages | **156** |
| Countries and territories of origin | **216** |

**Contributor network for commissioned collection** — registered, consented contributors who can
be activated for a specific brief: a language, a dialect region, a demographic, a recording
condition, a speech style. These are *not* active contributors to the corpus above; they are the
pool it is drawn from and extended through:

| | |
|---|---|
| Registered contributors | **2,000,000+** |
| Countries | **180+** |
| Languages reachable | **250+** |

The gap between the two is the point. Several East African languages we hold — Taita, Teso,
Kikuyu, Luo, Kamba, Luyia, Kalenjin — have no dedicated speech dataset on this Hub at all.

**In active collection:** 7,500 hours across Hiligaynon, Tagalog and Cebuano — 2,500 hours each,
split 1,000 hours single-speaker and 1,500 hours multi-speaker.

## Get in touch

Volume licensing, human-validated transcription over a larger subset, or commissioned collection
in a language not listed here — **info@silencio.network**

Licence on published samples is CC BY-NC 4.0. Non-commercial covers research, evaluation and
publication; benchmarking a commercial product model is a commercial use and needs a licence,
which is usually granted for evaluation. Ask.