README / README.md
ASMSIlencio's picture
Organization card
de64c45 verified
|
Raw
History Blame Contribute Delete
3.98 kB
metadata
title: README
emoji: 🎙️
colorFrom: indigo
colorTo: blue
sdk: static
pinned: false

Silencio Network

Consented, crowdsourced speech data. Contributors record on their own devices, in their own environments, under terms covering AI/ML training use, with a per-contributor consent record and a deletion pathway.

Everything published here is a sample. The cards state exactly what shipped — real hour counts, real speaker counts, and a Limitations section on every release. Where a sample has no transcripts, the card says so rather than implying otherwise.

Start here — human-validated transcription with word-level alignment

Our current standard. Human-checked transcript text, forced-aligned per token, speaker-level language proficiency, dialect and demographics.

Dataset Hours Speakers Notes
kenyan swahili 1.4 103 Spontaneous. One clip per speaker — 66 Nairobi Sheng-influenced, 26 Mombasa coastal
cebuano 3.5 15 Spontaneous long-form, mean clip over two minutes, 27,000+ timestamped tokens
tagalog / filipino 2.0 16 Spontaneous, 13,782 timestamped tokens

These are evaluation sets, not training corpora. They are sized for measuring how a model performs across a population of speakers rather than for fitting one.

Earlier samples — audio and metadata

Published April 2026, rebuilt in August against what the repos actually contain. Most carry no transcripts; human transcription with word-level alignment is available on request for any subset, in the format shown above.

complete-voiceai · amharic · yoruba · hausa · igbo · global french · medical

What sits behind the samples

Two figures, because they answer different questions.

Recorded and available off the shelf — already collected, with metadata, licensable today:

Hours recorded 127,793
Recordings 9,392,870
Contributors who recorded 222,145
Languages 156
Countries and territories of origin 216

Contributor network for commissioned collection — registered, consented contributors who can be activated for a specific brief: a language, a dialect region, a demographic, a recording condition, a speech style. These are not active contributors to the corpus above; they are the pool it is drawn from and extended through:

Registered contributors 2,000,000+
Countries 180+
Languages reachable 250+

The gap between the two is the point. Several East African languages we hold — Taita, Teso, Kikuyu, Luo, Kamba, Luyia, Kalenjin — have no dedicated speech dataset on this Hub at all.

In active collection: 7,500 hours across Hiligaynon, Tagalog and Cebuano — 2,500 hours each, split 1,000 hours single-speaker and 1,500 hours multi-speaker.

Get in touch

Volume licensing, human-validated transcription over a larger subset, or commissioned collection in a language not listed here — info@silencio.network

Licence on published samples is CC BY-NC 4.0. Non-commercial covers research, evaluation and publication; benchmarking a commercial product model is a commercial use and needs a licence, which is usually granted for evaluation. Ask.