Labari Voice
AI & ML interests
None defined yet.
Recent Activity
Labari Voice
Voice infrastructure ยท low-resource languages
The voice data your models have never heard.
Eight languages. Around 200 million speakers. Native-speaker recordings, transcription, 14 annotation layers, commercial rights included โ produced in studio with native speakers under contract, delivered via API.
labari.dev ยท hello@labari.dev
Why this data doesn't exist
Speech models were trained on the languages that were easiest to collect. Hundreds of millions of speakers fall outside that set.
This is not a demand problem. It is a production problem โ and it is why the gap has stayed open.
You cannot scrape this data. It has to be recorded: in country, in studio, with native speakers under contract, on terms that clear commercial rights at the source. Then transcribed, in languages whose written standards are still consolidating and whose competent transcribers are few.
That barrier is the reason the data isn't available. It is also what we built.
Languages
| Language | Primary countries |
|---|---|
| Fon | Benin |
| Ewe | Togo, Ghana |
| Yoruba | Nigeria, Benin |
| Hausa | Nigeria, Niger |
| Wolof | Senegal, Gambia |
| Pulaar | Senegal, Mauritania, Guinea |
| Bambara | Mali |
| Maninka | Guinea, Mali |
Need a language that isn't listed? Our sourcing network extends beyond the published catalog.
What a Labari dataset contains
- Native-speaker recordings โ studio-recorded, speakers under contract, documented recording chain.
- Transcription โ with orthographic validation for languages whose written standards are still consolidating.
- 14 annotation layers โ phonetics, semantics, timing and context.
- Quality reported, not claimed โ gold-standard multi-annotator protocol, with Krippendorff ฮฑ, WER and CER in the datasheet.
- Commercial rights secured at the source, GDPR compliant.
- API delivery.
How we produce
End-to-end, in-house: speaker sourcing, studio recording, segmentation, transcription, orthographic validation and QA.
Every corpus carries its own audit trail โ segmentation thresholds, per-segment quality measurements, SHA-256 checksums for sources and segments, and the full production parameters. Where a recording chain applies processing that affects a measurement, we say so and flag what it affects. We would rather publish a caveat than let a buyer find it after delivery.
That traceability is the product as much as the audio: it is what lets you filter a corpus on quality before training, and what lets you defend its provenance afterwards.
Licensing
- Catalog subscription โ access to published corpora.
- Custom datasets โ specify language, register, genre, speaker profile, volume and recording conditions; we produce to that brief.
- Exclusivity windows โ negotiated periods of exclusive commercial use.
Public samples
The corpora published on this page are pilot releases: small, unlabeled, and intended for method review rather than training. They exist so the segmentation, measurement and documentation behind the catalog can be audited before you evaluate the annotated product.
The commercial catalog โ transcribed, annotated, licensed โ is not distributed here.