README / README.md
labari-dev's picture
Terminology: sector-standard wording
8946330 verified
|
Raw
History Blame Contribute Delete
3.51 kB
---
title: README
emoji: πŸŽ™οΈ
colorFrom: green
colorTo: gray
sdk: static
pinned: false
---
# Labari Voice
**Voice infrastructure Β· low-resource languages**
*The voice data your models have never heard.*
Eight languages. Around 200 million speakers. Native-speaker recordings, transcription,
14 annotation layers, commercial rights included β€” produced in studio with
native speakers under contract, delivered via API.
[labari.dev](https://labari.dev) Β· `hello@labari.dev`
## Why this data doesn't exist
Speech models were trained on the languages that were easiest to collect.
Hundreds of millions of speakers fall outside that set.
This is not a demand problem. It is a production problem β€” and it is why the
gap has stayed open.
You cannot scrape this data. It has to be recorded: in country, in studio, with
native speakers under contract, on terms that clear commercial rights at the
source. Then transcribed, in languages whose written standards are still
consolidating and whose competent transcribers are few.
That barrier is the reason the data isn't available. It is also what we built.
## Languages
| Language | Primary countries |
|---|---|
| Fon | Benin |
| Ewe | Togo, Ghana |
| Yoruba | Nigeria, Benin |
| Hausa | Nigeria, Niger |
| Wolof | Senegal, Gambia |
| Pulaar | Senegal, Mauritania, Guinea |
| Bambara | Mali |
| Maninka | Guinea, Mali |
Need a language that isn't listed? Our sourcing network extends beyond the
published catalog.
## What a Labari dataset contains
- **Native-speaker recordings** β€” studio-recorded, speakers under contract, documented
recording chain.
- **Transcription** β€” with orthographic validation for languages whose written
standards are still consolidating.
- **14 annotation layers** β€” phonetics, semantics, timing and context.
- **Quality reported, not claimed** β€” gold-standard multi-annotator protocol,
with Krippendorff Ξ±, WER and CER in the datasheet.
- **Commercial rights secured at the source**, GDPR compliant.
- **API delivery.**
## How we produce
End-to-end, in-house: speaker sourcing, studio recording, segmentation,
transcription, orthographic validation and QA.
Every corpus carries its own audit trail β€” segmentation thresholds, per-segment
quality measurements, SHA-256 checksums for sources and segments, and the full
production parameters. Where a recording chain applies processing that affects a
measurement, we say so and flag what it affects. We would rather publish a
caveat than let a buyer find it after delivery.
That traceability is the product as much as the audio: it is what lets you
filter a corpus on quality before training, and what lets you defend its
provenance afterwards.
## Licensing
- **Catalog subscription** β€” access to published corpora.
- **Custom datasets** β€” specify language, register, genre, speaker profile,
volume and recording conditions; we produce to that brief.
- **Exclusivity windows** β€” negotiated periods of exclusive commercial use.
[Describe your needs β†’](https://labari.dev)
## Public samples
The corpora published on this page are pilot releases: small, unlabeled, and
intended for method review rather than training. They exist so the segmentation,
measurement and documentation behind the catalog can be audited before you
evaluate the annotated product.
The commercial catalog β€” transcribed, annotated, licensed β€” is not distributed
here.
[Browse the catalog β†’](https://labari.dev) Β· [Request a sample β†’](https://labari.dev)