Spaces:
Running
Running
| title: README | |
| emoji: ποΈ | |
| colorFrom: green | |
| colorTo: gray | |
| sdk: static | |
| pinned: false | |
| # Labari Voice | |
| **Voice infrastructure Β· low-resource languages** | |
| *The voice data your models have never heard.* | |
| Eight languages. Around 200 million speakers. Native-speaker recordings, transcription, | |
| 14 annotation layers, commercial rights included β produced in studio with | |
| native speakers under contract, delivered via API. | |
| [labari.dev](https://labari.dev) Β· `hello@labari.dev` | |
| ## Why this data doesn't exist | |
| Speech models were trained on the languages that were easiest to collect. | |
| Hundreds of millions of speakers fall outside that set. | |
| This is not a demand problem. It is a production problem β and it is why the | |
| gap has stayed open. | |
| You cannot scrape this data. It has to be recorded: in country, in studio, with | |
| native speakers under contract, on terms that clear commercial rights at the | |
| source. Then transcribed, in languages whose written standards are still | |
| consolidating and whose competent transcribers are few. | |
| That barrier is the reason the data isn't available. It is also what we built. | |
| ## Languages | |
| | Language | Primary countries | | |
| |---|---| | |
| | Fon | Benin | | |
| | Ewe | Togo, Ghana | | |
| | Yoruba | Nigeria, Benin | | |
| | Hausa | Nigeria, Niger | | |
| | Wolof | Senegal, Gambia | | |
| | Pulaar | Senegal, Mauritania, Guinea | | |
| | Bambara | Mali | | |
| | Maninka | Guinea, Mali | | |
| Need a language that isn't listed? Our sourcing network extends beyond the | |
| published catalog. | |
| ## What a Labari dataset contains | |
| - **Native-speaker recordings** β studio-recorded, speakers under contract, documented | |
| recording chain. | |
| - **Transcription** β with orthographic validation for languages whose written | |
| standards are still consolidating. | |
| - **14 annotation layers** β phonetics, semantics, timing and context. | |
| - **Quality reported, not claimed** β gold-standard multi-annotator protocol, | |
| with Krippendorff Ξ±, WER and CER in the datasheet. | |
| - **Commercial rights secured at the source**, GDPR compliant. | |
| - **API delivery.** | |
| ## How we produce | |
| End-to-end, in-house: speaker sourcing, studio recording, segmentation, | |
| transcription, orthographic validation and QA. | |
| Every corpus carries its own audit trail β segmentation thresholds, per-segment | |
| quality measurements, SHA-256 checksums for sources and segments, and the full | |
| production parameters. Where a recording chain applies processing that affects a | |
| measurement, we say so and flag what it affects. We would rather publish a | |
| caveat than let a buyer find it after delivery. | |
| That traceability is the product as much as the audio: it is what lets you | |
| filter a corpus on quality before training, and what lets you defend its | |
| provenance afterwards. | |
| ## Licensing | |
| - **Catalog subscription** β access to published corpora. | |
| - **Custom datasets** β specify language, register, genre, speaker profile, | |
| volume and recording conditions; we produce to that brief. | |
| - **Exclusivity windows** β negotiated periods of exclusive commercial use. | |
| [Describe your needs β](https://labari.dev) | |
| ## Public samples | |
| The corpora published on this page are pilot releases: small, unlabeled, and | |
| intended for method review rather than training. They exist so the segmentation, | |
| measurement and documentation behind the catalog can be audited before you | |
| evaluate the annotated product. | |
| The commercial catalog β transcribed, annotated, licensed β is not distributed | |
| here. | |
| [Browse the catalog β](https://labari.dev) Β· [Request a sample β](https://labari.dev) | |