my-cool-model / README.md
TLMsp's picture
Upload folder using huggingface_hub
38ad221 verified
|
Raw
History Blame Contribute Delete
5.83 kB
---
configs:
- config_name: agriculture
data_files:
- split: train
path: agriculture_train_tarred/sharded_manifests_with_image/manifest_*.json
- split: validation
path: agriculture_val_tarred/sharded_manifests_with_image/manifest_*.json
- split: test
path: agriculture_test_tarred/sharded_manifests_with_image/manifest_*.json
- config_name: education
data_files:
- split: train
path: education_train_tarred/sharded_manifests_with_image/manifest_*.json
- split: validation
path: education_val_tarred/sharded_manifests_with_image/manifest_*.json
- split: test
path: education_test_tarred/sharded_manifests_with_image/manifest_*.json
- config_name: financial
data_files:
- split: train
path: financial_train_tarred/sharded_manifests_with_image/manifest_*.json
- split: validation
path: financial_val_tarred/sharded_manifests_with_image/manifest_*.json
- split: test
path: financial_test_tarred/sharded_manifests_with_image/manifest_*.json
- config_name: government
data_files:
- split: train
path: goverment_train_tarred/sharded_manifests_with_image/manifest_*.json
- split: validation
path: goverment_val_tarred/sharded_manifests_with_image/manifest_*.json
- split: test
path: goverment_test_tarred/sharded_manifests_with_image/manifest_*.json
- config_name: health
data_files:
- split: train
path: health_train_tarred/sharded_manifests_with_image/manifest_*.json
- split: validation
path: health_val_tarred/sharded_manifests_with_image/manifest_*.json
- split: test
path: health_test_tarred/sharded_manifests_with_image/manifest_*.json
default: agriculture
license: cc-by-4.0
task_categories:
- automatic-speech-recognition
language:
- rw
tags:
- ASR
- STT
- RW
- KIN
pretty_name: Afrivoice Kinyarwanda ASR dataset
size_categories:
- 100K<n<1M
annotations_creators:
- crowdsourced
language_creators:
- crowdsourced
---
# Dataset Card for the image text and voice dataset
## Dataset Description
Each datapoint in this dataset consists of a JPEG image, a corresponding audio Webm file describing the image, and when available, the transcription of the audio file.
| Domain | Total Hours | Transcribed Hours | Number of Clips | Dataset Size (GB) |
|-------------|-------------|-------------------|-----------------|-------------------|
| Agriculture | 467.13 | 465.40 | 86,305 | 30.13 |
| Health | 994.32 | 992.87 | 179,219 | 58.53 |
| Finance | 564.21 | 563.11 | 103,159 | 38.55 |
| Government | 676.10 | 674.11 | 122,265 | 49.22 |
| Education | 198.49 | 197.52 | 37,377 | 13.70 |
| Scripted Education | 262.63 | 262.63 | 42,062 | 15.08 |
| **Total** | **3162.88** | **3155.64** | **570,387** | **205.21** |
## How to use
The `datasets` library allows you to load and pre-process your dataset in pure Python, at scale. The dataset can be downloaded and prepared in one call to your local drive in two ways:
Direct download of the whole dataset to a local directory
```
from huggingface_hub import snapshot_download
snapshot_download(repo_id="DigitalUmuganda/Afrivoice_Kinyarwanda",repo_type='dataset',local_dir='<destination_dir>')
```
Or, using load_dataset to download a particular domain
```
from datasets import load_dataset
data = load_dataset("DigitalUmuganda/Afrivoice_Kinyarwanda",name='health')
```
## Dataset Structure
### Data Instances
```python
{'creator_id': 'bcSXMYbErjM6pJyAwwLs7NAxA9v2',
'raw_text': 'Ingagi ihagaze yonyine. Ingagi ni nziza cyane, kuko zikurura ba mukerarugendo bakazana amadovize mu Gihugu cyacu.',
'duration': 15.06,
'LUFS': -25.5,
'image_category': 'Agriculture',
'image_sub_category': 'Wild Animals',
'text': 'ingagi ihagaze yonyine ingagi ni nziza cyane kuko zikurura ba mukerarugendo bakazana amadovize mu gihugu cyacu',
'audio_filepath': 'audio_1751479904-bcSXMYbErjM6pJyAwwLs7NAxA9v2.webm',
'image_filepath': 'restyf.jpg',
'age_group': '50+',
'gender': 'Male',
'location': 'Musanze',
'shard_id': 0,
'image_shard_id': 0}
```
### Data Fields
`creator` (`string`): An id for which client (voice) made the recording
`raw_text` (`string`): Original audio transcription with punctuation and capitalization
`image_filepath` (`string`): name of the image file inside the shard
`audio_filepath` (`string`): name of the audio file inside the shard
`text` (`string`): normalized audio transcription (i.e: without punctuation and capitalization)
`age_group` (`string`): age range of the audio recorder
`gender` (`string`): The gender of the speaker
`location` (`string`): geographical location of the audio recorder
`duration` (`int`): length in seconds of the audio file
`image_category` (`string`): domain of the image (eg: health, agriculture, finance), used as prompt during audio creation.
`image_sub_category` (`string`): Sub-domain label of the image (e.g., within agriculture: “seed farming” or “forestry”), used to guide audio creation.
`shard_id` (`int`): index of the shard containing the audio file in the `audio_filepath` column.
`image_shard_id` (`int`): index of the shard containing the image in the `image_filepath` column.
`LUFS` (`int`): Loudness of the audio in Loudness Units relative to Full Scale (LUFS). Lower values indicate quieter audio, while higher values indicate louder audio.
### Data Splits
Each domain in the dataset is divided into train, validation, and test splits.
# Licensing Information
All datasets are licensed under the [Creative Commons license (CC-BY-4)](https://creativecommons.org/licenses/).