Buckets:
dataset_info:
features:
- name: audio
dtype: audio
- name: text
dtype: string
- name: source
dtype: string
- name: sample_rate
dtype: int64
- name: speaker
dtype: string
splits:
- name: train
num_bytes: 2942869578
num_examples: 13000
download_size: 3303296950
dataset_size: 2942869578
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
license: cc
task_categories:
- text-to-speech
- automatic-speech-recognition
language:
- tr
size_categories:
- 10K<n<100K
Synthetic Turkish TTS Data
This dataset was created by generating synthetic Turkish text across multiple speech scenarios. The text was produced in the following domains: finance_master, cs_master, parcel_delivery, ecommerce, telecom, isp_support, technical_support, subscription, insurance, health_appointments, public_services, education_registration, and daily_speech.
These synthetic texts were then synthesized with a high-quality Turkish TTS model. The dataset is intended to be used as synthetic training data for future Turkish TTS model training.
Dataset Summary
- Total number of samples: 13,000
- Total number of speakers: 4
- Total duration: 28 hours 50 minutes 13 seconds
Speaker Distribution
- ali: 6 hours 50 minutes 57 seconds (3250 wav files)
- zeynep: 6 hours 50 minutes 35 seconds (3250 wav files)
- leyla: 7 hours 10 minutes 55 seconds (3250 wav files)
- alev: 7 hours 57 minutes 46 seconds (3250 wav files)
Columns
audio: audio filetext: corresponding transcriptionsource: scenario/domain from which the sample was generatedsample_rate: audio sampling ratespeaker: speaker name
License
This dataset is released under the CC BY 4.0 license. You may use, share, and adapt the dataset provided that proper attribution is given.
Audio Generation
The audio samples in this dataset were generated using the Freya AI API.
Reference: https://www.linkedin.com/company/107923818/
Xet Storage Details
- Size:
- 2.05 kB
- Xet hash:
- 88c2e2076414ae866801571a63427b1f83181a5c9f36a8c41f3a075b71082457
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.