230 GB
1,328 files
Updated about 2 months ago
Name
Size
data
.gitattributes2.5 kB
xet
README.md1.55 kB
xet
README.md

TTS-German

High-quality German speech dataset for TTS and ASR, derived from CML-TTS German.

Processing Pipeline

  1. Standardize → 24kHz mono WAV, loudness normalize
  2. Transcribe → WhisperX word-level timestamps
  3. Segment → ≤12s at word boundaries
  4. Denoise → DeepFilterNet
  5. Quality filter → DNSMOS ≥ 2.5
  6. G2P → IPA phonemes (custom dictionary)

Statistics

Metric Value
Samples 670,509
Hours 1250h
Sample rate 24kHz mono
Max duration 12s

Schema

Column Type Description
__key__ string Unique ID
audio Audio (24kHz FLAC) Lossless audio
text string Transcript
ipa string IPA phonemes
language string Language code
speaker_id string Speaker identifier
gender string male / female / unknown
dnsmos float Quality score (1–5)

Usage

from datasets import load_dataset
ds = load_dataset("datadriven-company/TTS-German", split="train")
sample = ds[0]
print(sample["text"])    # transcript
print(sample["ipa"])     # IPA phonemes
# sample["audio"] → {"array": np.ndarray, "sampling_rate": 24000}

License

cc-by-4.0 — derived from CML-TTS German.

Total size
230 GB
Files
1,328
Last updated
Aug 6
Pre-warmed CDN
US EU US EU

Contributors