Buckets:
230 GB
1,328 files
Updated about 2 months ago
Ctrl+K
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| data | 1,326 items | ||
| .gitattributes | 2.5 kB xet | 738f1125 | |
| README.md | 1.55 kB xet | 182ae441 |
TTS-German
High-quality German speech dataset for TTS and ASR, derived from CML-TTS German.
Processing Pipeline
- Standardize → 24kHz mono WAV, loudness normalize
- Transcribe → WhisperX word-level timestamps
- Segment → ≤12s at word boundaries
- Denoise → DeepFilterNet
- Quality filter → DNSMOS ≥ 2.5
- G2P → IPA phonemes (custom dictionary)
Statistics
| Metric | Value |
|---|---|
| Samples | 670,509 |
| Hours | 1250h |
| Sample rate | 24kHz mono |
| Max duration | 12s |
Schema
| Column | Type | Description |
|---|---|---|
__key__ |
string | Unique ID |
audio |
Audio (24kHz FLAC) | Lossless audio |
text |
string | Transcript |
ipa |
string | IPA phonemes |
language |
string | Language code |
speaker_id |
string | Speaker identifier |
gender |
string | male / female / unknown |
dnsmos |
float | Quality score (1–5) |
Usage
from datasets import load_dataset
ds = load_dataset("datadriven-company/TTS-German", split="train")
sample = ds[0]
print(sample["text"]) # transcript
print(sample["ipa"]) # IPA phonemes
# sample["audio"] → {"array": np.ndarray, "sampling_rate": 24000}
License
cc-by-4.0 — derived from CML-TTS German.
- Total size
- 230 GB
- Files
- 1,328
- Last updated
- Aug 6
- Pre-warmed CDN
- US EU US EU