Speech Sources — Unchunked Collection Long-form or minimally processed source speech datasets. • 11 items • Updated 4 days ago
TTS — Current Training Targets Collection Latest prepared _tts_train corpus for each source family. Use for standard TTS training. • 15 items • Updated 4 days ago
Model-specific Encodings and Artifacts Collection Tokenizer-, codec-, or model-specific derived representations. • 18 items • Updated 4 days ago
STT — Current Training Targets Collection Latest published speech-restored leaf for each of the 14 accepted STT source families. • 14 items • Updated 4 days ago
STT — Validated Upstream Inputs Collection The 14 upstream repositories used by the completed STT training run. These are not final immutable training releases. • 14 items • Updated 4 days ago
Speech Sources — Segmented and Chunked Collection Reusable segmented or chunked speech datasets. • 10 items • Updated 4 days ago
TTS — Raw Pair Plans (Not Train-Ready) Collection Raw clone-pair planning metadata. These repositories are inputs to consensus and packaging, not training releases. • 11 items • Updated 4 days ago
TTS — Voice-Clone Training Targets Collection Packaged non-raw reference/target pair shards for public voice-clone training. • 10 items • Updated 4 days ago
STT — Aligned Upstream Candidates Collection Restored or aligned STT repositories that are not in the validated 14-repository upstream set. • 9 items • Updated 4 days ago