edwixx/s2s-a0-en / README.md
edwixx's picture
|
download
raw
599 Bytes
# s2s-a0-en — Stage-A0 training grids for a small duplex S2S model
Format per `tokens/a0_*.npz`:
- `audio` int16 [T,16] — Mimi codes @12.5Hz, 8 codebooks: cols 0-7 = speech
("agent" channel), cols 8-15 = encoded silence ("user" channel)
- `text` int32 [T] — Qwen3-0.6B token ids, run-length from each utterance's
start frame; pad elsewhere
- `prompt_len` = 0, `source` = "a0_<dataset>"
Utterances packed to ~100s per grid with 0.2-0.4s gaps. No delay/lead applied.
`manifests/*.json` map each source parquet file to its grids + hours.
See ATTRIBUTION.md for source datasets and licenses.

Xet Storage Details

Size:
599 Bytes
·
Xet hash:
06fb6f8c06cda4d087e28b395a3483cfc94b8808dfb7a191f4b74bf957881dae

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.