| # s2s-a0-en — Stage-A0 training grids for a small duplex S2S model | |
| Format per `tokens/a0_*.npz`: | |
| - `audio` int16 [T,16] — Mimi codes @12.5Hz, 8 codebooks: cols 0-7 = speech | |
| ("agent" channel), cols 8-15 = encoded silence ("user" channel) | |
| - `text` int32 [T] — Qwen3-0.6B token ids, run-length from each utterance's | |
| start frame; pad elsewhere | |
| - `prompt_len` = 0, `source` = "a0_<dataset>" | |
| Utterances packed to ~100s per grid with 0.2-0.4s gaps. No delay/lead applied. | |
| `manifests/*.json` map each source parquet file to its grids + hours. | |
| See ATTRIBUTION.md for source datasets and licenses. | |
Xet Storage Details
- Size:
- 599 Bytes
- Xet hash:
- 06fb6f8c06cda4d087e28b395a3483cfc94b8808dfb7a191f4b74bf957881dae
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.