Buckets:
81.4 GB
1,812 files
Updated about 2 months ago
Ctrl+K
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| cond_step.mlmodelc | 4 items | ||
| cond_step.mlpackage | 3 items | ||
| cond_step_v2.mlmodelc | 5 items | ||
| cond_step_v3.mlmodelc | 5 items | ||
| cond_step_v3.mlpackage | 3 items | ||
| constants_bin | 67 items | ||
| flow_decoder.mlmodelc | 4 items | ||
| flow_decoder.mlpackage | 3 items | ||
| flowlm_step.mlmodelc | 4 items | ||
| flowlm_step.mlpackage | 3 items | ||
| mimi_decoder.mlmodelc | 4 items | ||
| mimi_decoder.mlpackage | 3 items | ||
| mimi_decoder_v2.mlmodelc | 10 items | ||
| mimi_decoder_v3.mlmodelc | 5 items | ||
| mimi_decoder_v3.mlpackage | 3 items | ||
| mimi_encoder.mlmodelc | 5 items | ||
| mimi_encoder.mlpackage | 3 items | ||
| mimi_encoderv2.mlmodelc | 5 items | ||
| mimi_encoderv2.mlpackage | 3 items | ||
| samples | 3 items | ||
| v2 | 776 items | ||
| v2.1 | 887 items | ||
| .gitattributes | 2.84 kB xet | dda559b4 | |
| README.md | 1.42 kB xet | 3b40f813 | |
| config.json | 2 Bytes xet | 56fc0e24 | |
| manifest.json | 44.5 kB xet | 62e3523e |
PocketTTS CoreML
CoreML conversion of kyutai/pocket-tts for on-device inference on Apple platforms.
Models
| Model | Description | Size |
|---|---|---|
| cond_step | KV cache prefill (voice + text conditioning) | ~200MB |
| flowlm_step | Autoregressive generation (transformer_out + EOS) | ~200MB |
| flow_decoder | Flow matching denoiser (8 Euler steps per frame) | ~190MB |
| mimi_decoder | Streaming audio codec (1920 samples per frame) | ~11MB |
Voices
4 pre-encoded voices in constants_bin/:
alba(default),azelma,cosette,javert
Voice cloning weights are not included — they are gated separately by Kyutai.
Usage
import FluidAudioTTS
let manager = PocketTtsManager()
try await manager.initialize()
let audio = try await manager.synthesize(text: "Hello, world!")
See https://github.com/FluidInference/FluidAudio for the full Swift framework.
License
CC-BY-4.0, inherited from https://huggingface.co/kyutai/pocket-tts. Attribution to Kyutai is required.
References
- https://huggingface.co/kyutai/pocket-tts
- https://arxiv.org/abs/2410.00037
- https://github.com/FluidInference/FluidAudio
- Total size
- 81.4 GB
- Files
- 1,812
- Last updated
- Jun 26
- Pre-warmed CDN
- US EU US EU