GGUF
vokra
wavtokenizer-large / README.md
ayousanz's picture
Upload folder using huggingface_hub
fa3dc6c verified
|
Raw
History Blame Contribute Delete
1.78 kB
---
license: mit
library_name: vokra
tags:
- vokra
- gguf
---
# wavtokenizer-large (Vokra GGUF)
Converted to the Vokra GGUF format for [Vokra](https://github.com/ayutaz/vokra), a zero-dependency speech-AI inference runtime.
**This is a conversion, not a new model.** The weights are the upstream ones; Vokra re-packages them so its runtime can memory-map them directly. Credit for the model belongs upstream — see *Source* below.
## Files
| File | Size | SHA-256 |
|---|---|---|
| `wavtokenizer-large.gguf` | 807.2 MB | `99b7dce0426266f7f2f6615091d832cea71387ce57edfae66666143a5c33a36b` |
## Usage
```bash
# Download (any HTTP client works — the file is a plain GGUF)
curl -L -o wavtokenizer-large.gguf \
https://huggingface.co/vokra/wavtokenizer-large/resolve/main/wavtokenizer-large.gguf
```
```bash
vokra-cli run --model wavtokenizer-large.gguf --input input.wav
```
## Provenance
| Field | Value |
|---|---|
| Architecture | `wavtokenizer` |
| Tensors | 1091 |
| Upstream source | novateur/WavTokenizer-large-speech-75token (single-codebook FSQ audio codec, 24 kHz, hop 320 → 75 tok/s, arXiv:2408.16532, MIT) |
| Upstream licence | `mit` |
| Licence class | `permissive` |
| Registry model id | `wavtokenizer-large-speech-75token` |
| Vokra GGUF schema | 1 |
| Converted by | vokra-core 0.1.0-alpha.0 |
Every row above is read out of this file's own `vokra.*` metadata, so the card cannot claim something the artifact does not carry.
## Licence
The weights are distributed under **mit**, unchanged from upstream. Conversion does not alter the licence, and your obligations run to the upstream author.
## Verifying this file
```bash
shasum -a 256 wavtokenizer-large.gguf
# expect: 99b7dce0426266f7f2f6615091d832cea71387ce57edfae66666143a5c33a36b
```