File size: 1,469 Bytes
91654d7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
---
license: mit
tags:
- webgpu
- wasm
- in-browser
- asr
- tts
- vad
- speech
- fluidaudio
---

# FluidAudio Web — raw model weights

Model weights for **[fluidaudio-web](https://github.com/FluidInference/fluidaudio-web)**:
fully in-browser inference (WebGPU + WebAssembly) for FluidAudio's speech models.

These are **raw weights only — no ONNX, no onnxruntime.** Each engine is a
hand-written forward pass (WebGPU/WASM/JS) that loads these tensors directly and,
where applicable, dequantizes them in-shader. Weights are extracted from the
source models and parity-verified against the originals before publishing.

## Layout (one folder per engine)

| folder | model | status |
|---|---|---|
| `vad/` | Silero VAD v5 (16 kHz) | ✅ available (fp32) |
| `parakeet/` | Parakeet TDT 0.6B v3 (encoder + decoder/joint) | 🚧 in progress (int4/palettized) |
| `kokoro/` | Kokoro TTS 82M | 🚧 planned |
| `nemotron/` · `eou/` · `sortformer/` · `whisper/` | streaming ASR / EOU / diarization / multilingual ASR | 🚧 planned |

## Format

- `*.bin` — concatenated little-endian tensors.
- `manifest.json``name -> { dims, offset, len }` (offset/len in **elements**, not bytes).

Load: fetch the `.bin`, slice each tensor per the manifest. Extraction scripts and
the forward passes live in the fluidaudio-web repo (`scripts/`, `src/engines/`).

Quantized engines additionally carry per-tensor scales / palettes (documented per
folder) for in-shader dequant.