File size: 5,456 Bytes
43fd74b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3d38133
 
43fd74b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
---
license: cc-by-4.0
library_name: flashvad
language:
  - ar
  - en
  - gu
  - hi
  - kn
  - pa
  - ta
  - te
  - ur
tags:
  - audio
  - voice-activity-detection
  - vad
  - streaming
  - onnx
  - pytorch
  - telephony
  - webrtc
---

# FlashVAD v0.1 model card

## Summary

FlashVAD v0.1 is a 46,170-parameter causal streaming voice-activity detector.
It consumes 16 kHz mono audio, produces one speech probability every 10 ms,
and keeps independent convolutional, recurrent, feature, and detector state
per call.

**Author:** Himanshu Maurya

This checkpoint is an **alpha research preview** for integration testing,
browser demonstrations, and shadow evaluation. It is not approved for
production use and is not validated as a general multilingual or India/GCC
call model.

## Artifacts

| Artifact | SHA-256 |
|---|---|
| `flashvad-v0.1.pt` | `ca9e35475518466b2a1f2e89b4953cd1e26e3d8c513cdcf265ab319e74e2b288` |
| `flashvad-stream.onnx` | `9a88e34bf3118d60e25a16cb622cb394e2f3ab71445b0aa5957df6f1d5f1b6ba` |
| `config.json` | `0b1ad372808f7c67cea5a1ca4b41a817714c9a0b0cf49b3baa56fe8d5f64ad2b` |
| `detector-calibration.json` | `b5d000e0406d81fbd87a9e66194a877fa8433a76783faac8275c7969c43051b4` |

The ONNX graph accepts precomputed 43-dimensional causal features. Use
`src/flashvad/features.py` or
`report-site/src/lib/vad-features.mjs`; it is not a raw-waveform graph.

The public ONNX file is self-contained and stripped of exporter stack traces,
local paths, and private build metadata.

## Download and source

Download the complete model repository:

```bash
hf download oss-codes/flashvad --local-dir flashvad-model
```

Source code and runtime integrations are published separately at
[`oss-codes/flashvad`](https://github.com/oss-codes/flashvad).

The ONNX graph does not accept raw waveform audio. It expects the causal
43-dimensional features described below, with independent feature and model
state for every call.

## Intended use

Appropriate current uses:

- research and architecture evaluation;
- functional integration with browser, LiveKit, Pipecat, SIP, or PSTN stacks;
- latency and concurrency measurement;
- shadow-mode comparison on consented, labelled call audio.

Do not use this checkpoint as the sole basis for:

- emergency, medical, legal, financial, or safety-critical decisions;
- call recording consent or compliance decisions;
- a production multilingual accuracy claim;
- semantic end-of-turn detection.

Reset all feature, model, resampler, and detector state when a call ends or an
audio discontinuity occurs.

## Architecture

- 25 ms causal analysis frame and 10 ms hop;
- 40 log-mel bands plus energy, zero-crossing rate, and spectral flatness;
- four causal depthwise temporal blocks with dilations 1, 2, 4, and 8;
- one 64-unit GRU;
- speech and auxiliary event heads;
- 184,680 bytes of FP32 parameters.

## Training inputs and provenance limit

Historical training notes associated with the retained checkpoint report:

- 558 derived clips from nine FLEURS configurations: Arabic, English,
  Gujarati, Hindi, Kannada, Punjabi, Tamil, Telugu, and Urdu;
- 288 AMI meeting clips with meeting-family-disjoint train/validation splits;
- 64 MUSAN noise clips;
- weak frame targets from the official Silero VAD model.

FLEURS, AMI, and MUSAN attribution is in `NOTICE`; dataset audio is not
distributed here.

Those counts and corpus names are not embedded as complete provenance in the
public checkpoint. The exact retained training manifests, their digests, source
revisions, and teacher-output digest were not preserved, so the checkpoint's
training run is **not bit-for-bit reproducible** from the public tree. This is a
provenance limitation, not evidence of broader accuracy. Future release
candidates must preserve those records before training begins.

## Evaluation

The retained checkpoint was repeatedly inspected on TEN VAD's public
30-recording set while research candidates were compared. At the retained
detector policy, the descriptive results are:

- 26,243 frames;
- ROC-AUC: 0.882;
- raw F1 at the configured threshold: 0.886;
- hysteresis-decision F1: 0.889;
- false-alarm rate: 26.3%;
- miss rate: 13.0%.

Because the public set influenced research decisions, this is an exploratory
external-set result—not an untouched test, independent benchmark, or
production generalization estimate. Language is recorded as `und`, and codec,
channel, device, and SNR are unknown.

The machine-readable report is
`benchmarks/flashvad-v0.1/ten-public-evaluation.json`. Its false-alarm rate is
too high for production.

## Known limitations

- Quiet speech, music, TTS leakage, echo, television, laughter, singing, and
  overlapping speakers may cause misses or false triggers.
- Read speech and meetings do not cover real carrier, device, packet-loss, and
  room conditions.
- Language presence in training does not prove per-language performance.
- Linear 8-to-16 kHz conversion prioritizes causal speed, not audio fidelity.
- VAD cannot determine whether a speaker has semantically completed a turn.

A production candidate needs consented, human-labelled, speaker-disjoint calls
with predeclared per-slice gates and a test set untouched until final
evaluation.

## Licences

Repository source code is MIT-licensed. The retained model artifacts are
separately available under CC BY 4.0; see
`MODEL_LICENSE.md` and `NOTICE`. Third-party datasets, models, and benchmark
materials retain their own terms.