File size: 32,272 Bytes
35d483e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0348402
d782513
35d483e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
# Technical report: Tiny Hinglish Turn Detector

**Challenge:** decide, from audio, whether a user is finished speaking or only
pausing.

**Author:** Suvradeep Das

**Repository state documented:** 2026-08-23

**Current verdict:** a fast, fully exported neural preview exists, but the
interpretable acoustic baseline is still stronger. Generalization remains
unestablished.

## Executive summary

I treated turn detection as a product decision under asymmetric error costs,
not as a generic balanced classification benchmark. A premature `END` causes a
visible interruption; a late `END` costs latency. The system therefore separates
three concerns:

1. an audio classifier estimates `p(END)` at a VAD pause checkpoint;
2. a threshold is selected to maximize endpoint recall under a predeclared
   false-interruption budget;
3. a deterministic controller applies minimum silence, long-pause relaxation,
   optional confirmation, and a hard maximum timeout.

The packaged preview is a 151,812-parameter, four-second log-mel TinyTCN trained
from scratch and then continued with a fresh optimizer. It has six causal
depthwise-separable residual blocks and mask-aware attentive statistics pooling,
with auxiliary heads for mid-turn and end-turn fillers. The FP32 ONNX file is
625,431 bytes and measures 1.629 ms p95 from a four-second waveform to
probability on the recorded Apple-arm/one-thread setup. A larger 396,164-
parameter configuration and a Whisper-tiny teacher remain proposed full-data
experiments, not packaged results.

The repository covers the full engineering path: pinned/resumable data download,
streaming audit, transitive leakage grouping, deterministic split generation,
lazy audio loading, training, calibration, sliced and operational metrics, ONNX
export/parity, quantization hooks, CPU benchmarking, runtime replay, Gradio, and
guarded Hugging Face packaging/publishing.

What it does **not** contain is equally important: only one of 83 train shards
is present locally; there is no full-corpus run; the official test has not been
downloaded or evaluated; and no human Hinglish recordings have been collected.
The upstream shard also lacks usable speaker/conversation/voice IDs, making the
IID split best-effort row-disjoint rather than speaker-disjoint.

On the 326-row IID validation split, the acoustic logistic baseline is the
current winner: AP 0.7310, AUROC 0.7522, and 21/160 END recall at 3/166 false
interruptions. The TinyTCN reaches AP 0.6881, AUROC 0.7390, and 10/160 END
recall at the same three false interruptions. Its constrained-recall delta is
−0.0688 with a paired bootstrap 95% interval of [−0.1330, −0.0122]. That is a
useful negative result. Both models and thresholds were selected adaptively on
the same development data, so the result is not an independent estimate.

## 1. Problem definition

At each pause checkpoint, the target is:

- `y = 0` (`HOLD`): responding now would interrupt an incomplete turn;
- `y = 1` (`END`): the turn is complete and the agent may respond.

The model consumes only the recent audio suffix. This is intentionally narrower
than end-to-end dialogue policy: VAD decides when speech has paused, the endpoint
model estimates conversational completeness, and the controller turns that
score into a time-bounded action.

### 1.1 Error economics

Let `FP` be an incomplete turn predicted as `END`, and `FN` a complete turn
predicted as `HOLD`.

- `FP` is a false interruption. It can cut off an address, order ID, correction,
  or filler-delayed continuation.
- `FN` adds waiting time, bounded by the controller's maximum timeout.

Because the first is normally more harmful, the primary operating point is not
the maximum-F1 threshold. It is the threshold with maximum END recall subject to
`FPR <= budget`, where `FPR = FP / (FP + TN)`. The code reports budgets of 1%,
2%, and 5%; the primary training configuration selects at 2%. A product owner
can change that budget without retraining the representation.

### 1.2 Product-level success criteria

The final scorecard should contain:

- false-positive rate and END recall at the frozen threshold;
- false interruptions per turn and per audio hour;
- p50/p90/p95 endpoint delay after a true completion;
- AUROC and average precision for threshold-independent ranking;
- Brier score, log loss, ECE, and reliability bins for calibration;
- language, source, synthetic/human, filler, duration, device/noise, and
  consented demographic slices;
- grouped bootstrap intervals by speaker or conversation;
- neural-only and end-to-end p50/p95/p99 batch-1 CPU latency, model size, load
  time, and peak RSS.

“Tiny + fast + accurate” is not satisfied by a single aggregate F1 value.

## 2. Data work

### 2.1 Upstream snapshot

The train source is
[`pipecat-ai/smart-turn-data-v3.2-train`](https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-train),
pinned to revision
`e564e2ac567f774d1880aa1db6ce97afb8c519b7`. The repository catalog describes
270,946 rows across 83 Parquet shards (about 41.4 GB). The download script is
resumable and verifies the shard count before success.

The official test source is a separate 10-shard snapshot, pinned to
`0500378e8ed6d38e37b016e24d261e8e6c6a6859`. Its downloader requires the exact
acknowledgement `I_HAVE_FROZEN_MODEL_AND_THRESHOLD`. It has not been downloaded
or inspected.

### 2.2 Measured local audit

Only `train-00010-of-00083.parquet` is local. The following numbers therefore
describe **that shard only**, not the full dataset:

| Audit field | Measured value |
|---|---:|
| Rows | 3,265 |
| Valid / invalid | 3,265 / 0 |
| Encoded duration | 24,685.517 s = 6.857 h |
| Duration range | 0.36–29.56 s |
| Audio format | 3,265 FLAC |
| Endpoint `HOLD` / `END` | 1,667 / 1,598 |
| Synthetic / human-tagged | 2,712 / 553 |
| Languages | 23 |
| English / Hindi / Marathi rows | 758 / 158 / 81 |
| Exact-audio duplicates | 0 |
| Duplicate record IDs | 0 |
| Conflicting duplicate labels | 0 |

The inspected shard is roughly 83% synthetic. Its sequence and source mix may
be non-random, so no count should be extrapolated to the other 82 shards.

The nullable filler labels were preserved rather than silently coerced to
`false`: `midfiller` has 1,339 true, 1,284 false, and 642 null values;
`endfiller` has 905 true, 1,718 false, and 642 null values. Auxiliary losses mask
the nulls.

The machine-readable evidence is in
[`reports/partial_shard_audit.json`](reports/partial_shard_audit.json).

### 2.3 Leakage control

Exact duplicates are only the simplest leakage route. Two rows may share a
conversation, speaker, TTS voice, recording, prompt, record ID, or repeated
transcript while having different audio hashes. The audit derives hashed linkage
keys for each available identifier. Repeated normalized text is linked globally
when it is long enough to be meaningful. A union-find pass computes transitive
connected components, so `A shares audio with B` and `B shares speaker with C`
forces all three into one group.

Splitting assigns whole groups with deterministic iterative multilabel
stratification. Validation checks all three independent crossing classes:

- component `group_id`;
- exact `audio_sha256`;
- every underlying metadata linkage key.

The 90/10 development split has 2,939 train and 326 validation rows, including
1,438/1,501 END/HOLD train and 160/166 END/HOLD validation examples. All crossing
maps are empty. See
[`reports/partial_iid_split.json`](reports/partial_iid_split.json).

That zero-crossing result is narrower than speaker-disjointness. The shard does
not expose usable speaker, conversation, or TTS-voice IDs; every one of the
3,265 base components used by the IID split is a singleton. (The source stress
split separately groups by its 12 dataset values.) The IID split is best-effort row-disjoint
and catches exact audio/observed-key leakage, while identity and repeated-voice
leakage remain unknown. Group-bootstrap intervals on this shard therefore reduce
to row-bootstrap intervals and must not be described as speaker-clustered.

### 2.4 Source-held-out stress test

A second split holds entire `dataset` values together: 2,617 train and 648
validation rows, again with zero detected crossings. It is useful as a joint
domain-shift test, but not as a clean estimate of source robustness:

- train is 2,614 synthetic and only 3 human-tagged rows;
- validation is 98 synthetic and 550 human-tagged rows;
- train spans 23 languages; validation contains only English and Spanish.

Any performance delta simultaneously measures source, language, human/synthetic,
and collection-process shift. Reporting it as “generalization to unseen source”
without this caveat would be misleading. Leave-one-source-out folds are a better
next analysis because they expose variability across individual sources, though
the confounding cannot be removed without better data. Evidence:
[`reports/partial_source_holdout_split.json`](reports/partial_source_holdout_split.json).

### 2.5 Why a separate Hinglish set is necessary

Single ISO tags such as `hin` and `eng` do not establish within-utterance
code-switching, Indian accents, natural hesitation, or logistics vocabulary.
The repository includes a purpose-built protocol with 900 balanced assignments:

- 30 speakers × 30 recordings;
- 450 HOLD and 450 END targets;
- 21/4/5 speaker-disjoint train/dev/test speakers, producing 630/120/150 examples;
- minimal pairs covering address, COD, pickup, reschedule, returns, tracking,
  support, phone/order identifiers, delivery instructions, and damaged parcels;
- natural 0.3–2.0 second internal pauses for HOLD recordings;
- independent “would responding now feel like an interruption?” judgments from
  three annotators.

The assignment file is real and deterministic. The recordings are not: **no
participants have been recorded locally**. Collection requires informed consent,
PII controls, retention/deletion rules, and a redistribution decision before any
audio leaves private storage. See
[`docs/data_collection_protocol.md`](docs/data_collection_protocol.md).

## 3. Model design

### 3.1 Canonical frontend

Audio is converted to mono float, resampled to 16 kHz, suffix-cropped, and
padded with an explicit validity mask. The packaged preview uses four seconds;
the proposed full-data configuration uses eight. Both use:

- 400-sample (25 ms) Hann window;
- 160-sample (10 ms) hop;
- 80 area-normalized HTK mel bands from 0–8 kHz;
- base-10 log energy with an eight-log-unit dynamic-range floor;
- Whisper-style `(log_mel + 4) / 4` scaling, without per-utterance
  mean/variance standardization in the primary configuration.

An explicit frame mask prevents padding from affecting convolutional states or
attentive pooling. The training and dependency-light NumPy serving frontends
must pass parity tests before release. The Whisper teacher uses Slaney mel bands
with the Whisper log convention and fixed encoder length; this is serialized
separately rather than silently reusing the student's HTK filterbank.

### 3.2 TinyTCN student

For packaged input `X` of shape `[batch, 80, frames]`:

1. a 1×1 projection maps 80 mel bins to 128 channels;
2. six residual depthwise-separable causal Conv1d blocks use kernel 5 and
   dilations `[1, 2, 4, 8, 16, 32]`;
3. per-frame LayerNorm avoids padding-dependent batch/time statistics;
4. masked learned attention produces weighted mean and standard deviation;
5. a 96-unit embedding feeds endpoint, midfiller, and endfiller logits.

This selected preview has 151,812 parameters. The configurable implementation
also supports the proposed 192-channel/eight-block/eight-second model with
396,164 parameters. That larger design has not been shown to improve current
development metrics and is not the packaged checkpoint.

Attentive pooling still recomputes a suffix-level summary at a checkpoint. A
future streaming version can cache convolutional states and maintain online
pooling statistics, but the current implementation makes no such latency claim.

### 3.3 Objective and optimization

The default loss is:

```text
L = BCE(endpoint) + 0.15 * masked_BCE(midfiller)
                  + 0.15 * masked_BCE(endfiller)
```

Endpoint label smoothing is 0.02 in the primary config. Positive weighting is a
single train-split statistic if needed; it is never recomputed per batch. The
trainer uses AdamW, gradient clipping, optional CUDA mixed precision,
deterministic kernels, configurable metric-based early stopping, and a seed
fixed before model construction. The default full-data run selects constrained
END recall under the 2% FPR budget. The bounded preview continuation
pre-registered average precision for checkpoint selection, while its decision
threshold was still selected under the 2% development FPR rule.

The checkpoint contains model configuration, model and optimizer state,
frontend configuration, maximum suffix length, selected threshold, run name,
framework version, and validation evidence. The resolved config and full epoch
history are written next to it.

### 3.4 Whisper teacher and distillation

The alternative starts with `openai/whisper-tiny`, uses the correct Whisper
log-mel convention and attention mask, pools encoder states with the same
mask-aware head, and can freeze all or selectively unfreeze final encoder
layers. The default experiment is head-only.

A hard/soft distillation objective is implemented, including temperature
scaling and optional embedding alignment. It is not yet wired into a completed
teacher-student run; presenting it as a measured gain would be false. The proper
sequence is teacher validation, cached teacher logits, student distillation, and
paired evaluation at the same false-interruption budget.

## 4. Experiments completed so far

Smoke and resolver runs first established deterministic mechanics. The rows
below are the bounded real-data experiments; all use one audited train shard.

| Run | Split/model | Result | Decision |
|---|---|---|---|
| Acoustic baseline | IID, ten four-second waveform statistics + class-balanced logistic regression | AP 0.7310, AUROC 0.7522, Brier 0.2038; threshold 0.7639 gives 21 TP / 3 FP / 163 TN / 139 FN | Current local winner and development floor |
| Initial TinyTCN | IID, 151,812 parameters, four seconds, three epochs, LR `3e-4` | AP 0.6874, AUROC 0.7291, recall 0.0938 at FPR 0.0181 | Viable ranking; continue under bounded plan |
| High-LR TinyTCN | Same architecture, planned eight epochs, LR `1e-3` | Stopped after epoch 4; best AP 0.6190, AUROC 0.6829, recall 0.0313 at FPR 0.0181 | Stop for poor trajectory |
| Packaged warm continuation | IID, weights-only continuation, fresh optimizer, five epochs at LR `3e-4` | Selected continuation epoch 1: AP 0.6881, AUROC 0.7390, Brier 0.2133; threshold 0.7410 gives 10 TP / 3 FP / 163 TN / 150 FN | Package as development preview, not as winner |
| Source-stress baseline | Source-held-out logistic baseline | AP 0.5368, AUROC 0.5539, recall 0.0279 at FPR 0.0185 | Severe joint-shift collapse |
| Source-stress TinyTCN | Source-held-out TinyTCN, three epochs | AP 0.5500, AUROC 0.5523, recall 0.0402 at FPR 0.0185 | No established advantage; severe joint-shift collapse |

The warm continuation retained its first epoch; later continuation epochs fell
to AP 0.6798, 0.6735, 0.6723, and 0.6726. The improvement over the original
checkpoint is marginal and the development split had already been inspected.
This is adaptive checkpoint selection, not a fresh replication.

### 4.1 Paired IID comparison

Both model thresholds were chosen on the same 326-row validation set. With the
selected thresholds held fixed, a 2,000-sample paired bootstrap gives TinyTCN
minus baseline:

| Metric | Difference | 95% interval |
|---|---:|---:|
| AP | −0.0429 | [−0.1195, 0.0352] |
| AUROC | −0.0132 | [−0.0782, 0.0512] |
| END recall | −0.0688 | [−0.1330, −0.0122] |
| FPR | 0.0000 | [−0.0284, 0.0298] |

The AP/AUROC intervals include zero; the constrained-recall interval favors the
baseline. Because all 326 validation groups are singletons, this “grouped”
bootstrap is numerically row bootstrap, not speaker/conversation-clustered
uncertainty. It also conditions on adaptively selected checkpoints and
thresholds.

The TinyTCN absolute FPR interval is `[0, 0.0407]` and its recall interval is
`[0.0261, 0.1019]`. Observing 3/166 false interruptions is compatible with more
than a 2% population FPR, so the nominal budget is not statistically certified.

### 4.2 Source-held-out stress

The 648-row validation side contains nine source groups. TinyTCN-minus-baseline
paired deltas are AP `+0.0131` (95% CI `[−0.2149, 0.1261]`), AUROC `−0.0016`
(`[−0.2713, 0.1273]`), and recall `+0.0124`
(`[−0.0112, 0.0948]`). The intervals are wide and both AUROCs are near 0.55.
Combined with the split's source/language/human-synthetic confounding, the only
defensible conclusion is that both models are brittle under this joint shift.

### 4.3 Silence perturbation

Appending valid zero-valued silence before suffix cropping changes the selected
model materially at its frozen threshold:

| Appended silence | AP | AUROC | FPR | END recall | Flips vs 0 ms |
|---:|---:|---:|---:|---:|---:|
| 0 ms | 0.6881 | 0.7390 | 0.0181 | 0.0625 | — |
| 200 ms | 0.6840 | 0.7403 | 0.0301 | 0.0688 | 5 |
| 400 ms | 0.6979 | 0.7428 | 0.0301 | 0.0813 | 13 |
| 800 ms | 0.7463 | 0.7672 | 0.0361 | 0.1438 | 26 |

At 800 ms, the mean absolute score shift is 0.1058; four false interruptions
are introduced and one is resolved. Since controller policy separately consumes
silence duration, the current interface can double-count silence. Trailing-
silence jitter or speech-end normalization, followed by controller
recalibration on real sequences, is now a priority rather than an optional
augmentation.

## 5. Evaluation and threshold protocol

### 5.1 Development sequence

For a final run, the intended sequence is:

1. Audit all train shards and fix hard schema/audio errors.
2. Freeze the transitive grouping policy before comparing models.
3. Train on the group-safe train split.
4. Select architecture/hyperparameters on development loss and the constrained
   operating point, never the official test.
5. Run source-held-out and leave-one-source-out stress tests as secondary
   robustness evidence.
6. Calibrate threshold on a reserved calibration partition. If INT8 changes
   scores, recalibrate after quantization.
7. Freeze model hash, frontend metadata, controller parameters, and threshold.
8. Open the official test once, publish all prespecified metrics and slices, and
   do not tune against it.

The one-shard preview did not fully satisfy that ideal: the same 326 examples
informed checkpoints, the two model thresholds, and follow-up experiments. Its
privately stored predictions support exact reproduction and exploratory
comparison, but not an unbiased holdout estimate.

### 5.2 Confidence and comparison

Rows from one conversation or speaker are correlated. Final intervals should
resample speakers or conversations—not individual rows. The current IID
manifest cannot do that because every observed leakage group is a singleton;
its paired bootstrap is row-level in effect. The source stress bootstrap uses
nine validation source groups, but source is heavily confounded. Small slices
must show raw numerator and denominator, not a visually confident percentage.

### 5.3 End-to-end replay

Offline clip classification is not the product. Real VAD checkpoints should be
logged as JSONL with timestamp, silence duration, score, target, and inference
time. `scripts/replay_stream.py --metadata <model_metadata.json>` feeds those
checkpoints through the exact serialized controller configuration. The Gradio
demo is a one-checkpoint visualizer and creates a fresh controller for each
click; stateful sequence semantics are exercised by replay. This exposes double-trigger, debounce,
minimum/maximum-wait, and latency behavior that a per-clip confusion matrix
cannot measure.

The repository also includes a hand-authored synthetic replay fixture and its
integration output. Its probabilities, targets, and latency values are not
model outputs or human annotations. It validates policy binding, threshold
relaxation, timeout, END latching, and duplicate-emission prevention only; it
is not model, conversation, latency, or product-quality evidence.

No genuine turn or conversation IDs exist in the local shard, so sequence
metrics are unavailable. The reported interruptions-per-audio-hour value is a
clip-normalized proxy and is not an online interruption rate.

## 6. Runtime and release design

The controller has three states:

- `SPEAKING` after observed voice activity;
- `HOLD` during an incomplete or too-short pause;
- `END` when the calibrated score is confirmed or the maximum timeout fires.

Before `min_silence_ms`, the system always holds. After `relax_after_ms`, the
threshold decreases linearly from the normal endpoint threshold toward a
long-pause threshold. At `max_silence_ms`, it ends regardless of score. This
bounds both interruption and indefinite waiting behavior and makes policy
changes testable without retraining.

END is terminal for the current turn. The transition emits one
`emit_response=true` edge; later pause callbacks remain in the latched END
state but emit false until `observe_speech()` or `reset()` starts a new turn.
The replay summary reports duplicate response emissions explicitly.

ONNX export produces probability—not logits—and a self-describing JSON contract
with input names/types/shapes, frontend, threshold, dynamic/fixed frame behavior,
file hashes, sizes, and parity error. FP32 parity must be within `1e-4` in the
export script. Static INT8 requires representative calibration log-mels; any
quantized model requires new score calibration.

The selected FP32 export has exact measured parity (`max_abs_error = 0.0`) and
is 625,431 bytes. Batch-1, one-thread, 20-warmup/200-iteration measurements on
the recorded Apple-arm machine are:

| Scope | Runtime | p50 | p95 | p99 | Load | Peak RSS |
|---|---|---:|---:|---:|---:|---:|
| Neural only, 400 frames | PyTorch 2.11 | 2.855 ms | 2.949 ms | 3.038 ms | 4.303 ms | 224.8 MB |
| Neural only, 400 frames | ONNX Runtime 1.26 | 0.919 ms | 0.956 ms | 0.974 ms | 6.041 ms | 67.0 MB |
| Four-second waveform → probability | ONNX Runtime 1.26 | 1.572 ms | 1.629 ms | 1.651 ms | 5.814 ms | 78.2 MB |

The end-to-end row includes canonical waveform preprocessing and neural
inference, but excludes VAD, transport, controller waiting, and concurrent
production load. It is a machine-specific implementation measurement, not a
service SLA.

The Gradio demo accepts microphone or uploaded audio and exposes assumed silence,
threshold, state, reason, score, and inference latency. If the model is absent,
it clearly labels a heuristic development fallback. The release builder rejects
smoke models and missing test metrics by default. The publisher checks the
authenticated Hugging Face identity before creating or uploading repositories,
never logs the token, verifies stable remote HEAD snapshots, and requires an
observed build/start transition after a new Space upload. The Hub runtime API
does not expose a deployed commit SHA, so the receipt records
`runtime_commit_bound=false` even when content and `RUNNING` are verified.

The release inventory binds the exact 2.7 MB split manifest by SHA-256 but does
not redistribute it: the manifest contains upstream `spoken_text` and record
identifiers whose redistribution rights are unresolved. Reproduction instead
ships deterministic preparation code, immutable data revision, manifest hash,
and aggregate reports. This is a deliberate privacy/rights trade-off, not an
omitted provenance link.

## 7. Failure analysis

The most important failure families are:

- **filled pauses:** “uh”, “matlab”, “haan”, elongated vowels;
- **syntactic continuation:** conjunctions, subordinate clauses, enumeration;
- **slot continuation:** incomplete phone, PIN code, order ID, address;
- **self-repair:** corrections after a short pause;
- **backchannels:** “haan”, “achha”, “right” used to keep the floor;
- **prosody mismatch:** rising/level intonation at apparent sentence boundaries;
- **acoustic shift:** phone codecs, traffic, far-field speech, clipping, music;
- **duration policy:** short true endings and long deliberate holds;
- **domain/source artifacts:** synthetic cadence or recording pipeline shortcuts.

At the selected threshold, the TinyTCN has 3 false interruptions and 150 missed
ends. All three false interruptions are synthetic-tagged; two occur in the
134-row `midfiller` slice. Those are counts for prioritization, not causal
explanations: no blinded listening review has assigned acoustic or linguistic
failure tags. Privacy-safe case IDs are retained for authorized local review;
the public aggregate contains no audio, transcript, raw record ID, or source
path.

The source-held-out split reveals a data-design failure mode: source identity is
almost a proxy for human/synthetic status and language in the inspected shard.
Both models drop to roughly 0.55 AUROC. The correct response is to label it a
joint stress test and improve collection—not to hide the slice or overinterpret
it. The silence perturbation also reveals a concrete interface risk: score and
controller policy can both react to the same trailing silence.

Detailed review fields and promotion gates are in
[`docs/failure_analysis.md`](docs/failure_analysis.md).

## 8. Fourteen-day execution map

The code was assembled in an accelerated implementation pass; “Day” below is
the intended dependency order for a two-week hiring sprint, not a claim that 14
calendar days or the missing external data collection occurred.

| Day | Intended outcome | Repository evidence | Current state |
|---:|---|---|---|
| 1 | Reframe task and define error economics | This report; FPR-budget metric contract | Complete |
| 2 | Pin and download data safely | `download_dataset.sh`, revisions, shard verification | Implemented; 1/83 train shards local |
| 3 | Audit schema/audio/labels | `src/turn_detection/data/audit.py`, partial audit report | Complete on local shard |
| 4 | Eliminate duplicate and identity leakage | privacy-preserving keys + transitive union-find | Implemented and tested |
| 5 | Build IID and domain stress splits | iterative group stratification, holdouts, LOO APIs | Complete on local shard |
| 6 | Establish reproducible smoke baseline | deterministic generated-feature run | Complete; integration only |
| 7 | Train from-scratch TinyTCN | model, loss, trainer, partial-shard checkpoints | 151,812-param preview trained; full run pending |
| 8 | Run Whisper teacher baseline | teacher and config | Implemented; run pending |
| 9 | Add distillation/ablation path | distillation objective | Primitive implemented; run pending |
| 10 | Evaluate product metrics and slices | IID/source reports, paired bootstrap, privacy-safe failures | Complete for partial-shard development; sequence/final evidence pending |
| 11 | Collect Hinglish hard cases | 900 assignments and consent protocol | Protocol complete; recordings pending |
| 12 | Export/quantize/benchmark | FP32 ONNX, exact parity, PyTorch/ONNX/end-to-end CPU reports | FP32 preview complete; INT8/final artifact pending |
| 13 | Integrate controller and demo | replay, three-state controller, Gradio, Hinglish recording prompts | Preview model and synthetic controller integration complete; real sequence replay pending |
| 14 | Freeze test, package, publish | sealed-test guard, atomic release builder, exact-mirror publisher | Development packaging path ready; official-test/final gates not satisfied |

The repository quality gate passed Ruff 0.15.16 and all 136 automated tests
(131 passed via Python `unittest` discovery). That validates the implemented behavior under test; it does not fill
the missing empirical datasets or rights review.

## 9. Decision log and next experiments

1. **Keep the acoustic baseline as the evidence winner.** The preview TinyTCN is
   the deployable neural research artifact, but it has not earned promotion on
   IID constrained recall.
2. **Use filler heads as auxiliary supervision.** Their labels are useful but
   nullable; masked loss avoids inventing negatives.
3. **Retain a four-second preview window, then ablate.** Endpoint evidence is at
   the end, but the current experiment does not prove four or eight seconds
   optimal. Compare 2/4/8 seconds on a fresh development design.
4. **Treat source holdout as stress, not causality.** Current confounding is too
   large for a clean source claim.
5. **Optimize under an FPR budget.** Interruptions deserve an explicit bound.
6. **Keep policy outside the network.** Silence and timeout trade-offs remain
   inspectable and replayable.
7. **Do not open test early.** A guarded script turns this from advice into an
   operational control.
8. **Remove silence double-counting.** Train with trailing-silence jitter or
   speech-end normalization, then calibrate the silence-aware controller on
   genuine multi-checkpoint turns.

The priority experiment matrix is:

- TinyTCN channels/blocks: `96×6`, `128×8`, `192×8`;
- suffix duration: 2, 4, 8 seconds;
- endpoint-only versus filler multitask loss;
- no normalization versus mask-aware per-mel normalization;
- IID, leave-one-source-out, and human-only validation;
- Whisper frozen-head baseline, selective unfreezing, then distillation;
- FP32 versus static INT8 with post-quantization threshold calibration;
- controller sweeps over threshold, minimum silence, relaxation, maximum timeout,
  and confirmation count using real pause sequences.

Selection should use a Pareto frontier over false interruptions, endpoint delay,
model bytes, and end-to-end CPU p95—not a leaderboard of accuracy alone.

## 10. Limitations, ethics, and release risk

- The neural preview underperforms the simple baseline on current IID
  constrained recall.
- Development data was reused adaptively for checkpoint and threshold selection;
  its intervals do not remove that optimism.
- The local audit covers 1/83 of train and may be systematically unrepresentative.
- The upstream mixture is mostly synthetic in the inspected shard.
- Language tags do not validate natural Hinglish code-switching.
- No local Hinglish audio, device/noise diversity, inter-annotator agreement, or
  consented demographic fairness result exists yet.
- A clip-level classifier cannot by itself prove streaming conversational quality.
- Appended silence shifts scores enough to breach the nominal measured 2% FPR
  operating point, while the controller also consumes silence duration.
- A maximum timeout can still interrupt a very long thinking pause; that is a
  deliberate policy trade-off and should be disclosed.
- Audio may contain biometric information and PII. Logs should prefer scores and
  anonymous IDs over raw speech, with retention and deletion controls.
- The inspected upstream dataset card does not declare an explicit dataset
  license. Repository code is Apache-2.0, but that does not grant rights to
  upstream audio, collected recordings, or necessarily derived artifacts.
  Redistribution and commercial-use terms require review before publication.

This detector is intended only for conversational timing. It must not be
repurposed for speaker identification, emotion or health inference, surveillance,
or consequential decisions.

## Conclusion

The main contribution is not a premature benchmark number. It is a small,
inspectable system that makes the dangerous parts of turn detection explicit:
leakage, filler-label missingness, domain confounding, asymmetric errors,
calibration, streaming policy, test contamination, preprocessing parity, and
release provenance.

The concrete result is a small and fast exported system whose present neural
model does not yet beat a ten-feature baseline. That negative result, the
source-shift collapse, and the silence sensitivity define the next experiments
more clearly than a polished demo alone could. The repository is ready for a
full-data, newly partitioned experiment; the preview weights are not
production-ready.