Technical report: Tiny Hinglish Turn Detector
Challenge: decide, from audio, whether a user is finished speaking or only pausing.
Author: Suvradeep Das
Repository state documented: 2026-08-23
Current verdict: a fast, fully exported neural preview exists, but the interpretable acoustic baseline is still stronger. Generalization remains unestablished.
Executive summary
I treated turn detection as a product decision under asymmetric error costs,
not as a generic balanced classification benchmark. A premature END causes a
visible interruption; a late END costs latency. The system therefore separates
three concerns:
- an audio classifier estimates
p(END)at a VAD pause checkpoint; - a threshold is selected to maximize endpoint recall under a predeclared false-interruption budget;
- a deterministic controller applies minimum silence, long-pause relaxation, optional confirmation, and a hard maximum timeout.
The packaged preview is a 151,812-parameter, four-second log-mel TinyTCN trained from scratch and then continued with a fresh optimizer. It has six causal depthwise-separable residual blocks and mask-aware attentive statistics pooling, with auxiliary heads for mid-turn and end-turn fillers. The FP32 ONNX file is 625,431 bytes and measures 1.629 ms p95 from a four-second waveform to probability on the recorded Apple-arm/one-thread setup. A larger 396,164- parameter configuration and a Whisper-tiny teacher remain proposed full-data experiments, not packaged results.
The repository covers the full engineering path: pinned/resumable data download, streaming audit, transitive leakage grouping, deterministic split generation, lazy audio loading, training, calibration, sliced and operational metrics, ONNX export/parity, quantization hooks, CPU benchmarking, runtime replay, Gradio, and guarded Hugging Face packaging/publishing.
What it does not contain is equally important: only one of 83 train shards is present locally; there is no full-corpus run; the official test has not been downloaded or evaluated; and no human Hinglish recordings have been collected. The upstream shard also lacks usable speaker/conversation/voice IDs, making the IID split best-effort row-disjoint rather than speaker-disjoint.
On the 326-row IID validation split, the acoustic logistic baseline is the current winner: AP 0.7310, AUROC 0.7522, and 21/160 END recall at 3/166 false interruptions. The TinyTCN reaches AP 0.6881, AUROC 0.7390, and 10/160 END recall at the same three false interruptions. Its constrained-recall delta is −0.0688 with a paired bootstrap 95% interval of [−0.1330, −0.0122]. That is a useful negative result. Both models and thresholds were selected adaptively on the same development data, so the result is not an independent estimate.
1. Problem definition
At each pause checkpoint, the target is:
y = 0(HOLD): responding now would interrupt an incomplete turn;y = 1(END): the turn is complete and the agent may respond.
The model consumes only the recent audio suffix. This is intentionally narrower than end-to-end dialogue policy: VAD decides when speech has paused, the endpoint model estimates conversational completeness, and the controller turns that score into a time-bounded action.
1.1 Error economics
Let FP be an incomplete turn predicted as END, and FN a complete turn
predicted as HOLD.
FPis a false interruption. It can cut off an address, order ID, correction, or filler-delayed continuation.FNadds waiting time, bounded by the controller's maximum timeout.
Because the first is normally more harmful, the primary operating point is not
the maximum-F1 threshold. It is the threshold with maximum END recall subject to
FPR <= budget, where FPR = FP / (FP + TN). The code reports budgets of 1%,
2%, and 5%; the primary training configuration selects at 2%. A product owner
can change that budget without retraining the representation.
1.2 Product-level success criteria
The final scorecard should contain:
- false-positive rate and END recall at the frozen threshold;
- false interruptions per turn and per audio hour;
- p50/p90/p95 endpoint delay after a true completion;
- AUROC and average precision for threshold-independent ranking;
- Brier score, log loss, ECE, and reliability bins for calibration;
- language, source, synthetic/human, filler, duration, device/noise, and consented demographic slices;
- grouped bootstrap intervals by speaker or conversation;
- neural-only and end-to-end p50/p95/p99 batch-1 CPU latency, model size, load time, and peak RSS.
“Tiny + fast + accurate” is not satisfied by a single aggregate F1 value.
2. Data work
2.1 Upstream snapshot
The train source is
pipecat-ai/smart-turn-data-v3.2-train,
pinned to revision
e564e2ac567f774d1880aa1db6ce97afb8c519b7. The repository catalog describes
270,946 rows across 83 Parquet shards (about 41.4 GB). The download script is
resumable and verifies the shard count before success.
The official test source is a separate 10-shard snapshot, pinned to
0500378e8ed6d38e37b016e24d261e8e6c6a6859. Its downloader requires the exact
acknowledgement I_HAVE_FROZEN_MODEL_AND_THRESHOLD. It has not been downloaded
or inspected.
2.2 Measured local audit
Only train-00010-of-00083.parquet is local. The following numbers therefore
describe that shard only, not the full dataset:
| Audit field | Measured value |
|---|---|
| Rows | 3,265 |
| Valid / invalid | 3,265 / 0 |
| Encoded duration | 24,685.517 s = 6.857 h |
| Duration range | 0.36–29.56 s |
| Audio format | 3,265 FLAC |
Endpoint HOLD / END |
1,667 / 1,598 |
| Synthetic / human-tagged | 2,712 / 553 |
| Languages | 23 |
| English / Hindi / Marathi rows | 758 / 158 / 81 |
| Exact-audio duplicates | 0 |
| Duplicate record IDs | 0 |
| Conflicting duplicate labels | 0 |
The inspected shard is roughly 83% synthetic. Its sequence and source mix may be non-random, so no count should be extrapolated to the other 82 shards.
The nullable filler labels were preserved rather than silently coerced to
false: midfiller has 1,339 true, 1,284 false, and 642 null values;
endfiller has 905 true, 1,718 false, and 642 null values. Auxiliary losses mask
the nulls.
The machine-readable evidence is in
reports/partial_shard_audit.json.
2.3 Leakage control
Exact duplicates are only the simplest leakage route. Two rows may share a
conversation, speaker, TTS voice, recording, prompt, record ID, or repeated
transcript while having different audio hashes. The audit derives hashed linkage
keys for each available identifier. Repeated normalized text is linked globally
when it is long enough to be meaningful. A union-find pass computes transitive
connected components, so A shares audio with B and B shares speaker with C
forces all three into one group.
Splitting assigns whole groups with deterministic iterative multilabel stratification. Validation checks all three independent crossing classes:
- component
group_id; - exact
audio_sha256; - every underlying metadata linkage key.
The 90/10 development split has 2,939 train and 326 validation rows, including
1,438/1,501 END/HOLD train and 160/166 END/HOLD validation examples. All crossing
maps are empty. See
reports/partial_iid_split.json.
That zero-crossing result is narrower than speaker-disjointness. The shard does not expose usable speaker, conversation, or TTS-voice IDs; every one of the 3,265 base components used by the IID split is a singleton. (The source stress split separately groups by its 12 dataset values.) The IID split is best-effort row-disjoint and catches exact audio/observed-key leakage, while identity and repeated-voice leakage remain unknown. Group-bootstrap intervals on this shard therefore reduce to row-bootstrap intervals and must not be described as speaker-clustered.
2.4 Source-held-out stress test
A second split holds entire dataset values together: 2,617 train and 648
validation rows, again with zero detected crossings. It is useful as a joint
domain-shift test, but not as a clean estimate of source robustness:
- train is 2,614 synthetic and only 3 human-tagged rows;
- validation is 98 synthetic and 550 human-tagged rows;
- train spans 23 languages; validation contains only English and Spanish.
Any performance delta simultaneously measures source, language, human/synthetic,
and collection-process shift. Reporting it as “generalization to unseen source”
without this caveat would be misleading. Leave-one-source-out folds are a better
next analysis because they expose variability across individual sources, though
the confounding cannot be removed without better data. Evidence:
reports/partial_source_holdout_split.json.
2.5 Why a separate Hinglish set is necessary
Single ISO tags such as hin and eng do not establish within-utterance
code-switching, Indian accents, natural hesitation, or logistics vocabulary.
The repository includes a purpose-built protocol with 900 balanced assignments:
- 30 speakers × 30 recordings;
- 450 HOLD and 450 END targets;
- 21/4/5 speaker-disjoint train/dev/test speakers, producing 630/120/150 examples;
- minimal pairs covering address, COD, pickup, reschedule, returns, tracking, support, phone/order identifiers, delivery instructions, and damaged parcels;
- natural 0.3–2.0 second internal pauses for HOLD recordings;
- independent “would responding now feel like an interruption?” judgments from three annotators.
The assignment file is real and deterministic. The recordings are not: no
participants have been recorded locally. Collection requires informed consent,
PII controls, retention/deletion rules, and a redistribution decision before any
audio leaves private storage. See
docs/data_collection_protocol.md.
3. Model design
3.1 Canonical frontend
Audio is converted to mono float, resampled to 16 kHz, suffix-cropped, and padded with an explicit validity mask. The packaged preview uses four seconds; the proposed full-data configuration uses eight. Both use:
- 400-sample (25 ms) Hann window;
- 160-sample (10 ms) hop;
- 80 area-normalized HTK mel bands from 0–8 kHz;
- base-10 log energy with an eight-log-unit dynamic-range floor;
- Whisper-style
(log_mel + 4) / 4scaling, without per-utterance mean/variance standardization in the primary configuration.
An explicit frame mask prevents padding from affecting convolutional states or attentive pooling. The training and dependency-light NumPy serving frontends must pass parity tests before release. The Whisper teacher uses Slaney mel bands with the Whisper log convention and fixed encoder length; this is serialized separately rather than silently reusing the student's HTK filterbank.
3.2 TinyTCN student
For packaged input X of shape [batch, 80, frames]:
- a 1×1 projection maps 80 mel bins to 128 channels;
- six residual depthwise-separable causal Conv1d blocks use kernel 5 and
dilations
[1, 2, 4, 8, 16, 32]; - per-frame LayerNorm avoids padding-dependent batch/time statistics;
- masked learned attention produces weighted mean and standard deviation;
- a 96-unit embedding feeds endpoint, midfiller, and endfiller logits.
This selected preview has 151,812 parameters. The configurable implementation also supports the proposed 192-channel/eight-block/eight-second model with 396,164 parameters. That larger design has not been shown to improve current development metrics and is not the packaged checkpoint.
Attentive pooling still recomputes a suffix-level summary at a checkpoint. A future streaming version can cache convolutional states and maintain online pooling statistics, but the current implementation makes no such latency claim.
3.3 Objective and optimization
The default loss is:
L = BCE(endpoint) + 0.15 * masked_BCE(midfiller)
+ 0.15 * masked_BCE(endfiller)
Endpoint label smoothing is 0.02 in the primary config. Positive weighting is a single train-split statistic if needed; it is never recomputed per batch. The trainer uses AdamW, gradient clipping, optional CUDA mixed precision, deterministic kernels, configurable metric-based early stopping, and a seed fixed before model construction. The default full-data run selects constrained END recall under the 2% FPR budget. The bounded preview continuation pre-registered average precision for checkpoint selection, while its decision threshold was still selected under the 2% development FPR rule.
The checkpoint contains model configuration, model and optimizer state, frontend configuration, maximum suffix length, selected threshold, run name, framework version, and validation evidence. The resolved config and full epoch history are written next to it.
3.4 Whisper teacher and distillation
The alternative starts with openai/whisper-tiny, uses the correct Whisper
log-mel convention and attention mask, pools encoder states with the same
mask-aware head, and can freeze all or selectively unfreeze final encoder
layers. The default experiment is head-only.
A hard/soft distillation objective is implemented, including temperature scaling and optional embedding alignment. It is not yet wired into a completed teacher-student run; presenting it as a measured gain would be false. The proper sequence is teacher validation, cached teacher logits, student distillation, and paired evaluation at the same false-interruption budget.
4. Experiments completed so far
Smoke and resolver runs first established deterministic mechanics. The rows below are the bounded real-data experiments; all use one audited train shard.
| Run | Split/model | Result | Decision |
|---|---|---|---|
| Acoustic baseline | IID, ten four-second waveform statistics + class-balanced logistic regression | AP 0.7310, AUROC 0.7522, Brier 0.2038; threshold 0.7639 gives 21 TP / 3 FP / 163 TN / 139 FN | Current local winner and development floor |
| Initial TinyTCN | IID, 151,812 parameters, four seconds, three epochs, LR 3e-4 |
AP 0.6874, AUROC 0.7291, recall 0.0938 at FPR 0.0181 | Viable ranking; continue under bounded plan |
| High-LR TinyTCN | Same architecture, planned eight epochs, LR 1e-3 |
Stopped after epoch 4; best AP 0.6190, AUROC 0.6829, recall 0.0313 at FPR 0.0181 | Stop for poor trajectory |
| Packaged warm continuation | IID, weights-only continuation, fresh optimizer, five epochs at LR 3e-4 |
Selected continuation epoch 1: AP 0.6881, AUROC 0.7390, Brier 0.2133; threshold 0.7410 gives 10 TP / 3 FP / 163 TN / 150 FN | Package as development preview, not as winner |
| Source-stress baseline | Source-held-out logistic baseline | AP 0.5368, AUROC 0.5539, recall 0.0279 at FPR 0.0185 | Severe joint-shift collapse |
| Source-stress TinyTCN | Source-held-out TinyTCN, three epochs | AP 0.5500, AUROC 0.5523, recall 0.0402 at FPR 0.0185 | No established advantage; severe joint-shift collapse |
The warm continuation retained its first epoch; later continuation epochs fell to AP 0.6798, 0.6735, 0.6723, and 0.6726. The improvement over the original checkpoint is marginal and the development split had already been inspected. This is adaptive checkpoint selection, not a fresh replication.
4.1 Paired IID comparison
Both model thresholds were chosen on the same 326-row validation set. With the selected thresholds held fixed, a 2,000-sample paired bootstrap gives TinyTCN minus baseline:
| Metric | Difference | 95% interval |
|---|---|---|
| AP | −0.0429 | [−0.1195, 0.0352] |
| AUROC | −0.0132 | [−0.0782, 0.0512] |
| END recall | −0.0688 | [−0.1330, −0.0122] |
| FPR | 0.0000 | [−0.0284, 0.0298] |
The AP/AUROC intervals include zero; the constrained-recall interval favors the baseline. Because all 326 validation groups are singletons, this “grouped” bootstrap is numerically row bootstrap, not speaker/conversation-clustered uncertainty. It also conditions on adaptively selected checkpoints and thresholds.
The TinyTCN absolute FPR interval is [0, 0.0407] and its recall interval is
[0.0261, 0.1019]. Observing 3/166 false interruptions is compatible with more
than a 2% population FPR, so the nominal budget is not statistically certified.
4.2 Source-held-out stress
The 648-row validation side contains nine source groups. TinyTCN-minus-baseline
paired deltas are AP +0.0131 (95% CI [−0.2149, 0.1261]), AUROC −0.0016
([−0.2713, 0.1273]), and recall +0.0124
([−0.0112, 0.0948]). The intervals are wide and both AUROCs are near 0.55.
Combined with the split's source/language/human-synthetic confounding, the only
defensible conclusion is that both models are brittle under this joint shift.
4.3 Silence perturbation
Appending valid zero-valued silence before suffix cropping changes the selected model materially at its frozen threshold:
| Appended silence | AP | AUROC | FPR | END recall | Flips vs 0 ms |
|---|---|---|---|---|---|
| 0 ms | 0.6881 | 0.7390 | 0.0181 | 0.0625 | — |
| 200 ms | 0.6840 | 0.7403 | 0.0301 | 0.0688 | 5 |
| 400 ms | 0.6979 | 0.7428 | 0.0301 | 0.0813 | 13 |
| 800 ms | 0.7463 | 0.7672 | 0.0361 | 0.1438 | 26 |
At 800 ms, the mean absolute score shift is 0.1058; four false interruptions are introduced and one is resolved. Since controller policy separately consumes silence duration, the current interface can double-count silence. Trailing- silence jitter or speech-end normalization, followed by controller recalibration on real sequences, is now a priority rather than an optional augmentation.
5. Evaluation and threshold protocol
5.1 Development sequence
For a final run, the intended sequence is:
- Audit all train shards and fix hard schema/audio errors.
- Freeze the transitive grouping policy before comparing models.
- Train on the group-safe train split.
- Select architecture/hyperparameters on development loss and the constrained operating point, never the official test.
- Run source-held-out and leave-one-source-out stress tests as secondary robustness evidence.
- Calibrate threshold on a reserved calibration partition. If INT8 changes scores, recalibrate after quantization.
- Freeze model hash, frontend metadata, controller parameters, and threshold.
- Open the official test once, publish all prespecified metrics and slices, and do not tune against it.
The one-shard preview did not fully satisfy that ideal: the same 326 examples informed checkpoints, the two model thresholds, and follow-up experiments. Its privately stored predictions support exact reproduction and exploratory comparison, but not an unbiased holdout estimate.
5.2 Confidence and comparison
Rows from one conversation or speaker are correlated. Final intervals should resample speakers or conversations—not individual rows. The current IID manifest cannot do that because every observed leakage group is a singleton; its paired bootstrap is row-level in effect. The source stress bootstrap uses nine validation source groups, but source is heavily confounded. Small slices must show raw numerator and denominator, not a visually confident percentage.
5.3 End-to-end replay
Offline clip classification is not the product. Real VAD checkpoints should be
logged as JSONL with timestamp, silence duration, score, target, and inference
time. scripts/replay_stream.py --metadata <model_metadata.json> feeds those
checkpoints through the exact serialized controller configuration. The Gradio
demo is a one-checkpoint visualizer and creates a fresh controller for each
click; stateful sequence semantics are exercised by replay. This exposes double-trigger, debounce,
minimum/maximum-wait, and latency behavior that a per-clip confusion matrix
cannot measure.
The repository also includes a hand-authored synthetic replay fixture and its integration output. Its probabilities, targets, and latency values are not model outputs or human annotations. It validates policy binding, threshold relaxation, timeout, END latching, and duplicate-emission prevention only; it is not model, conversation, latency, or product-quality evidence.
No genuine turn or conversation IDs exist in the local shard, so sequence metrics are unavailable. The reported interruptions-per-audio-hour value is a clip-normalized proxy and is not an online interruption rate.
6. Runtime and release design
The controller has three states:
SPEAKINGafter observed voice activity;HOLDduring an incomplete or too-short pause;ENDwhen the calibrated score is confirmed or the maximum timeout fires.
Before min_silence_ms, the system always holds. After relax_after_ms, the
threshold decreases linearly from the normal endpoint threshold toward a
long-pause threshold. At max_silence_ms, it ends regardless of score. This
bounds both interruption and indefinite waiting behavior and makes policy
changes testable without retraining.
END is terminal for the current turn. The transition emits one
emit_response=true edge; later pause callbacks remain in the latched END
state but emit false until observe_speech() or reset() starts a new turn.
The replay summary reports duplicate response emissions explicitly.
ONNX export produces probability—not logits—and a self-describing JSON contract
with input names/types/shapes, frontend, threshold, dynamic/fixed frame behavior,
file hashes, sizes, and parity error. FP32 parity must be within 1e-4 in the
export script. Static INT8 requires representative calibration log-mels; any
quantized model requires new score calibration.
The selected FP32 export has exact measured parity (max_abs_error = 0.0) and
is 625,431 bytes. Batch-1, one-thread, 20-warmup/200-iteration measurements on
the recorded Apple-arm machine are:
| Scope | Runtime | p50 | p95 | p99 | Load | Peak RSS |
|---|---|---|---|---|---|---|
| Neural only, 400 frames | PyTorch 2.11 | 2.855 ms | 2.949 ms | 3.038 ms | 4.303 ms | 224.8 MB |
| Neural only, 400 frames | ONNX Runtime 1.26 | 0.919 ms | 0.956 ms | 0.974 ms | 6.041 ms | 67.0 MB |
| Four-second waveform → probability | ONNX Runtime 1.26 | 1.572 ms | 1.629 ms | 1.651 ms | 5.814 ms | 78.2 MB |
The end-to-end row includes canonical waveform preprocessing and neural inference, but excludes VAD, transport, controller waiting, and concurrent production load. It is a machine-specific implementation measurement, not a service SLA.
The Gradio demo accepts microphone or uploaded audio and exposes assumed silence,
threshold, state, reason, score, and inference latency. If the model is absent,
it clearly labels a heuristic development fallback. The release builder rejects
smoke models and missing test metrics by default. The publisher checks the
authenticated Hugging Face identity before creating or uploading repositories,
never logs the token, verifies stable remote HEAD snapshots, and requires an
observed build/start transition after a new Space upload. The Hub runtime API
does not expose a deployed commit SHA, so the receipt records
runtime_commit_bound=false even when content and RUNNING are verified.
The release inventory binds the exact 2.7 MB split manifest by SHA-256 but does
not redistribute it: the manifest contains upstream spoken_text and record
identifiers whose redistribution rights are unresolved. Reproduction instead
ships deterministic preparation code, immutable data revision, manifest hash,
and aggregate reports. This is a deliberate privacy/rights trade-off, not an
omitted provenance link.
7. Failure analysis
The most important failure families are:
- filled pauses: “uh”, “matlab”, “haan”, elongated vowels;
- syntactic continuation: conjunctions, subordinate clauses, enumeration;
- slot continuation: incomplete phone, PIN code, order ID, address;
- self-repair: corrections after a short pause;
- backchannels: “haan”, “achha”, “right” used to keep the floor;
- prosody mismatch: rising/level intonation at apparent sentence boundaries;
- acoustic shift: phone codecs, traffic, far-field speech, clipping, music;
- duration policy: short true endings and long deliberate holds;
- domain/source artifacts: synthetic cadence or recording pipeline shortcuts.
At the selected threshold, the TinyTCN has 3 false interruptions and 150 missed
ends. All three false interruptions are synthetic-tagged; two occur in the
134-row midfiller slice. Those are counts for prioritization, not causal
explanations: no blinded listening review has assigned acoustic or linguistic
failure tags. Privacy-safe case IDs are retained for authorized local review;
the public aggregate contains no audio, transcript, raw record ID, or source
path.
The source-held-out split reveals a data-design failure mode: source identity is almost a proxy for human/synthetic status and language in the inspected shard. Both models drop to roughly 0.55 AUROC. The correct response is to label it a joint stress test and improve collection—not to hide the slice or overinterpret it. The silence perturbation also reveals a concrete interface risk: score and controller policy can both react to the same trailing silence.
Detailed review fields and promotion gates are in
docs/failure_analysis.md.
8. Fourteen-day execution map
The code was assembled in an accelerated implementation pass; “Day” below is the intended dependency order for a two-week hiring sprint, not a claim that 14 calendar days or the missing external data collection occurred.
| Day | Intended outcome | Repository evidence | Current state |
|---|---|---|---|
| 1 | Reframe task and define error economics | This report; FPR-budget metric contract | Complete |
| 2 | Pin and download data safely | download_dataset.sh, revisions, shard verification |
Implemented; 1/83 train shards local |
| 3 | Audit schema/audio/labels | src/turn_detection/data/audit.py, partial audit report |
Complete on local shard |
| 4 | Eliminate duplicate and identity leakage | privacy-preserving keys + transitive union-find | Implemented and tested |
| 5 | Build IID and domain stress splits | iterative group stratification, holdouts, LOO APIs | Complete on local shard |
| 6 | Establish reproducible smoke baseline | deterministic generated-feature run | Complete; integration only |
| 7 | Train from-scratch TinyTCN | model, loss, trainer, partial-shard checkpoints | 151,812-param preview trained; full run pending |
| 8 | Run Whisper teacher baseline | teacher and config | Implemented; run pending |
| 9 | Add distillation/ablation path | distillation objective | Primitive implemented; run pending |
| 10 | Evaluate product metrics and slices | IID/source reports, paired bootstrap, privacy-safe failures | Complete for partial-shard development; sequence/final evidence pending |
| 11 | Collect Hinglish hard cases | 900 assignments and consent protocol | Protocol complete; recordings pending |
| 12 | Export/quantize/benchmark | FP32 ONNX, exact parity, PyTorch/ONNX/end-to-end CPU reports | FP32 preview complete; INT8/final artifact pending |
| 13 | Integrate controller and demo | replay, three-state controller, Gradio, Hinglish recording prompts | Preview model and synthetic controller integration complete; real sequence replay pending |
| 14 | Freeze test, package, publish | sealed-test guard, atomic release builder, exact-mirror publisher | Development packaging path ready; official-test/final gates not satisfied |
The repository quality gate passed Ruff 0.15.16 and all 136 automated tests
(131 passed via Python unittest discovery). That validates the implemented behavior under test; it does not fill
the missing empirical datasets or rights review.
9. Decision log and next experiments
- Keep the acoustic baseline as the evidence winner. The preview TinyTCN is the deployable neural research artifact, but it has not earned promotion on IID constrained recall.
- Use filler heads as auxiliary supervision. Their labels are useful but nullable; masked loss avoids inventing negatives.
- Retain a four-second preview window, then ablate. Endpoint evidence is at the end, but the current experiment does not prove four or eight seconds optimal. Compare 2/4/8 seconds on a fresh development design.
- Treat source holdout as stress, not causality. Current confounding is too large for a clean source claim.
- Optimize under an FPR budget. Interruptions deserve an explicit bound.
- Keep policy outside the network. Silence and timeout trade-offs remain inspectable and replayable.
- Do not open test early. A guarded script turns this from advice into an operational control.
- Remove silence double-counting. Train with trailing-silence jitter or speech-end normalization, then calibrate the silence-aware controller on genuine multi-checkpoint turns.
The priority experiment matrix is:
- TinyTCN channels/blocks:
96×6,128×8,192×8; - suffix duration: 2, 4, 8 seconds;
- endpoint-only versus filler multitask loss;
- no normalization versus mask-aware per-mel normalization;
- IID, leave-one-source-out, and human-only validation;
- Whisper frozen-head baseline, selective unfreezing, then distillation;
- FP32 versus static INT8 with post-quantization threshold calibration;
- controller sweeps over threshold, minimum silence, relaxation, maximum timeout, and confirmation count using real pause sequences.
Selection should use a Pareto frontier over false interruptions, endpoint delay, model bytes, and end-to-end CPU p95—not a leaderboard of accuracy alone.
10. Limitations, ethics, and release risk
- The neural preview underperforms the simple baseline on current IID constrained recall.
- Development data was reused adaptively for checkpoint and threshold selection; its intervals do not remove that optimism.
- The local audit covers 1/83 of train and may be systematically unrepresentative.
- The upstream mixture is mostly synthetic in the inspected shard.
- Language tags do not validate natural Hinglish code-switching.
- No local Hinglish audio, device/noise diversity, inter-annotator agreement, or consented demographic fairness result exists yet.
- A clip-level classifier cannot by itself prove streaming conversational quality.
- Appended silence shifts scores enough to breach the nominal measured 2% FPR operating point, while the controller also consumes silence duration.
- A maximum timeout can still interrupt a very long thinking pause; that is a deliberate policy trade-off and should be disclosed.
- Audio may contain biometric information and PII. Logs should prefer scores and anonymous IDs over raw speech, with retention and deletion controls.
- The inspected upstream dataset card does not declare an explicit dataset license. Repository code is Apache-2.0, but that does not grant rights to upstream audio, collected recordings, or necessarily derived artifacts. Redistribution and commercial-use terms require review before publication.
This detector is intended only for conversational timing. It must not be repurposed for speaker identification, emotion or health inference, surveillance, or consequential decisions.
Conclusion
The main contribution is not a premature benchmark number. It is a small, inspectable system that makes the dangerous parts of turn detection explicit: leakage, filler-label missingness, domain confounding, asymmetric errors, calibration, streaming policy, test contamination, preprocessing parity, and release provenance.
The concrete result is a small and fast exported system whose present neural model does not yet beat a ten-feature baseline. That negative result, the source-shift collapse, and the silence sensitivity define the next experiments more clearly than a polished demo alone could. The repository is ready for a full-data, newly partitioned experiment; the preview weights are not production-ready.