suvradeepp's picture
Publish Tiny Hinglish Turn Detector development preview
0348402 verified
|
Raw
History Blame Contribute Delete
14.8 kB
metadata
language:
  - hi
  - en
license: other
library_name: onnxruntime
pipeline_tag: audio-classification
tags:
  - audio
  - turn-detection
  - endpointing
  - hinglish
  - voice-ai
  - onnx
datasets:
  - pipecat-ai/smart-turn-data-v3.2-train

Tiny Hinglish Turn Detector

Tiny Hinglish Turn Detector is an audio-native HOLD/END classifier intended to run at voice-activity-detection pause checkpoints. END means a voice agent may respond; HOLD means responding now risks interrupting the speaker.

Development preview, 2026-08-23. The packaged model was trained on one of 83 upstream training shards. It is not the local winner: an interpretable acoustic logistic baseline performs better on the current IID development split. The official test has not been downloaded or evaluated, and no verified Hinglish benchmark recordings exist. Do not present these numbers as production, official-test, or Hinglish performance.

Packaged artifact

Field Packaged development preview
Task / labels Pause-checkpoint audio classification; HOLD = 0, END = 1
Architecture From-scratch causal depthwise-separable TinyTCN
Parameters 151,812
Blocks / channels 6 / 128, dilations [1, 2, 4, 8, 16, 32]
Input Mono 16 kHz audio; most recent 4 seconds
Features 80-bin HTK log-mel, 25 ms window, 10 ms hop, Whisper-style log scaling
Output Scalar p(END); filler heads are training-only
Frozen preview threshold 0.7410007715
Serving format FP32 ONNX Runtime, dynamic frame axis, batch 1
ONNX size 625,431 bytes (about 611 KiB)
Evidence status preview-only, development data only
License Code: Apache-2.0; weights: no license grant pending upstream-rights review

The network uses a 1Γ—1 input projection, six residual causal depthwise-separable convolution blocks, mask-aware attentive mean/standard- deviation pooling, and a 96-unit endpoint head. Nullable midfiller and endfiller targets are masked rather than converted to negatives.

The convolutions are causal, but attentive pooling summarizes the supplied suffix. This is a pause-checkpoint clip classifier, not a state-cached sample-by-sample streaming network. A larger 396,164-parameter/eight-second configuration remains in the repository as a proposed full-data experiment; it is not the packaged artifact.

Run the packaged model

From a downloaded model repository, the included app automatically resolves the root model.onnx artifact:

python -m pip install -r space/requirements.txt
python app.py

For application integration, use src/turn_detection/runtime/predictor.py so waveform normalization, resampling, log-mel extraction, input names, and the stored threshold come from the adjacent model_metadata.json. Supplying raw ONNX tensors without that contract can change the score semantics.

The stronger ten-feature acoustic logistic model is also shipped as reference_models/acoustic_baseline.json. It is not disguised as the neural submission: it is an auditable evidence winner and a fallback candidate for the next fresh holdout. Load it with LogisticBaseline.from_dict() and compute its inputs with extract_audio_statistics() from turn_detection.baselines.

Intended use

Use the model as one signal in a bounded voice-agent controller:

  1. VAD observes speech followed by a pause;
  2. the current turn's recent audio suffix is scored;
  3. p(END) is compared with a deployment-calibrated threshold;
  4. minimum silence, confirmation, threshold relaxation, and maximum wait are applied outside the network.

The adjacent metadata serializes that controller policy. END latches for the current turn and produces exactly one emit_response=true transition; new speech or an explicit reset starts the next turn.

The packaged controller replay report uses eight hand-authored checkpoints across three synthetic turns and observes three response edges with zero duplicates. Its scores, targets, and latency values are fabricated integration inputs, not ONNX outputs or human annotations; it is plumbing evidence only.

The target application is Indian logistics/customer-support speech, including Hinglish, fillers, self-corrections, identifiers, and natural pauses. That is a design target, not a validated domain claim.

Do not use this artifact for speaker identification, accent/emotion/health or demographic inference, surveillance, employee or customer scoring, or any legal, credit, employment, medical, or other consequential decision. Do not use it for autonomous turn-taking without a controller, monitoring, and a safe timeout policy.

Training data and split boundary

The source is pipecat-ai/smart-turn-data-v3.2-train at revision e564e2ac567f774d1880aa1db6ce97afb8c519b7. Current local evidence covers only train-00010-of-00083.parquet:

  • 3,265 valid rows / 6.857 audio hours;
  • 1,667 HOLD and 1,598 END;
  • 2,712 synthetic and 553 human-tagged rows;
  • 23 language tags, including 158 Hindi and 758 English rows;
  • no exact-audio, record-ID, or conflicting-label duplicates in that shard.

The IID manifest has 2,939 train and 326 validation rows. All observed linkage checks pass, but the raw shard exposes no usable speaker, conversation, or voice identity: all 3,265 base linkage groups are singletons. The split is therefore best-effort row-disjoint and exact-duplicate-safe, not speaker-disjoint. Bootstrap resampling on the IID split is effectively row bootstrap.

The source stress split has 2,617 train and 648 validation rows. Its train side is 2,614/2,617 synthetic while validation is 550/648 human-tagged and contains only English and Spanish. It is a deliberately harsh joint source/language/ human-synthetic shift, not a clean source-causality estimate.

Hindi and English language tags do not prove within-recording Hinglish. A separate 900-recording, 30-speaker consent-first Hinglish protocol is included, but zero participant recordings have been collected, trained on, or evaluated.

Training procedure and selection

The packaged checkpoint is a weights-only continuation of the selected three-epoch LR=3e-4 preview. AdamW state and schedule were intentionally restarted; the continuation used learning rate 3e-4, weight decay 1e-3, batch size 16, seed 17, gradient clipping at 1.0, endpoint label smoothing 0.02, and 0.15 weights for each masked filler loss. Five continuation epochs were run; continuation epoch 1 was retained by the pre-registered average-precision criterion. Later epochs overfit the same validation split.

This was an adaptive development process. Earlier checkpoints and a higher- learning-rate run had already been inspected on the same 326 validation rows; the final threshold was also selected there under an FPR <= 0.02 rule. Reported intervals condition on fixed selected models and thresholds. They do not remove selection or calibration optimism.

Development evaluation

IID split

Both rows below use the same 326 validation examples (160 END, 166 HOLD), with each threshold selected on that same split.

Model AP AUROC Brier Threshold TP / FP / TN / FN FPR END recall
Acoustic logistic baseline 0.7310 0.7522 0.2038 0.7639 21 / 3 / 163 / 139 0.0181 0.1313
Packaged TinyTCN 0.6881 0.7390 0.2133 0.7410 10 / 3 / 163 / 150 0.0181 0.0625

For the TinyTCN, the bootstrap 95% interval is [0, 0.0407] for FPR and [0.0261, 0.1019] for END recall. Three observed false positives therefore do not establish a population 2% interruption bound.

Paired TinyTCN-minus-baseline bootstrap differences are:

Metric Difference 95% interval
Average precision βˆ’0.0429 [βˆ’0.1195, 0.0352]
AUROC βˆ’0.0132 [βˆ’0.0782, 0.0512]
END recall at fixed selected thresholds βˆ’0.0688 [βˆ’0.1330, βˆ’0.0122]
FPR at fixed selected thresholds 0.0000 [βˆ’0.0284, 0.0298]

The baseline is the current local winner. The neural candidate's ranking deltas are uncertain, while its constrained-recall interval is below the baseline's. Because thresholds and checkpoints were selected on this development set, this comparison is descriptive rather than independent confirmatory evidence.

The Hindi-tagged validation slice has only 16 examples (8 END, 8 HOLD). The TinyTCN recovered 1/8 END examples at the global threshold. Its favorable ranking metrics on that tiny slice are unstable and are not evidence of Hinglish behavior.

Source-held-out stress split

Model AP AUROC Brier TP / FP / TN / FN FPR END recall
Acoustic logistic 0.5368 0.5539 0.3265 9 / 6 / 319 / 314 0.0185 0.0279
TinyTCN 0.5500 0.5523 0.2820 13 / 6 / 319 / 310 0.0185 0.0402

The paired AP difference is +0.0131 with a nine-source bootstrap 95% interval of [βˆ’0.2149, 0.1261]; AUROC difference is βˆ’0.0016 [βˆ’0.2713, 0.1273]; recall difference is +0.0124 [βˆ’0.0112, 0.0948]. Both models' ranking is near chance under this confounded joint shift, and the intervals are too wide to establish a winner.

Trailing-silence sensitivity

The frozen TinyTCN was rescored on the same 326 examples after appending valid zero-valued silence before the four-second suffix crop. The threshold stayed fixed.

Appended silence AP AUROC FPR END recall Decision flips vs 0 ms
0 ms 0.6881 0.7390 0.0181 0.0625 β€”
200 ms 0.6840 0.7403 0.0301 0.0688 5 / 326
400 ms 0.6979 0.7428 0.0301 0.0813 13 / 326
800 ms 0.7463 0.7672 0.0361 0.1438 26 / 326

At 800 ms the mean absolute probability change is 0.1058 and four false interruptions are introduced while one is resolved. Silence duration also enters the external controller, so an uncorrected deployment can double-count silence. The next model iteration should jitter trailing silence or normalize the speech endpoint and then recalibrate the controller on real pause sequences.

Clip-normalized false interruptions per audio hour are available in the report, but genuine turn/conversation IDs are absent. No online interruptions-per-turn, endpoint-delay, debounce, or repeated-END result is claimed.

Runtime and export

The FP32 ONNX export has a PyTorch/ONNX maximum absolute error of 0.0 on the parity check. Benchmarks use batch 1, one CPU thread, 20 warmups, and 200 measured iterations on the recorded Apple-arm machine:

Scope Runtime p50 p95 p99 Load Peak RSS
Neural only, 400 log-mel frames PyTorch 2.11 2.855 ms 2.949 ms 3.038 ms 4.303 ms 224.8 MB
Neural only, 400 log-mel frames ONNX Runtime 1.26 0.919 ms 0.956 ms 0.974 ms 6.041 ms 67.0 MB
4 s waveform to probability ONNX Runtime 1.26 1.572 ms 1.629 ms 1.651 ms 5.814 ms 78.2 MB

These measurements exclude VAD, microphone/network transport, the controller's intentional waiting time, and a production concurrency workload. They are not a device-independent SLA. No INT8 model is published; quantization would require new calibration and a new threshold.

The model is distributed with adjacent model_metadata.json and export_manifest.json. Those files bind preprocessing, I/O names and shapes, probability semantics, threshold, parameter count, checkpoint/config/split hashes, executable-source hashes, and parity. Ignoring them can silently invalidate results.

What remains unmeasured

  • the other 82 training shards and a full-corpus neural run;
  • an untouched neural model-selection holdout;
  • natural speaker-disjoint Hinglish minimal pairs;
  • the 31,527-row pinned official test;
  • real multi-checkpoint controller replay and endpoint-delay distributions;
  • device, codec, noise, regional, gender, disability, and speech-rate coverage;
  • FP32/INT8 decision parity and calibrated INT8 performance;
  • production concurrency, sustained load, and target-hardware latency.

No model-index metrics are provided because there is no release-quality test result.

Limitations and risk controls

  • False END predictions interrupt users; calibrate and monitor on the actual deployment population rather than copying the preview threshold.
  • False HOLD predictions are frequent at the strict preview threshold and add latency until controller relaxation or timeout.
  • Appended silence materially changes model scores; coordinate audio framing with the silence-aware controller.
  • Acoustic-only classification can miss lexical incompleteness that an ASR-aware system catches.
  • The inspected data is mostly synthetic and may expose generator or collection shortcuts.
  • Source, language, and human/synthetic status are confounded in the stress split.
  • A clip classifier cannot establish stateful conversational behavior.
  • Calibration can shift by language, device, noise, codec, quantization, and VAD policy.

Production monitoring should prefer anonymous turn/checkpoint IDs, score, threshold, silence, decision reason, and latency over raw speech. Raw audio or transcripts require a documented purpose, access controls, retention, deletion, and PII handling.

License, privacy, and provenance

Repository-authored source is Apache-2.0. The inspected upstream dataset card did not state an explicit dataset license. The license: other metadata and this card do not grant rights to upstream audio, speaker likeness, customer speech, local recordings, or the preview weights. Derived-weight redistribution and commercial-use rights remain unresolved pending upstream and legal/privacy review.

The repository does not redistribute upstream audio. The local collection protocol prohibits real customer names, phone numbers, addresses, order IDs, and support-call audio and requires informed consent and deletion controls.

Reproducibility

The repository includes the pinned data revision, audit and split reports, resolved training config, checkpoint and ONNX hashes, paired comparison reports, failure aggregates, silence perturbation report, exact environment files, CPU benchmarks, and guarded release tooling. Per-example validation predictions and the text-bearing split manifest remain in the private local workspace; the public package contains aggregate reports and hashes, not those sensitive rows. The quality gate completed with Ruff 0.15.16 passing and 136/136 automated tests passing via Python unittest discovery. The configured Hugging Face owner is suvradeepp.

See REPORT.md for the experiment narrative and DATA_CARD.md for data scope and governance.