suvradeepp's picture
Publish Tiny Hinglish Turn Detector development preview
0348402 verified
|
Raw
History Blame Contribute Delete
2.36 kB

Tiny Hinglish Turn Detector — ONNX preview

This is a flat, dependency-light Kaggle Models bundle for audio-native HOLD/END decisions at VAD pause checkpoints. Upload every file in this folder as one ONNX model variation.

Development preview: the model was trained on one of 83 upstream training shards. The official test remains sealed, the acoustic logistic baseline is stronger on the current development split, and no verified Hinglish benchmark recordings have been evaluated. Do not claim production or Hinglish accuracy.

Kaggle model settings

  • Framework: ONNX
  • Suggested variation: tiny-tcn-fp32-preview
  • Fine-tunable: No
  • Visibility: Private until upstream-derived-weight redistribution rights have been reviewed
  • License: Other (specified in description). Apache-2.0 covers authored code, not the upstream data or derived-weight rights.

Files needed for inference

  • model.onnx: 151,812-parameter FP32 TinyTCN
  • model_metadata.json: frontend, tensor names, threshold, controller policy
  • turn_detector.py: standalone NumPy + ONNX Runtime inference
  • requirements.txt: three runtime dependencies

MODEL_CARD.md, DATA_CARD.md, development_metrics.json, and benchmark.json document the limited evidence. SHA256SUMS binds the payload.

Use inside a Kaggle Notebook

from pathlib import Path
import kagglehub

model_dir = Path(kagglehub.model_download(
    "YOUR_USERNAME/tiny-hinglish-turn-detector/onnx/tiny-tcn-fp32-preview"
))

import sys
sys.path.insert(0, str(model_dir))
from turn_detector import TurnDetector

detector = TurnDetector(model_dir)
result = detector.predict_file("/kaggle/input/your-audio/example.wav", silence_ms=300)
print(result)

For a downloaded folder outside Kaggle:

python -m pip install -r requirements.txt
python smoke_test.py
python example_inference.py path/to/audio.wav --silence-ms 300

Input audio may be mono or stereo and is resampled deterministically to 16 kHz. The model uses the most recent four seconds. p_end is compared with the serialized threshold and bounded by the serialized minimum/maximum silence policy. This helper makes one stateless checkpoint decision; production callers should retain the stateful controller from the full GitHub repository to latch END and prevent duplicate responses.