tiny-turn-detector / README.md
Yash-V1002's picture
Deploy Tiny Turn Detector ZeroGPU
824cd37 verified
|
Raw
History Blame Contribute Delete
11 kB

A newer version of the Gradio SDK is available: 6.26.0

Upgrade
metadata
title: Tiny Turn Detector
emoji: πŸŽ™οΈ
colorFrom: blue
colorTo: purple
sdk: gradio
app_file: app.py
pinned: false
python_version: 3.12.12

Tiny Turn Detector

A small, audio-based turn detection model for conversational voice AI β€” built for the Shiprocket Data Scientist assessment.

Project

This project builds a tiny, audio-native model that decides whether a speaker has finished their turn (END) or is mid-pause / hesitating and likely to continue (CONTINUE). The chosen final architecture is a frozen Whisper Tiny encoder, mean-pooled, feeding a small Logistic Regression classifier β€” arrived at through a documented ladder of experiments (energy/silence baseline β†’ classical acoustic features β†’ temporal features β†’ Whisper Tiny), not assumed from the outset. See docs/RESULTS.md for the full evidence trail and experiments/EXPERIMENTS.md for the complete experiment log.

Problem

Turn detection is different from voice activity detection (VAD). VAD answers "is there speech energy right now?" β€” a memoryless, frame-level question. Turn detection answers "given everything heard so far, has the speaker's intent shifted to yielding the floor?" This requires distinguishing a mid-utterance pause (silence, but the speaker intends to continue β€” hesitation, a filler word, a connective word like "and"/ "but") from a true end-of-turn silence. Getting this wrong in either direction has a real cost in a live voice agent: ending a turn too early interrupts the user (premature END); waiting too long makes the agent feel slow and unresponsive (delayed END / false CONTINUE). Both error types are measured and reported separately throughout this project.

Dataset

pipecat-ai/smart-turn-data-v3.2-train — the real training dataset behind Pipecat's open-source Smart Turn v3.2 model (270,946 rows, 41.4GB, Whisper Tiny + linear classifier reference architecture). Schema, label semantics (endpoint_bool=true→END, false→CONTINUE), and provenance were confirmed directly from the dataset's own README and the upstream project's documentation — see docs/INITIAL_ANALYSIS.md.

The official test set, pipecat-ai/smart-turn-data-v3.2-test, was reserved throughout this project and was never used for tuning, model selection, or any experiment reported here. All development and validation splits used in this project were carved from the training distribution β€” this is a deliberate, stated methodological choice, not an oversight, and it means no result in this repository represents true held-out-test performance in the strictest sense.

Our small development/validation samples (90–300 clips) do not represent the complete 270,946-row dataset. Every result in this project states its exact sample size and source. See docs/LANGUAGE_ANALYSIS.md for a Known/Inferred/Not-measurable breakdown of what could and couldn't be determined about language coverage (including Hindi/Hinglish) from the available metadata.

Approach

  1. 16kHz mono audio normalization β€” matches the dataset's native format (confirmed uniform across every clip inspected) and Whisper's expected input.
  2. Whisper Tiny frozen encoder β€” chosen based on measured evidence (EXP-004), not assumed from the assessment brief's suggestion alone.
  3. Mean pooling over the encoder's time axis, producing a fixed 384-dimensional representation per clip.
  4. Logistic Regression β€” the smallest classifier that performed well in testing (92 parameters, ~2.9KB).
  5. Threshold / debounce / hysteresis decision layer β€” prevents committing to END on a single uncertain pause; configurable, not a second trained model. See docs/architecture.md.

Experiments

Experiment F1 Key finding
Energy baseline 0.400 Weak
Acoustic features (Logistic Regression) 0.575 Strong improvement over energy baseline
Temporal acoustic features +0.112 (delta) Recent context matters
Whisper Tiny + Logistic Regression 0.693 Strongest measured approach

Phase 3 (acoustic experiments) and Phase 4 (Whisper experiment) used different development samples β€” same dataset, same general sampling methodology, independent draws, not a paired comparison. The +0.118 F1 difference between the acoustic and Whisper approaches is best described as a directional improvement on independent samples drawn from the same Smart Turn training distribution, not a controlled paired experiment. Full numbers, sample sizes, and this caveat repeated in context: docs/RESULTS.md.

Filler analysis

The acoustic-only model's error analysis (Phase 3) found that 73% of its false-END errors had filler metadata present (midfiller or endfiller = True) β€” its dominant failure mode was exactly the assessment brief's named hard case: a pause that acoustically resembles an ending but linguistically isn't, because it's preceded by a filler word. The Whisper-based model's filler-flagged-clip F1 (0.722, n=35) was notably higher than the acoustic model's (0.529, n=39) on this specific slice, and its false-END errors were less filler-associated (56% vs. 73%). This filler-subset result should be treated as directional due to the modest sample sizes involved (35–39 clips) and the fact that it compares different samples β€” see docs/ERROR_ANALYSIS.md for the full, unglossed breakdown, including what was explicitly too small to trust (the no_filler_known slice, n=7).

Latency

~15.6ms per clip, measured on a Colab T4 GPU (preprocessing + Whisper encoder forward pass + Logistic Regression classifier, warm calls). This is hardware-dependent β€” the acoustic-only baseline's 12.2ms figure was CPU-measured, so the two are not directly comparable. No CPU-measured Whisper latency exists yet for this project; do not assume the GPU figure transfers to CPU deployment.

Limitations

  • Development/validation sample sizes are small (75–300 clips per experiment) β€” not enough for tight statistical confidence.
  • Phase 3 (acoustic) and Phase 4 (Whisper) experiments used different samples β€” not a paired, controlled comparison.
  • No transcripts exist anywhere in the dataset (spoken_text is null for every row) β€” no claim in this project relies on knowing what was actually said.
  • No true conversation-level endpoint timestamps exist in the dataset β€” this project measures classification accuracy and false-END/ false-CONTINUE rates, not true wall-clock endpoint latency.
  • Hindi/Hinglish code-switching could not be directly verified from dataset metadata (no transcripts, no code-switch label) β€” this project does not claim Hinglish robustness.
  • GPU-measured Whisper latency differs from CPU deployment latency, which hasn't been separately measured.
  • The official test set was never touched β€” no number here represents true held-out-test performance in the strictest sense.

Future work

  • Streaming incremental inference (avoid full-buffer recomputation on every triggering event)
  • Training/evaluating on a much larger sample of the full training set
  • A controlled Hinglish challenge set with ASR-verified code-switched content
  • Final evaluation on the official held-out test set, once the architecture is fully locked

Project structure

turn-detection/
β”œβ”€β”€ app.py                          # Gradio demo
β”œβ”€β”€ README.md
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ .gitignore
β”‚
β”œβ”€β”€ src/turn_detector/
β”‚   β”œβ”€β”€ __init__.py
β”‚   β”œβ”€β”€ inference.py                 # TurnDetector β€” final inference module
β”‚   β”œβ”€β”€ audio_io.py                  # audio loading/resampling (soundfile or ffmpeg fallback)
β”‚   β”œβ”€β”€ features.py                  # classical acoustic features (EXP-001/002/003b)
β”‚   β”œβ”€β”€ baseline_energy.py           # EXP-001
β”‚   β”œβ”€β”€ baseline_classifier.py       # EXP-002
β”‚   β”œβ”€β”€ splits.py                    # random / source-aware split utilities
β”‚   β”œβ”€β”€ evaluation.py                # shared metrics
β”‚   β”œβ”€β”€ data.py                      # HF streaming + stratified sampling
β”‚   └── parquet_reader.py            # pure-Python Parquet reader (built due to sandbox network limits)
β”‚
β”œβ”€β”€ models/
β”‚   β”œβ”€β”€ whisper_classifier.joblib    # trained Logistic Regression head (real, from EXP-004)
β”‚   └── model_metadata.json
β”‚
β”œβ”€β”€ notebooks/
β”‚   └── EXP004_whisper_baseline.ipynb  # self-contained Colab notebook (Whisper Tiny run)
β”‚
β”œβ”€β”€ scripts/
β”‚   β”œβ”€β”€ inspect_dataset.py
β”‚   β”œβ”€β”€ create_dev_subset.py
β”‚   β”œβ”€β”€ phase3_inventory_and_sample.py
β”‚   └── phase3_run_baselines.py
β”‚
β”œβ”€β”€ experiments/
β”‚   └── EXPERIMENTS.md               # full experiment log
β”‚
β”œβ”€β”€ artifacts/exp004/                # real EXP-004 outputs (results, embeddings, sample metadata)
β”‚
β”œβ”€β”€ docs/
β”‚   β”œβ”€β”€ INITIAL_ANALYSIS.md
β”‚   β”œβ”€β”€ LANGUAGE_ANALYSIS.md
β”‚   β”œβ”€β”€ PHASE2_REPORT.md
β”‚   β”œβ”€β”€ PHASE3_REAL_AUDIO_VALIDATION.md
β”‚   β”œβ”€β”€ ERROR_ANALYSIS.md
β”‚   β”œβ”€β”€ RESULTS.md
β”‚   └── architecture.md
β”‚
└── tests/
    β”œβ”€β”€ test_features_and_baselines.py   # 31 tests, synthetic audio
    └── test_inference.py                # 20 tests, real classifier + real embeddings

Whisper weights

Whisper Tiny's weights (openai/whisper-tiny) are not committed to this repository β€” they're loaded at runtime via transformers.WhisperModel.from_pretrained("openai/whisper-tiny"), which downloads and caches them (~151MB) on first use. This requires network access to Hugging Face. Only the trained Logistic Regression head (models/whisper_classifier.joblib, ~14KB) is committed, since that's the part actually trained in this project.

Running the demo

Hugging Face Spaces

This project is configured for a Gradio ZeroGPU Space. The GPU is allocated only when the predict_turn inference callback runs; the Whisper Tiny model and classifier are used for real inference. The Space does not require a paid dedicated GPU tier.

For local/cloud environments:

pip install -r requirements.txt
python app.py

Whisper Tiny weights are downloaded from Hugging Face on first inference.

Running tests

python -m pytest tests/ -v

51 tests total: 31 against synthetic audio (feature-extraction math, baseline logic), 20 against the real trained classifier and real Whisper embeddings from the actual EXP-004 Colab run (audio validation, decision logic, and end-to-end classifier-stage prediction) β€” the Whisper encoder stage itself is not testable in this project's own dev sandbox (see above), and tests that would require it fail informatively rather than being silently skipped or faked.