--- title: Tiny Turn Detector emoji: πŸŽ™οΈ colorFrom: blue colorTo: purple sdk: gradio app_file: app.py pinned: false python_version: "3.12.12" --- # Tiny Turn Detector A small, audio-based turn detection model for conversational voice AI β€” built for the Shiprocket Data Scientist assessment. ## Project This project builds a tiny, audio-native model that decides whether a speaker has finished their turn (END) or is mid-pause / hesitating and likely to continue (CONTINUE). The chosen final architecture is a frozen Whisper Tiny encoder, mean-pooled, feeding a small Logistic Regression classifier β€” arrived at through a documented ladder of experiments (energy/silence baseline β†’ classical acoustic features β†’ temporal features β†’ Whisper Tiny), not assumed from the outset. See `docs/RESULTS.md` for the full evidence trail and `experiments/EXPERIMENTS.md` for the complete experiment log. ## Problem Turn detection is different from voice activity detection (VAD). VAD answers "is there speech energy right now?" β€” a memoryless, frame-level question. Turn detection answers "given everything heard so far, has the speaker's intent shifted to yielding the floor?" This requires distinguishing a **mid-utterance pause** (silence, but the speaker intends to continue β€” hesitation, a filler word, a connective word like "and"/ "but") from a **true end-of-turn silence**. Getting this wrong in either direction has a real cost in a live voice agent: ending a turn too early interrupts the user (premature END); waiting too long makes the agent feel slow and unresponsive (delayed END / false CONTINUE). Both error types are measured and reported separately throughout this project. ## Dataset [`pipecat-ai/smart-turn-data-v3.2-train`](https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-train) β€” the real training dataset behind Pipecat's open-source Smart Turn v3.2 model (270,946 rows, 41.4GB, Whisper Tiny + linear classifier reference architecture). Schema, label semantics (`endpoint_bool=true`β†’END, `false`β†’CONTINUE), and provenance were confirmed directly from the dataset's own README and the upstream project's documentation β€” see `docs/INITIAL_ANALYSIS.md`. **The official test set, [`pipecat-ai/smart-turn-data-v3.2-test`](https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-test), was reserved throughout this project and was never used for tuning, model selection, or any experiment reported here.** All development and validation splits used in this project were carved from the training distribution β€” this is a deliberate, stated methodological choice, not an oversight, and it means no result in this repository represents true held-out-test performance in the strictest sense. **Our small development/validation samples (90–300 clips) do not represent the complete 270,946-row dataset.** Every result in this project states its exact sample size and source. See `docs/LANGUAGE_ANALYSIS.md` for a Known/Inferred/Not-measurable breakdown of what could and couldn't be determined about language coverage (including Hindi/Hinglish) from the available metadata. ## Approach 1. **16kHz mono audio normalization** β€” matches the dataset's native format (confirmed uniform across every clip inspected) and Whisper's expected input. 2. **Whisper Tiny frozen encoder** β€” chosen based on measured evidence (EXP-004), not assumed from the assessment brief's suggestion alone. 3. **Mean pooling** over the encoder's time axis, producing a fixed 384-dimensional representation per clip. 4. **Logistic Regression** β€” the smallest classifier that performed well in testing (92 parameters, ~2.9KB). 5. **Threshold / debounce / hysteresis decision layer** β€” prevents committing to END on a single uncertain pause; configurable, not a second trained model. See `docs/architecture.md`. ## Experiments | Experiment | F1 | Key finding | |---|---:|---| | Energy baseline | 0.400 | Weak | | Acoustic features (Logistic Regression) | 0.575 | Strong improvement over energy baseline | | Temporal acoustic features | +0.112 (delta) | Recent context matters | | Whisper Tiny + Logistic Regression | 0.693 | Strongest measured approach | **Phase 3 (acoustic experiments) and Phase 4 (Whisper experiment) used different development samples** β€” same dataset, same general sampling methodology, independent draws, not a paired comparison. The +0.118 F1 difference between the acoustic and Whisper approaches is best described as a **directional improvement on independent samples drawn from the same Smart Turn training distribution**, not a controlled paired experiment. Full numbers, sample sizes, and this caveat repeated in context: `docs/RESULTS.md`. ## Filler analysis The acoustic-only model's error analysis (Phase 3) found that **73% of its false-END errors had filler metadata present** (`midfiller` or `endfiller` = True) β€” its dominant failure mode was exactly the assessment brief's named hard case: a pause that acoustically resembles an ending but linguistically isn't, because it's preceded by a filler word. The Whisper-based model's filler-flagged-clip F1 (0.722, n=35) was notably higher than the acoustic model's (0.529, n=39) on this specific slice, and its false-END errors were less filler-associated (56% vs. 73%). **This filler-subset result should be treated as directional due to the modest sample sizes involved (35–39 clips) and the fact that it compares different samples** β€” see `docs/ERROR_ANALYSIS.md` for the full, unglossed breakdown, including what was explicitly too small to trust (the `no_filler_known` slice, n=7). ## Latency **~15.6ms per clip**, measured on a Colab T4 GPU (preprocessing + Whisper encoder forward pass + Logistic Regression classifier, warm calls). This is hardware-dependent β€” the acoustic-only baseline's 12.2ms figure was CPU-measured, so the two are not directly comparable. No CPU-measured Whisper latency exists yet for this project; do not assume the GPU figure transfers to CPU deployment. ## Limitations - Development/validation sample sizes are small (75–300 clips per experiment) β€” not enough for tight statistical confidence. - Phase 3 (acoustic) and Phase 4 (Whisper) experiments used different samples β€” not a paired, controlled comparison. - No transcripts exist anywhere in the dataset (`spoken_text` is null for every row) β€” no claim in this project relies on knowing what was actually said. - No true conversation-level endpoint timestamps exist in the dataset β€” this project measures classification accuracy and false-END/ false-CONTINUE rates, not true wall-clock endpoint latency. - Hindi/Hinglish code-switching could not be directly verified from dataset metadata (no transcripts, no code-switch label) β€” **this project does not claim Hinglish robustness.** - GPU-measured Whisper latency differs from CPU deployment latency, which hasn't been separately measured. - The official test set was never touched β€” no number here represents true held-out-test performance in the strictest sense. ## Future work - Streaming incremental inference (avoid full-buffer recomputation on every triggering event) - Training/evaluating on a much larger sample of the full training set - A controlled Hinglish challenge set with ASR-verified code-switched content - Final evaluation on the official held-out test set, once the architecture is fully locked --- ## Project structure ``` turn-detection/ β”œβ”€β”€ app.py # Gradio demo β”œβ”€β”€ README.md β”œβ”€β”€ requirements.txt β”œβ”€β”€ .gitignore β”‚ β”œβ”€β”€ src/turn_detector/ β”‚ β”œβ”€β”€ __init__.py β”‚ β”œβ”€β”€ inference.py # TurnDetector β€” final inference module β”‚ β”œβ”€β”€ audio_io.py # audio loading/resampling (soundfile or ffmpeg fallback) β”‚ β”œβ”€β”€ features.py # classical acoustic features (EXP-001/002/003b) β”‚ β”œβ”€β”€ baseline_energy.py # EXP-001 β”‚ β”œβ”€β”€ baseline_classifier.py # EXP-002 β”‚ β”œβ”€β”€ splits.py # random / source-aware split utilities β”‚ β”œβ”€β”€ evaluation.py # shared metrics β”‚ β”œβ”€β”€ data.py # HF streaming + stratified sampling β”‚ └── parquet_reader.py # pure-Python Parquet reader (built due to sandbox network limits) β”‚ β”œβ”€β”€ models/ β”‚ β”œβ”€β”€ whisper_classifier.joblib # trained Logistic Regression head (real, from EXP-004) β”‚ └── model_metadata.json β”‚ β”œβ”€β”€ notebooks/ β”‚ └── EXP004_whisper_baseline.ipynb # self-contained Colab notebook (Whisper Tiny run) β”‚ β”œβ”€β”€ scripts/ β”‚ β”œβ”€β”€ inspect_dataset.py β”‚ β”œβ”€β”€ create_dev_subset.py β”‚ β”œβ”€β”€ phase3_inventory_and_sample.py β”‚ └── phase3_run_baselines.py β”‚ β”œβ”€β”€ experiments/ β”‚ └── EXPERIMENTS.md # full experiment log β”‚ β”œβ”€β”€ artifacts/exp004/ # real EXP-004 outputs (results, embeddings, sample metadata) β”‚ β”œβ”€β”€ docs/ β”‚ β”œβ”€β”€ INITIAL_ANALYSIS.md β”‚ β”œβ”€β”€ LANGUAGE_ANALYSIS.md β”‚ β”œβ”€β”€ PHASE2_REPORT.md β”‚ β”œβ”€β”€ PHASE3_REAL_AUDIO_VALIDATION.md β”‚ β”œβ”€β”€ ERROR_ANALYSIS.md β”‚ β”œβ”€β”€ RESULTS.md β”‚ └── architecture.md β”‚ └── tests/ β”œβ”€β”€ test_features_and_baselines.py # 31 tests, synthetic audio └── test_inference.py # 20 tests, real classifier + real embeddings ``` ## Whisper weights Whisper Tiny's weights (`openai/whisper-tiny`) are **not** committed to this repository β€” they're loaded at runtime via `transformers.WhisperModel.from_pretrained("openai/whisper-tiny")`, which downloads and caches them (~151MB) on first use. This requires network access to Hugging Face. Only the trained Logistic Regression head (`models/whisper_classifier.joblib`, ~14KB) is committed, since that's the part actually trained in this project. ## Running the demo ### Hugging Face Spaces This project is configured for a **Gradio ZeroGPU Space**. The GPU is allocated only when the `predict_turn` inference callback runs; the Whisper Tiny model and classifier are used for real inference. The Space does not require a paid dedicated GPU tier. For local/cloud environments: ```bash pip install -r requirements.txt python app.py ``` Whisper Tiny weights are downloaded from Hugging Face on first inference. ## Running tests ```bash python -m pytest tests/ -v ``` 51 tests total: 31 against synthetic audio (feature-extraction math, baseline logic), 20 against the real trained classifier and real Whisper embeddings from the actual EXP-004 Colab run (audio validation, decision logic, and end-to-end classifier-stage prediction) β€” the Whisper encoder stage itself is not testable in this project's own dev sandbox (see above), and tests that would require it fail informatively rather than being silently skipped or faked.