tiny-turn-detector / docs /PHASE2_REPORT.md
Yash-V1002's picture
Deploy Tiny Turn Detector
875e4af verified
|
Raw
History Blame Contribute Delete
8.69 kB

A newer version of the Gradio SDK is available: 6.26.0

Upgrade

Phase 2 Report

Dataset

  • Accessible rows in this sandbox: 0 of 270,946 (train) / 31,473 (test). Confirmed, tested, and documented blocker: bash network egress returns HTTP 403 (x-deny-reason: host_not_allowed) for huggingface.co, the datasets/huggingface_hub Python packages are not installed, and pip install cannot reach PyPI from this sandbox either. load_dataset(..., streaming=True) was attempted per the Phase 2 instructions (src/turn_detector/data.py::load_hf_dataset) and fails with a clear, caught DatasetAccessError rather than crashing or hanging β€” verified by actually running scripts/create_dev_subset.py, which exits cleanly with exit code 1 and a diagnostic message.
  • Development subset size: 0 (not yet created). scripts/create_dev_subset.py is complete and implements reproducible stratified sampling (seed-based, streaming-compatible, stratifies on endpoint_bool, language, dataset, synthetic, midfiller, endfiller, and a derived duration_bucket), but cannot run without dataset access.
  • Distributions: everything knowable without full access is already in docs/INITIAL_ANALYSIS.md and docs/LANGUAGE_ANALYSIS.md from the n=79 hand-parsed HF-viewer sample (not re-derived here to avoid duplication). No new distributional data was obtained this phase β€” further datasets-server API endpoints were attempted and blocked by this session's fetch-tool URL allowlist (see docs/LANGUAGE_ANALYSIS.md, top).

Baseline 1 (Energy/Silence)

  • Architecture: trailing-silence-duration rule against a tunable RMS threshold and duration threshold. Implemented in src/turn_detector/baseline_energy.py, including a tune_threshold() grid-search function that must be called with dev/val data only (the brief's "never tune on test" rule is stated in the docstring; not mechanically enforceable, so it's a process discipline, documented as such).
  • Metrics: not available. No real data to evaluate against.
  • Latency: not measured. evaluation.py::measure_inference_latency is implemented and generic (works on any predict_fn); running it against this baseline on real clips is a Phase 3 action item.
  • What IS verified: 5 unit tests against synthetic audio pass β€” correctly predicts END given long trailing silence, CONTINUE given none, and tune_threshold selects a sensible threshold and achieves >80% accuracy on cleanly-separable synthetic data (a sanity check on the mechanism, not a claim about real-world accuracy).

Baseline 2 (Classical features + Logistic Regression)

  • Features: RMS/ZCR/spectral centroid/bandwidth/rolloff/13 MFCCs (summary stats: mean/std/min/max each) + energy slope + trailing silence duration + duration, all implemented from scratch in numpy/scipy (src/turn_detector/features.py) since librosa isn't installed here and can't be installed without network access. This is a real engineering tradeoff, stated in the module docstring β€” not a compromise on feature quality, since MFCC math is standard (mel filterbank + DCT) and unit-tested directly (test_mfcc_shape and related tests verify finite, correctly-shaped output).
  • Classifier: sklearn Pipeline (median-impute β†’ standardize β†’ LogisticRegression, class-balanced). src/turn_detector/baseline_classifier.py.
  • Metrics: not available. No real data.
  • Latency/model size: not measured against real data, but the plumbing exists: param_count() and model_size_bytes() are implemented and unit-tested (return positive, sane values on a synthetic-data fit).
  • What IS verified: fits and predicts correctly on synthetic cleanly-separable data (>70% accuracy β€” expected for a sanity check, not a real-world number), exposes inspectable top-weighted features via top_features() for future error analysis.

Temporal feature experiment

  • Whether recent-window features help: not yet determined β€” needs real data.
  • What's implemented: extract_recent_window_features() computes RMS mean and slope over the last 100/250/500/750/1000ms of each clip, with an explicit <window>_available flag for clips shorter than a given window (never silently zero-padded, which would fabricate signal that isn't there). Unit-tested for correct availability-flagging on both short and long synthetic clips.
  • This is real, working infrastructure to answer the question the Phase 2 brief asks β€” the answer itself doesn't exist yet.

Source-aware evaluation

  • Random vs. source-aware split comparison: not yet run β€” needs real labeled data to compare accuracy on.
  • What's implemented and verified: src/turn_detector/splits.py provides split_random (may leave dataset-source overlap between train and val) and split_source_aware (group-disjoint by dataset source, verified by a passing unit test to produce zero source overlap). Both operate on lightweight metadata records, not audio, so they're cheap to run against a real dev subset once one exists.
  • The docstrings are explicit that this is a source-disjoint split, not a speaker-disjoint split (the dataset has no speaker IDs β€” see docs/INITIAL_ANALYSIS.md Β§4) β€” this distinction is preserved end-to-end rather than glossed over.

Language analysis

New document: docs/LANGUAGE_ANALYSIS.md. Structured exactly as requested β€” Known (schema facts + hand-counted n=79 sample facts, each sourced), Inferred (reasonable but unconfirmed readings, e.g. "Hindi content in this dataset may skew synthetic" based on 6/6 sampled Hindi rows being TTS-sourced β€” explicitly flagged as a weak, small-n signal), and Not measurable from available metadata (whether code-switching is actually present; the true full-scale language distribution). No claim that the dataset "is Hinglish" merely because Hindi exists β€” the document is explicit that this remains unverified pending EXP-000 (an ASR-based content check), which requires real audio access this sandbox doesn't yet have.

Error analysis

Not produced this phase. There is nothing to analyze β€” no model has been evaluated against real data, so there are no real misclassifications to inspect. docs/ERROR_ANALYSIS.md was deliberately not created with placeholder content; creating it now would either be empty or would invite fabricated-looking failure-mode claims, which the project rules explicitly prohibit. It will be created once EXP-001/EXP-002 actually run against real data.

Key conclusion

"How much can we achieve without Whisper?" β€” cannot be answered yet. This is stated plainly rather than filled in with a plausible-sounding guess. Both non-Whisper baselines are fully implemented, correctly functioning (verified via 31 passing unit tests against synthetic audio), and ready to run β€” but "ready to run" and "evaluated" are different things, and only the latter would justify a real conclusion about whether Whisper Tiny is empirically necessary. Presenting a confident answer here without real-data results would violate this project's most basic rule ("never fabricate metrics").

What Phase 2 does establish, honestly:

  • The dataset is well-understood at the schema/provenance level (Phase 1 + this phase's language analysis).
  • The full non-Whisper pipeline is built, modular, and correctness-tested.
  • The single blocking issue for getting real numbers is data access in this specific sandbox β€” not a design or methodology gap.

Recommended next experiment (exactly one)

Get a real, even if small, slice of actual audio into this sandbox, and run EXP-001 (energy/silence baseline) against it.γ€€This is deliberately chosen over jumping to EXP-002 or EXP-000: EXP-001 is the cheapest possible real signal about whether this whole pipeline behaves sensibly on real data, requires the least amount of real audio to be informative, and directly unblocks tuning a defensible baseline before any heavier comparison (classical classifier, Whisper) is worth attempting. Concretely, this needs one of:

  1. A person-provided upload of a small representative batch of real audio files (even a few hundred clips would be enough to run and sanity-check the full EXP-001 pipeline end-to-end), or
  2. Continued small-scale sampling of individual signed audio URLs via the HF dataset viewer (web_fetch), accepting the same small-sample caveats already documented in docs/LANGUAGE_ANALYSIS.md, or
  3. Running scripts/create_dev_subset.py and scripts/inspect_dataset.py in an environment with real network access to Hugging Face.