# Results All results below are measured, not estimated. Sample sizes and sources are stated for every number — see `experiments/EXPERIMENTS.md` for the full experiment log with hypotheses and decisions. ## Summary table | Experiment | F1 | Sample | Key finding | |---|---:|---|---| | EXP-001 Energy/silence baseline | 0.400 | 90-clip val (Phase 3, 300-clip sample, one training shard) | Weak — fixed silence-duration threshold can't distinguish filler-preceded pauses from real endpoints | | EXP-002 Acoustic features + Logistic Regression | 0.575 | Same 90-clip val as EXP-001 | Strong improvement over EXP-001; still limited by lack of any lexical/semantic signal | | EXP-003b Temporal (recent-window) acoustic features | +0.112 F1 delta vs. global-only | Same 90-clip val, feature-set ablation on EXP-002's classifier | Recent-window (100–1000ms tail) features meaningfully help; kept in the default feature set | | EXP-004 Whisper Tiny frozen encoder + Logistic Regression | 0.693 | 75-clip val (fresh, independent 250-clip sample, same dataset) | Strongest measured approach; specifically reduces filler-associated false-END errors | **On the EXP-002 vs. EXP-004 comparison**: EXP-002 and EXP-004 used **different development samples** — both drawn from `pipecat-ai/smart-turn-data-v3.2-train` using the same general methodology (stratified sampling across `endpoint_bool`, language, `dataset` source, `synthetic`), but not the same clips, and not a paired before/after comparison on identical data. The correct way to describe the +0.118 F1 difference is: **a directional improvement on independent samples drawn from the same Smart Turn training distribution** — not a controlled, paired experiment. This is stated explicitly wherever this comparison appears in this project's documentation. ## Full metrics ### EXP-001 — Energy/silence baseline - n=90 (validation split) - Accuracy: 0.567 - F1: 0.400 - False-END rate: 0.309 - False-CONTINUE rate: 0.629 - Params: 2 (thresholds) ### EXP-002 — Acoustic features + Logistic Regression - n=90 (same validation split as EXP-001) - Accuracy: 0.622 - F1: 0.575 - False-END rate: 0.400 - False-CONTINUE rate: 0.343 - Params: 92 (91 features + intercept), ~2,920 bytes - Inference latency (CPU, Claude sandbox): mean 12.2ms / p95 22.4ms per clip ### EXP-003b — Temporal (recent-window) features ablation - Same 90-clip val set, same classifier family (Logistic Regression) - Global-only: accuracy 0.511, F1 0.463, false-END rate 0.509 - Global + recent-window: accuracy 0.622, F1 0.575, false-END rate 0.400 - **Delta F1: +0.112**; delta false-END rate: −0.109; delta false-CONTINUE rate: −0.114 - Both error rates improved simultaneously — not a precision/recall tradeoff in one direction ### EXP-004 — Whisper Tiny frozen encoder + Logistic Regression - n=75 (validation split from a fresh, independent 250-clip sample: 175 dev / 75 val, seed=42) - Accuracy: 0.693 - F1: **0.693** - False-END rate: **0.257** - False-CONTINUE rate: 0.350 - Encoder params: 8,208,384 (~32.8MB fp32) — matches the confirmed reference Smart Turn architecture's documented ~8M params almost exactly - Classifier params: 92 (~2,920 bytes) - Inference latency (Colab T4 GPU): preprocessing ~13.5ms + encoder ~1.7ms + classifier ~0.45ms ≈ **~15.6ms/clip total** ## Filler analysis Phase 3's error analysis on EXP-002 found that **73% of false-END errors had filler metadata present** (`midfiller` or `endfiller` = True), versus 50% of false-CONTINUE errors — meaning the acoustic-only model's dominant failure mode was exactly the case the assessment brief names as hard: a pause that acoustically resembles an ending but linguistically isn't (filler-preceded). Phase 4's EXP-004 evaluation targeted this directly: | | Acoustic (EXP-002, Phase 3, n=39, `midfiller` only) | Whisper (EXP-004, Phase 4, n=35, `midfiller` OR `endfiller`) | |---|---|---| | F1 on filler-flagged clips | 0.529 | **0.722** | And at the individual-error level: Whisper's false-END errors had a filler flag present 56% of the time (5/9), down from the acoustic model's 73% — consistent with, though not definitive proof of, Whisper's representation helping specifically on filler-associated pauses. **This filler-subset result should be treated as directional, not conclusive** — sample sizes are modest (35–39 clips), the two measurements come from different samples, and the slice definitions differ slightly (Phase 3 used `midfiller` alone, Phase 4 used `midfiller` OR `endfiller`). The `no_filler_known` slice in Phase 4 (n=7) was too small to draw any conclusion from and is explicitly excluded from this comparison — see `docs/ERROR_ANALYSIS.md`. ## Latency **~15.6ms per clip**, measured on a Colab T4 GPU, for the full pipeline (preprocessing + Whisper encoder forward pass + classifier). This is **hardware-dependent** and should not be read as a CPU-deployment number — the acoustic baseline's 12.2ms figure was measured on CPU, so the two latency numbers are not directly comparable to each other. No CPU-measured Whisper latency exists yet for this project. ## Limitations - **Development sample sizes are small.** EXP-001/002/003b: 90-clip validation set. EXP-004: 75-clip validation set. Neither approaches the scale needed for tight confidence intervals on F1 or slice metrics. - **EXP-002 and EXP-004 used different samples** — not a paired, controlled comparison. See note above. - **No transcripts exist anywhere in the dataset** (`spoken_text` is null for every row, confirmed at the schema level) — no result in this project relies on or claims access to what was actually said. - **No true conversation-level endpoint timestamps exist** — this project measures classification accuracy and false-END/false-CONTINUE rates, not true wall-clock endpoint latency (time from actual turn boundary to system decision). This is stated explicitly everywhere latency is discussed. - **Hindi/Hinglish code-switching cannot be directly verified from dataset metadata** — Hindi is present as a language tag (`docs/LANGUAGE_ANALYSIS.md` documents real, measured Hindi proportions: 4.2% of the full inspected shard), but whether it's naturally code-switched Hindi-English speech or monolingual Hindi (plausibly TTS-generated) was never independently verified via ASR in this project (EXP-000, proposed in Phase 2, was never run). **This project does not claim Hinglish robustness.** - **GPU-measured Whisper latency differs from CPU deployment latency** — a production CPU deployment would likely see meaningfully higher encoder latency than the ~1.7ms GPU figure above; this hasn't been measured and shouldn't be assumed. - **No official test-set evaluation has been run.** `pipecat-ai/smart-turn-data-v3.2-test` was reserved throughout this project and was never used for tuning or evaluation — this is a deliberate methodological choice, not an oversight, but it does mean no number in this document represents true held-out generalization performance in the strictest sense (the "held-out" splits used were carved from the training distribution, not the official test set). ## Future work - Streaming incremental inference (avoid full-buffer recomputation on every VAD-triggered call — see `docs/INITIAL_ANALYSIS.md` §5 for the design tension this raises with the reference Smart Turn architecture) - Training/evaluating on a larger sample (the full 270,946-row training set, or at least a much larger subset than 250–300 clips) - A controlled Hinglish challenge set with verified code-switched content (via ASR-based verification, EXP-000, never run in this project) - Final evaluation on the official held-out test set (`pipecat-ai/smart-turn-data-v3.2-test`), once the architecture is fully locked