Spaces:
Running on Zero
Running on Zero
File size: 7,851 Bytes
875e4af | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 | # Results
All results below are measured, not estimated. Sample sizes and sources
are stated for every number β see `experiments/EXPERIMENTS.md` for the
full experiment log with hypotheses and decisions.
## Summary table
| Experiment | F1 | Sample | Key finding |
|---|---:|---|---|
| EXP-001 Energy/silence baseline | 0.400 | 90-clip val (Phase 3, 300-clip sample, one training shard) | Weak β fixed silence-duration threshold can't distinguish filler-preceded pauses from real endpoints |
| EXP-002 Acoustic features + Logistic Regression | 0.575 | Same 90-clip val as EXP-001 | Strong improvement over EXP-001; still limited by lack of any lexical/semantic signal |
| EXP-003b Temporal (recent-window) acoustic features | +0.112 F1 delta vs. global-only | Same 90-clip val, feature-set ablation on EXP-002's classifier | Recent-window (100β1000ms tail) features meaningfully help; kept in the default feature set |
| EXP-004 Whisper Tiny frozen encoder + Logistic Regression | 0.693 | 75-clip val (fresh, independent 250-clip sample, same dataset) | Strongest measured approach; specifically reduces filler-associated false-END errors |
**On the EXP-002 vs. EXP-004 comparison**: EXP-002 and EXP-004 used
**different development samples** β both drawn from
`pipecat-ai/smart-turn-data-v3.2-train` using the same general
methodology (stratified sampling across `endpoint_bool`, language,
`dataset` source, `synthetic`), but not the same clips, and not a paired
before/after comparison on identical data. The correct way to describe the
+0.118 F1 difference is: **a directional improvement on independent
samples drawn from the same Smart Turn training distribution** β not a
controlled, paired experiment. This is stated explicitly wherever this
comparison appears in this project's documentation.
## Full metrics
### EXP-001 β Energy/silence baseline
- n=90 (validation split)
- Accuracy: 0.567
- F1: 0.400
- False-END rate: 0.309
- False-CONTINUE rate: 0.629
- Params: 2 (thresholds)
### EXP-002 β Acoustic features + Logistic Regression
- n=90 (same validation split as EXP-001)
- Accuracy: 0.622
- F1: 0.575
- False-END rate: 0.400
- False-CONTINUE rate: 0.343
- Params: 92 (91 features + intercept), ~2,920 bytes
- Inference latency (CPU, Claude sandbox): mean 12.2ms / p95 22.4ms per clip
### EXP-003b β Temporal (recent-window) features ablation
- Same 90-clip val set, same classifier family (Logistic Regression)
- Global-only: accuracy 0.511, F1 0.463, false-END rate 0.509
- Global + recent-window: accuracy 0.622, F1 0.575, false-END rate 0.400
- **Delta F1: +0.112**; delta false-END rate: β0.109; delta false-CONTINUE rate: β0.114
- Both error rates improved simultaneously β not a precision/recall tradeoff in one direction
### EXP-004 β Whisper Tiny frozen encoder + Logistic Regression
- n=75 (validation split from a fresh, independent 250-clip sample: 175 dev / 75 val, seed=42)
- Accuracy: 0.693
- F1: **0.693**
- False-END rate: **0.257**
- False-CONTINUE rate: 0.350
- Encoder params: 8,208,384 (~32.8MB fp32) β matches the confirmed reference Smart Turn architecture's documented ~8M params almost exactly
- Classifier params: 92 (~2,920 bytes)
- Inference latency (Colab T4 GPU): preprocessing ~13.5ms + encoder ~1.7ms + classifier ~0.45ms β **~15.6ms/clip total**
## Filler analysis
Phase 3's error analysis on EXP-002 found that **73% of false-END errors
had filler metadata present** (`midfiller` or `endfiller` = True), versus
50% of false-CONTINUE errors β meaning the acoustic-only model's dominant
failure mode was exactly the case the assessment brief names as hard: a
pause that acoustically resembles an ending but linguistically isn't
(filler-preceded).
Phase 4's EXP-004 evaluation targeted this directly:
| | Acoustic (EXP-002, Phase 3, n=39, `midfiller` only) | Whisper (EXP-004, Phase 4, n=35, `midfiller` OR `endfiller`) |
|---|---|---|
| F1 on filler-flagged clips | 0.529 | **0.722** |
And at the individual-error level: Whisper's false-END errors had a
filler flag present 56% of the time (5/9), down from the acoustic model's
73% β consistent with, though not definitive proof of, Whisper's
representation helping specifically on filler-associated pauses.
**This filler-subset result should be treated as directional, not
conclusive** β sample sizes are modest (35β39 clips), the two
measurements come from different samples, and the slice definitions
differ slightly (Phase 3 used `midfiller` alone, Phase 4 used
`midfiller` OR `endfiller`). The `no_filler_known` slice in Phase 4
(n=7) was too small to draw any conclusion from and is explicitly
excluded from this comparison β see `docs/ERROR_ANALYSIS.md`.
## Latency
**~15.6ms per clip**, measured on a Colab T4 GPU, for the full pipeline
(preprocessing + Whisper encoder forward pass + classifier). This is
**hardware-dependent** and should not be read as a CPU-deployment number
β the acoustic baseline's 12.2ms figure was measured on CPU, so the two
latency numbers are not directly comparable to each other. No CPU-measured
Whisper latency exists yet for this project.
## Limitations
- **Development sample sizes are small.** EXP-001/002/003b: 90-clip
validation set. EXP-004: 75-clip validation set. Neither approaches
the scale needed for tight confidence intervals on F1 or slice metrics.
- **EXP-002 and EXP-004 used different samples** β not a paired,
controlled comparison. See note above.
- **No transcripts exist anywhere in the dataset** (`spoken_text` is
null for every row, confirmed at the schema level) β no result in this
project relies on or claims access to what was actually said.
- **No true conversation-level endpoint timestamps exist** β this
project measures classification accuracy and false-END/false-CONTINUE
rates, not true wall-clock endpoint latency (time from actual turn
boundary to system decision). This is stated explicitly everywhere
latency is discussed.
- **Hindi/Hinglish code-switching cannot be directly verified from
dataset metadata** β Hindi is present as a language tag (`docs/LANGUAGE_ANALYSIS.md`
documents real, measured Hindi proportions: 4.2% of the full inspected
shard), but whether it's naturally code-switched Hindi-English speech or
monolingual Hindi (plausibly TTS-generated) was never independently
verified via ASR in this project (EXP-000, proposed in Phase 2, was
never run). **This project does not claim Hinglish robustness.**
- **GPU-measured Whisper latency differs from CPU deployment latency** β
a production CPU deployment would likely see meaningfully higher
encoder latency than the ~1.7ms GPU figure above; this hasn't been
measured and shouldn't be assumed.
- **No official test-set evaluation has been run.**
`pipecat-ai/smart-turn-data-v3.2-test` was reserved throughout this
project and was never used for tuning or evaluation β this is a
deliberate methodological choice, not an oversight, but it does mean
no number in this document represents true held-out generalization
performance in the strictest sense (the "held-out" splits used were
carved from the training distribution, not the official test set).
## Future work
- Streaming incremental inference (avoid full-buffer recomputation on
every VAD-triggered call β see `docs/INITIAL_ANALYSIS.md` Β§5 for the
design tension this raises with the reference Smart Turn architecture)
- Training/evaluating on a larger sample (the full 270,946-row training
set, or at least a much larger subset than 250β300 clips)
- A controlled Hinglish challenge set with verified code-switched content
(via ASR-based verification, EXP-000, never run in this project)
- Final evaluation on the official held-out test set
(`pipecat-ai/smart-turn-data-v3.2-test`), once the architecture is
fully locked
|