Spaces:
Running on Zero
Running on Zero
| # Experiment Log | |
| Rule for this file: a row only gets metrics once the corresponding script has | |
| actually been run against real data in this environment. Until then the | |
| `Status` column stays `NOT RUN` and the metric columns stay blank β never | |
| estimated, never "expected." | |
| **Context as of end of Phase 3:** the user uploaded a REAL Parquet shard | |
| (`train-00000-of-00010.parquet`, 464MB, 3,153 rows) from | |
| `pipecat-ai/smart-turn-data-v3.2-train`. No `pyarrow`/`fastparquet` are | |
| installed and none are installable (no network β same `host_not_allowed` | |
| block covers `pypi.org`), so a minimal pure-Python Parquet reader was | |
| built from scratch (`src/turn_detector/parquet_reader.py`: Thrift | |
| compact-protocol footer parser, hand-written Snappy decompressor, RLE/ | |
| bit-packed definition-level decoder) to actually read this file β see | |
| `docs/PHASE3_REAL_AUDIO_VALIDATION.md` for the full account, including | |
| real bugs caught and fixed while building it against this exact file | |
| (wrong definition-level bit-width for the nested `audio` struct; missing | |
| INT32 dictionary-page support). Audio is FLAC-encoded (`fLaC` magic, | |
| confirmed by direct byte inspection); `ffmpeg` (already present in this | |
| sandbox) is used for decoding instead of `soundfile` (not installed, | |
| not installable). | |
| **All 3,153 rows of this shard were successfully decoded β 0 corrupted | |
| rows.** A reproducible stratified sample of 300 real clips (seed=42) was | |
| drawn and materialized as real 16kHz mono WAV files | |
| (`data/raw/phase3_sample/`). EXP-001/002/003 below were run against a | |
| 90-clip validation split held out from those 300 (210 dev / 90 val, | |
| random split, seed=42) β tuning done on dev only. **Every result below is | |
| labeled `SMALL REAL-AUDIO VALIDATION` and must not be read as full-dataset | |
| or full-shard performance** β it's 90 held-out clips from one of ten | |
| shards. | |
| | ID | Name | Hypothesis | Status | Accuracy | F1 | Premature-END rate | Endpoint latency | Params | CPU latency | Decision | | |
| |----|------|-----------|--------|----------|-----|--------------------|--------------------|--------|--------------|----------| | |
| | EXP-000 | Audio content check (Hinglish verification) | The dataset has no transcripts and no code-switch label; running Whisper ASR over a sample of `hin`-tagged clips (vs. a comparison sample of `eng`-tagged clips) will reveal whether genuine Hindi-English code-switching is present, or whether `hin` rows are monolingual Hindi. This determines whether "Hinglish robustness" can be meaningfully trained/evaluated from this dataset at all. | NOT RUN β needs sampled audio in-sandbox first | | | | | | | | | |
| | EXP-001 | Energy/silence endpointing baseline | A fixed-threshold silence-duration rule (classic VAD-style endpointing) will catch many obvious ends-of-turn but will fail badly on filler-preceded pauses, given the label methodology explicitly defines "incomplete" as often ending in trailing silence *after* a filler word β acoustically similar to a genuine end. This gives a real lower-bound baseline and a first look at real premature-END behavior. | **SMALL REAL-AUDIO VALIDATION β RUN.** Threshold tuned on 210-clip dev set (best: silence_duration_threshold_sec=0.25, dev accuracy=0.638), evaluated on 90-clip held-out val set. | 0.567 | 0.400 | 0.309 | N/A (dataset lacks conversation-level timestamps) | 2 | not separately benchmarked (rule-based, negligible) | Confirmed weak, as hypothesized; kept as floor baseline | | |
| | EXP-002 | Classical audio features + lightweight classifier | Adding RMS/ZCR/MFCC/spectral stats/pause-duration/energy-slope features to a small classifier (e.g. logistic regression / small GBM) should beat EXP-001 by capturing trend information a fixed threshold can't, while remaining extremely cheap. | **SMALL REAL-AUDIO VALIDATION β RUN.** Logistic regression, median-imputeβstandardizeβLR pipeline, fit on 210-clip dev set, evaluated on same 90-clip val set as EXP-001. | 0.622 | 0.575 | 0.400 | N/A (dataset lacks conversation-level timestamps) | 92 (91 features + intercept), 2920 bytes | mean 12.2ms / p95 22.4ms per clip, this dev machine, single-clip calls, 90 repetitions | Beats EXP-001 on this val set; kept as current best acoustic baseline | | |
| | EXP-004 | Whisper Tiny frozen encoder + lightweight classifier | Whisper Tiny's encoder representations, trained on large-scale multilingual speech, should encode prosodic/phonetic cues (trailing intonation, breath, filler acoustics) that classical features miss β specifically reducing false-END errors on filler-associated pauses, the exact failure mode identified in Phase 3's error analysis. | **SMALL REAL-AUDIO VALIDATION β RUN (Colab, GPU).** Frozen `openai/whisper-tiny` encoder, mean-pooled over time, + logistic regression head. Fresh 250-clip HF-streaming sample (175 dev / 75 val, seed=42), independent of Phase 3's 300-clip sample β same dataset/method, not the same clips. | 0.693 | 0.693 | 0.257 | N/A (dataset lacks conversation-level timestamps) | 8,208,384 (encoder) + 152 (LR head) | preprocessing 13.5ms + encoder ~1.7ms + classifier 0.45ms β 15.6ms/clip, Colab T4 GPU (not directly comparable to EXP-002's CPU latency) | **Clear improvement over EXP-002 (ΞF1 +0.118, Ξfalse-END β0.143); filler-slice F1 0.722 vs. acoustic's 0.529 supports the hypothesis.** Adopted as the leading candidate architecture β see decision below. | | |
| | EXP-003b | Global-only vs. global+recent-window (temporal) features | The acoustic state near the end of the observed audio (last 100β1000ms) contains more useful turn-completion signal than global whole-clip statistics alone. | **SMALL REAL-AUDIO VALIDATION β RUN.** Same classifier (logistic regression), same 210/90 dev/val split; only the feature set changes (global-only vs. global+recent-window). | global-only: 0.511 / both: 0.622 | global-only: 0.463 / both: 0.575 (Ξ **+0.112**) | global-only: 0.509 / both: 0.400 (Ξ **β0.109**) | N/A | n/a (feature-set ablation, same classifier family) | not separately benchmarked | **Hypothesis supported on this val set**: adding recent-window features improved F1 by +0.112 and reduced false-END rate by 0.109 and false-CONTINUE rate by 0.114 β all three moved in the favorable direction simultaneously. Kept in the default feature set. | | |
| | EXP-004b | Source-aware vs. random split comparison | Because there are no speaker/session IDs, `dataset` (source/voice-engine) is the strongest available grouping variable; a random split may let a model learn source-specific acoustic shortcuts, making random-split evaluation look artificially better than source-aware evaluation. | **Code implemented & unit-tested (synthetic grouped records) β NOT yet run on the real shard.** `src/turn_detector/splits.py` verified correct; Phase 3/4's actual runs used a plain random dev/val split (not source-aware) for simplicity given the small samples involved β the random-vs-source-aware accuracy-gap comparison itself is still open. | | | | | | | Deferred | | |
| | EXP-005 | Hybrid acoustic + Whisper representation | Combining cheap classical features (fast, precise about energy/silence timing) with Whisper embeddings (semantic/prosodic richness) may outperform either alone, since they likely capture complementary error patterns β to be checked via error-mode overlap analysis, not assumed. | NOT RUN β depends on EXP-004 (Whisper, now run β see above), candidate next experiment. | | | | | | | Deferred to Phase 5 | | |
| | EXP-006 | Temporal decision logic (hysteresis/debounce) | Raw per-invocation P(end) will be noisy call-to-call (especially if re-invoked on a growing buffer as the reference design does); adding a smoothing/hysteresis rule should reduce false END flips without meaningfully increasing true-END latency. | NOT RUN β depends on having a working classifier evaluated on real streaming-style sequences first. | | | | | | | Deferred | | |
| | EXP-007 | Hinglish / filler robustness pass | Contingent on EXP-000: if genuine Hinglish content is confirmed present, targeted analysis + augmentation/reweighting on those cases should close part of the accuracy gap on code-switched input. Filler-label supervision (`midfiller`/`endfiller`) is confirmed only populated for a subset of data sources (`chirp3_1`/`chirp3_2` in the sample inspected), which constrains how this can be trained. | NOT RUN β blocked on EXP-000 and real dataset access. | | | | | | | Deferred | | |
| | EXP-008 | Latency/model-size optimization | The winning model from EXP-003β007 is likely oversized for the accuracy it needs; quantization/distillation/pruning should recover most of the accuracy at a fraction of the CPU latency and model size β directly comparable against the confirmed reference numbers (8MB int8 CPU version, 32MB fp32 GPU version, ~65ms Pipecat Cloud inference). | NOT RUN β depends on having a trained model first. | | | | | | | Deferred | | |
| **New in Phase 2 β a "temporal/recent-window features" sub-experiment** | |
| (part of EXP-002's scope per the Phase 2 brief Β§8): `features.py`'s | |
| `extract_recent_window_features()` computes RMS mean/slope over the last | |
| 100/250/500/750/1000ms of each clip, with explicit `*_available` flags for | |
| clips shorter than a given window (never silently zero-filled/padded). | |
| Unit-tested (`test_recent_window_features_*`, 2/2 pass) for correct | |
| availability-flagging behavior. Whether recent-window features actually | |
| outperform whole-clip statistics is an empirical question that needs real | |
| data β the code to test it exists and is verified correct, the finding | |
| does not exist yet. | |
| ## Notes | |
| - **Numbering note (updated Phase 4):** the temporal/recent-window feature | |
| comparison is labeled **EXP-003b** (not EXP-003, which is Whisper Tiny β | |
| now actually run and reported as **EXP-004**). The source-aware vs. | |
| random split comparison is labeled **EXP-004b** to avoid colliding with | |
| the Whisper row. Numbering diverges slightly from any single phase | |
| brief's own local numbering by design, to keep one consistent ID per | |
| experiment across the whole project. | |
| - EXP-000 is new since the first Phase-1 draft: it exists specifically | |
| because we confirmed the dataset has no transcripts and no code-switch | |
| label, so Hinglish coverage is currently unverified and must be checked | |
| empirically before EXP-007 can be meaningfully planned. Still not run β | |
| would need ASR (e.g. Whisper) over the Hindi-tagged clips in the real | |
| sample now available (`data/raw/phase3_sample/`), which is feasible next | |
| since real audio now exists in-sandbox, but wasn't part of this phase's | |
| scope. | |
| - Experiments will be run in order, but a later experiment may be skipped if | |
| an earlier one already provides conclusive evidence against it β the | |
| brief explicitly says not to assume the final architecture before seeing | |
| data. | |
| - "Premature-END rate" and "Endpoint latency" columns: per | |
| `docs/INITIAL_ANALYSIS.md` Β§7, this dataset does not support true | |
| wall-clock endpoint-latency measurement (no conversation-level | |
| timestamps). Populated rows use `false_end_rate` (premature-endpoint | |
| proxy, computable) and mark endpoint latency `N/A` as documented. | |
| - No row in this table will be edited to look better after the fact. If a | |
| result is worse than expected, that's the result, and it drives the next | |
| hypothesis. EXP-001's 0.567 accuracy / 0.400 F1 on real validation data | |
| is a weak result and is reported as such, not smoothed over. | |
| ## Status as of end of Phase 4 | |
| **EXP-004 (Whisper Tiny frozen encoder) run successfully via the generated Colab notebook** | |
| (`notebooks/EXP004_whisper_baseline.ipynb`), on a fresh 250-clip HF-streaming sample (75-clip val), | |
| independent of Phase 3's 300-clip sample. Real artifacts saved: `artifacts/exp004/`. | |
| **Decision (Phase 4, per the A/B/C framework): Option A β strong improvement.** ΞF1 = +0.118, | |
| false-END rate improved by β0.143, and the improvement concentrates specifically on the | |
| filler-associated failure mode Phase 3's error analysis identified (filler-slice F1 0.722 vs. | |
| acoustic's 0.529) β a coherent, multi-signal result, not a single-metric artifact. Caveated | |
| honestly: val n=75 is modest, the sample differs from Phase 3's exact clips (same dataset/method), | |
| and the `no_filler_known` slice (n=7) is too small to trust and is excluded from the "supports the | |
| hypothesis" claim. | |