# Phase 3: Real-Audio Validation Report ## How real audio got into this sandbox The user uploaded a real Parquet shard: `/mnt/user-data/uploads/train-00000-of-00010.parquet` (464 MB, confirmed via `du`), one of ten training shards of `pipecat-ai/smart-turn-data-v3.2-train`. Verified as a genuine Parquet file by its magic bytes (`PAR1` at both start and end) before doing anything else. **No `pyarrow`, `fastparquet`, or `datasets` library is installed in this sandbox, and none are installable** — `pip`/`uv pip install pyarrow` fails with the same `403 host_not_allowed` block that covers `huggingface.co` (confirmed again this phase: `pypi.org` returns the identical proxy denial). So reading this real file required writing a **pure-Python Parquet reader from scratch**: `src/turn_detector/parquet_reader.py`. This is the single biggest piece of new engineering this phase, and it's worth documenting honestly, including the real bugs it had and how they were caught — because "trust but verify" matters more than usual when the verification code and the code under test were written by the same process in the same session. ## What was built, and the real bugs caught while building it 1. **Thrift compact-protocol decoder** (generic struct/list/map/binary reader) — needed because Parquet's footer metadata is Thrift-encoded. 2. **Footer parser** — extracts schema, row groups, column chunk metadata (codec, sizes, offsets) from the file's trailing metadata block. **First real result**: this shard has **3,153 rows across 32 row groups**, not the ~27,000 I'd have guessed by linearly scaling 270,946/10 — a good reminder not to estimate when you can just read the actual footer. 3. **Hand-written Snappy decompressor** (raw block format) — the shard's columns are all SNAPPY-compressed (confirmed from the footer, not assumed). Verified correct via round-trip: decompressed sizes matched the footer's declared `uncompressed_size` exactly for every page decoded. 4. **RLE/bit-packed hybrid decoder** for definition levels (needed because every column here is `OPTIONAL`/nullable, including scalars). - **Real bug #1, caught and fixed**: my first version hardcoded a 1-bit definition-level width for every column. This happened to work for top-level scalar columns (`language`, `endpoint_bool`, etc.) but silently corrupted `audio.bytes`, which is nested two levels deep inside the `OPTIONAL` `audio` struct and therefore needs a 2-bit definition level (0 = audio missing, 1 = audio present but bytes null, 2 = value present). The symptom was concrete and easy to catch: every decoded `audio.bytes` value came back `None`, even though `language` for the same rows decoded correctly. Fixed by computing the true max-definition-level per column from a real schema-tree walk (`_build_max_def_levels`), not an assumption. - **Real bug #2, caught and fixed**: dictionary-page decoding only handled `BYTE_ARRAY` values; `spoken_text` (typed `INT32`, entirely null) hit a dictionary page and raised `ValueError: Unsupported dictionary value type: INT32`. Fixed by adding an INT32 dictionary decode path. Re-verified afterward: `spoken_text` decodes cleanly to 100% `None` across the row group tested — consistent with the schema-level `dtype: null` fact already confirmed in Phase 1, now independently re-confirmed by actually decoding real values, not just trusting the declared schema. 5. **FLAC STREAMINFO parser** (`parse_flac_streaminfo`) — once `audio.bytes` decoded correctly, the first 4 bytes were `fLaC`: **the audio is FLAC, not raw WAV.** Rather than write a full FLAC decoder (LPC prediction + Rice coding — a much bigger undertaking), I parse just the 34-byte `STREAMINFO` metadata block to get sample rate, channel count, and total samples (hence duration) without decoding any audio — cheap enough to run across the entire 3,153-row shard in a few seconds. Cross-checked against `ffmpeg`'s own reported duration for a sample file: **exact match** (7.62775s both ways). 6. **`ffmpeg`-based audio decoding** for actual PCM samples (real decoding, not header-only): `ffmpeg` is already installed in this sandbox (unlike `soundfile`), so `src/turn_detector/audio_io.py` gained an `ffmpeg` subprocess fallback path, used transparently when `soundfile` is unavailable. Verified end-to-end: FLAC bytes → `ffmpeg` → WAV → loaded via the same fallback path → duration matches the STREAMINFO-derived duration exactly. Every one of these was tested against the real file at each step, not just written and assumed correct — see the tool-call history in this conversation for the actual intermediate outputs (row counts, sample values, byte-for-byte size checks) that caught bugs #1 and #2 above. ## Step 1: Full-shard inventory (real, all 3,153 rows) Computed by fully decoding every scalar metadata column plus STREAMINFO-derived duration for **every row in the shard** (not a sample — this was cheap enough, ~5s, to just do exhaustively): | Property | Value | |---|---| | Total rows | 3,153 | | Corrupted/unreadable rows | **0** | | `endpoint_bool` | 1,589 True (END) / 1,564 False (CONTINUE) — 50.4%/49.6%, matching the dataset authors' stated 50:50 design target almost exactly | | Sample rate | 16,000 Hz for all 3,153 rows (uniform) | | Channels | 1 (mono) for all 3,153 rows (uniform) | | Duration | min 0.36s, max 30.0s, mean 7.73s, median 7.32s | | `synthetic` | 2,575 True (82%) / 578 False (18%) | | `midfiller` | 1,293 True / 1,187 False / **673 None** (null ≠ false — unavailable for those rows) | | `endfiller` | 810 True / 1,670 False / **673 None** | **Language distribution (real, full shard, 23 languages present):** English dominates (789, 25%); Hindi (`hin`) = **132 rows (4.2%)** — a real measurement, not the ~7.6% rough estimate from Phase 1's 79-row hand sample (which was too small and non-random to trust for this — this is exactly why that estimate was labeled provisional at the time). **Source (`dataset` column) distribution — 12 distinct values, all now seen** (Phase 1's 79-row sample only ever saw 6 of these): `chirp3_1` (1,617), `chirp3_2` (852), `liva_1` (386), `midcentury_1` (108), `mundo_1` (40), `rime_2` (39), `human_5` (35), `orpheus_grammar_1` (24), `orpheus_endfiller_1` (19), `orpheus_midfiller_1` (13), `chirp3_3_short` (11), `human_convcollector_1` (9). The `orpheus_*` and `human_*` sources are new information this phase — not visible in Phase 1's smaller sample. ## Step 2: Verify label mapping Confirmed the same way documented in Phase 1 (upstream contribution guide: `endpoint_bool=true`→END, `false`→CONTINUE), plus a real smoke test on 60 real clips this phase: mean trailing silence for `endpoint_bool=True` clips was 0.300s vs. 0.114s for `endpoint_bool=False` clips — directionally consistent with the documented semantics (not proof, but nothing suggests an inverted or meaningless label). ## Step 3: Development sample materialized Stratified reservoir sample (seed=42, `src/turn_detector/data.py`), target 300, actual 300, stratified across `endpoint_bool` × `language` × `dataset` × `synthetic` × `midfiller` × `endfiller` × duration bucket. Materialized as real 16kHz mono WAV files at `data/raw/phase3_sample/audio/.wav` + `data/raw/phase3_sample/metadata.csv`. **300/300 clips converted successfully, 0 ffmpeg failures.** ## Steps 4-6: EXP-001 / EXP-002 / EXP-003(b) results All **SMALL REAL-AUDIO VALIDATION** — 210 dev / 90 val clips (random 70/30 split, seed=42), threshold/model tuned on dev only, all metrics below on held-out val: | | EXP-001 (energy/silence) | EXP-002 (classical features + LR) | EXP-003b Δ (temporal features) | |---|---|---|---| | Accuracy | 0.567 | 0.622 | global-only 0.511 → both 0.622 | | F1 | 0.400 | 0.575 | **+0.112** | | False END rate | 0.309 | 0.400 | **−0.109** (improved) | | False CONTINUE rate | 0.629 | 0.343 | **−0.114** (improved) | | Params / size | 2 thresholds | 92 params / 2,920 bytes | same classifier, feature-set ablation | | Latency (this machine) | negligible (rule-based) | mean 12.2ms / p95 22.4ms per clip | n/a | **EXP-003b's hypothesis is supported on this validation set**: adding recent-window (100–1000ms tail) features to the global whole-clip statistics improved F1 by +0.112 and reduced *both* error rates simultaneously — not just a precision/recall tradeoff in one direction. This is a real, if small-sample, result in favor of temporal features mattering for this task, consistent with what streaming-relevant intuition would predict. ## Step 7: Slice / filler analysis (real, on EXP-002's val predictions) - **By language**: only `eng` (n=22) met the `min_samples=10` bar in the 90-clip val set; every other language slice (including Hindi, n=3 in val) was too small and excluded rather than reported as if meaningful. `eng`-only: accuracy 0.636, F1 0.600. - **By synthetic flag**: non-synthetic (n=12): accuracy 0.583, F1 0.667. Synthetic (n=78): accuracy 0.628, F1 0.554. Interesting but n=12 for the non-synthetic slice is thin — flagged as exploratory, not conclusive. - **Filler metadata**: 72/90 val rows had `midfiller` populated (18 `None` — correctly excluded, not treated as "no filler"). `has_midfiller` (n=39): F1 0.529. `no_midfiller` (n=33): F1 0.593 — filler-present clips were somewhat harder for EXP-002, consistent with the qualitative error pattern below. ## Step 8: Error analysis (real false END / false CONTINUE cases) Full write-up: `docs/ERROR_ANALYSIS.md`. Headline, real, measured finding: **16 of 22 (73%) false-END errors have a filler flag present**, versus 50% for false-CONTINUE errors — and false-END clips have longer mean trailing silence (0.332s) than false-CONTINUE clips (0.217s). This is exactly the failure mode the assessment brief names as the hard case: a pause that acoustically resembles an ending but linguistically isn't (filler-preceded), which an audio-only, transcript-blind model has no way to catch beyond acoustic residue. ## Step 10/11: Decision **Question:** is the acoustic-only approach promising enough to continue, per the three-option framework? **Answer: Option B** — acoustic-only provides useful signal but is insufficient. Evidence: - EXP-002 (0.622 accuracy / 0.575 F1) clearly beats EXP-001 (0.567 / 0.400) and both clearly beat chance-level performance on a roughly balanced task, so acoustic features are picking up **real** signal — this isn't noise. - But 0.622 accuracy on a near-50/50 binary task is not close to production-usable on its own, and the error analysis identifies a specific, well-understood, real ceiling: **filler-word-preceded pauses are systematically harder**, and no acoustic-only feature set can fully resolve that without something closer to lexical/semantic content — the exact thing Whisper Tiny's speech representations (not just its transcription output) are trained to encode implicitly, which is the actual empirical justification the brief asked for, now grounded in a real measured failure mode rather than "the brief suggested it." - EXP-003b gives a concrete, real reason to believe **temporal/recency features help** (+0.112 F1), which is a design element worth carrying forward into whichever architecture comes next, Whisper-based or not. This is 90 validation clips from one of ten shards — a first real signal, not a final verdict. It's a directionally clear enough result to justify moving toward EXP-003 (Whisper Tiny) as the next architecture to actually try, while keeping EXP-002 as the benchmark it needs to beat. ## One recommended next experiment **EXP-003: Whisper Tiny encoder representations + the same lightweight classifier head, evaluated on the identical 210/90 dev/val split used here**, so the comparison against EXP-002's 0.575 F1 / 0.400 false-END rate / 0.343 false-CONTINUE rate is apples-to-apples. This directly tests whether Whisper's representations close the filler-word gap identified in error analysis, which is the specific, evidenced reason to expect it might help — not just "the brief suggested Whisper Tiny."