File size: 12,324 Bytes
875e4af
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
# Phase 3: Real-Audio Validation Report

## How real audio got into this sandbox

The user uploaded a real Parquet shard:
`/mnt/user-data/uploads/train-00000-of-00010.parquet` (464 MB, confirmed
via `du`), one of ten training shards of
`pipecat-ai/smart-turn-data-v3.2-train`. Verified as a genuine Parquet
file by its magic bytes (`PAR1` at both start and end) before doing
anything else.

**No `pyarrow`, `fastparquet`, or `datasets` library is installed in this
sandbox, and none are installable** β€” `pip`/`uv pip install pyarrow`
fails with the same `403 host_not_allowed` block that covers
`huggingface.co` (confirmed again this phase: `pypi.org` returns the
identical proxy denial). So reading this real file required writing a
**pure-Python Parquet reader from scratch**:
`src/turn_detector/parquet_reader.py`. This is the single biggest piece of
new engineering this phase, and it's worth documenting honestly, including
the real bugs it had and how they were caught β€” because "trust but verify"
matters more than usual when the verification code and the code under test
were written by the same process in the same session.

## What was built, and the real bugs caught while building it

1. **Thrift compact-protocol decoder** (generic struct/list/map/binary
   reader) β€” needed because Parquet's footer metadata is Thrift-encoded.
2. **Footer parser** β€” extracts schema, row groups, column chunk metadata
   (codec, sizes, offsets) from the file's trailing metadata block.
   **First real result**: this shard has **3,153 rows across 32 row
   groups**, not the ~27,000 I'd have guessed by linearly scaling
   270,946/10 β€” a good reminder not to estimate when you can just read the
   actual footer.
3. **Hand-written Snappy decompressor** (raw block format) β€” the shard's
   columns are all SNAPPY-compressed (confirmed from the footer, not
   assumed). Verified correct via round-trip: decompressed sizes matched
   the footer's declared `uncompressed_size` exactly for every page
   decoded.
4. **RLE/bit-packed hybrid decoder** for definition levels (needed because
   every column here is `OPTIONAL`/nullable, including scalars).
   - **Real bug #1, caught and fixed**: my first version hardcoded a
     1-bit definition-level width for every column. This happened to work
     for top-level scalar columns (`language`, `endpoint_bool`, etc.) but
     silently corrupted `audio.bytes`, which is nested two levels deep
     inside the `OPTIONAL` `audio` struct and therefore needs a 2-bit
     definition level (0 = audio missing, 1 = audio present but bytes
     null, 2 = value present). The symptom was concrete and easy to
     catch: every decoded `audio.bytes` value came back `None`, even
     though `language` for the same rows decoded correctly. Fixed by
     computing the true max-definition-level per column from a real
     schema-tree walk (`_build_max_def_levels`), not an assumption.
   - **Real bug #2, caught and fixed**: dictionary-page decoding only
     handled `BYTE_ARRAY` values; `spoken_text` (typed `INT32`, entirely
     null) hit a dictionary page and raised `ValueError: Unsupported
     dictionary value type: INT32`. Fixed by adding an INT32 dictionary
     decode path. Re-verified afterward: `spoken_text` decodes cleanly to
     100% `None` across the row group tested β€” consistent with the
     schema-level `dtype: null` fact already confirmed in Phase 1, now
     independently re-confirmed by actually decoding real values, not
     just trusting the declared schema.
5. **FLAC STREAMINFO parser** (`parse_flac_streaminfo`) β€” once
   `audio.bytes` decoded correctly, the first 4 bytes were `fLaC`: **the
   audio is FLAC, not raw WAV.** Rather than write a full FLAC decoder
   (LPC prediction + Rice coding β€” a much bigger undertaking), I parse
   just the 34-byte `STREAMINFO` metadata block to get sample rate,
   channel count, and total samples (hence duration) without decoding any
   audio β€” cheap enough to run across the entire 3,153-row shard in a few
   seconds. Cross-checked against `ffmpeg`'s own reported duration for a
   sample file: **exact match** (7.62775s both ways).
6. **`ffmpeg`-based audio decoding** for actual PCM samples (real decoding,
   not header-only): `ffmpeg` is already installed in this sandbox (unlike
   `soundfile`), so `src/turn_detector/audio_io.py` gained an `ffmpeg`
   subprocess fallback path, used transparently when `soundfile` is
   unavailable. Verified end-to-end: FLAC bytes β†’ `ffmpeg` β†’ WAV β†’ loaded
   via the same fallback path β†’ duration matches the STREAMINFO-derived
   duration exactly.

Every one of these was tested against the real file at each step, not
just written and assumed correct β€” see the tool-call history in this
conversation for the actual intermediate outputs (row counts, sample
values, byte-for-byte size checks) that caught bugs #1 and #2 above.

## Step 1: Full-shard inventory (real, all 3,153 rows)

Computed by fully decoding every scalar metadata column plus
STREAMINFO-derived duration for **every row in the shard** (not a sample β€”
this was cheap enough, ~5s, to just do exhaustively):

| Property | Value |
|---|---|
| Total rows | 3,153 |
| Corrupted/unreadable rows | **0** |
| `endpoint_bool` | 1,589 True (END) / 1,564 False (CONTINUE) β€” 50.4%/49.6%, matching the dataset authors' stated 50:50 design target almost exactly |
| Sample rate | 16,000 Hz for all 3,153 rows (uniform) |
| Channels | 1 (mono) for all 3,153 rows (uniform) |
| Duration | min 0.36s, max 30.0s, mean 7.73s, median 7.32s |
| `synthetic` | 2,575 True (82%) / 578 False (18%) |
| `midfiller` | 1,293 True / 1,187 False / **673 None** (null β‰  false β€” unavailable for those rows) |
| `endfiller` | 810 True / 1,670 False / **673 None** |

**Language distribution (real, full shard, 23 languages present):**
English dominates (789, 25%); Hindi (`hin`) = **132 rows (4.2%)** β€” a real
measurement, not the ~7.6% rough estimate from Phase 1's 79-row hand
sample (which was too small and non-random to trust for this β€” this is
exactly why that estimate was labeled provisional at the time).

**Source (`dataset` column) distribution β€” 12 distinct values, all now
seen** (Phase 1's 79-row sample only ever saw 6 of these):
`chirp3_1` (1,617), `chirp3_2` (852), `liva_1` (386), `midcentury_1` (108),
`mundo_1` (40), `rime_2` (39), `human_5` (35), `orpheus_grammar_1` (24),
`orpheus_endfiller_1` (19), `orpheus_midfiller_1` (13),
`chirp3_3_short` (11), `human_convcollector_1` (9). The `orpheus_*` and
`human_*` sources are new information this phase β€” not visible in Phase 1's
smaller sample.

## Step 2: Verify label mapping

Confirmed the same way documented in Phase 1 (upstream contribution guide:
`endpoint_bool=true`β†’END, `false`β†’CONTINUE), plus a real smoke test on 60
real clips this phase: mean trailing silence for `endpoint_bool=True`
clips was 0.300s vs. 0.114s for `endpoint_bool=False` clips β€” directionally
consistent with the documented semantics (not proof, but nothing suggests
an inverted or meaningless label).

## Step 3: Development sample materialized

Stratified reservoir sample (seed=42, `src/turn_detector/data.py`),
target 300, actual 300, stratified across `endpoint_bool` Γ— `language` Γ—
`dataset` Γ— `synthetic` Γ— `midfiller` Γ— `endfiller` Γ— duration bucket.
Materialized as real 16kHz mono WAV files at
`data/raw/phase3_sample/audio/<id>.wav` + `data/raw/phase3_sample/metadata.csv`.
**300/300 clips converted successfully, 0 ffmpeg failures.**

## Steps 4-6: EXP-001 / EXP-002 / EXP-003(b) results

All **SMALL REAL-AUDIO VALIDATION** β€” 210 dev / 90 val clips (random
70/30 split, seed=42), threshold/model tuned on dev only, all metrics
below on held-out val:

| | EXP-001 (energy/silence) | EXP-002 (classical features + LR) | EXP-003b Ξ” (temporal features) |
|---|---|---|---|
| Accuracy | 0.567 | 0.622 | global-only 0.511 β†’ both 0.622 |
| F1 | 0.400 | 0.575 | **+0.112** |
| False END rate | 0.309 | 0.400 | **βˆ’0.109** (improved) |
| False CONTINUE rate | 0.629 | 0.343 | **βˆ’0.114** (improved) |
| Params / size | 2 thresholds | 92 params / 2,920 bytes | same classifier, feature-set ablation |
| Latency (this machine) | negligible (rule-based) | mean 12.2ms / p95 22.4ms per clip | n/a |

**EXP-003b's hypothesis is supported on this validation set**: adding
recent-window (100–1000ms tail) features to the global whole-clip
statistics improved F1 by +0.112 and reduced *both* error rates
simultaneously β€” not just a precision/recall tradeoff in one direction.
This is a real, if small-sample, result in favor of temporal features
mattering for this task, consistent with what streaming-relevant intuition
would predict.

## Step 7: Slice / filler analysis (real, on EXP-002's val predictions)

- **By language**: only `eng` (n=22) met the `min_samples=10` bar in the
  90-clip val set; every other language slice (including Hindi, n=3 in
  val) was too small and excluded rather than reported as if meaningful.
  `eng`-only: accuracy 0.636, F1 0.600.
- **By synthetic flag**: non-synthetic (n=12): accuracy 0.583, F1 0.667.
  Synthetic (n=78): accuracy 0.628, F1 0.554. Interesting but n=12 for the
  non-synthetic slice is thin β€” flagged as exploratory, not conclusive.
- **Filler metadata**: 72/90 val rows had `midfiller` populated (18 `None`
  β€” correctly excluded, not treated as "no filler"). `has_midfiller`
  (n=39): F1 0.529. `no_midfiller` (n=33): F1 0.593 β€” filler-present clips
  were somewhat harder for EXP-002, consistent with the qualitative error
  pattern below.

## Step 8: Error analysis (real false END / false CONTINUE cases)

Full write-up: `docs/ERROR_ANALYSIS.md`. Headline, real, measured finding:
**16 of 22 (73%) false-END errors have a filler flag present**, versus
50% for false-CONTINUE errors β€” and false-END clips have longer mean
trailing silence (0.332s) than false-CONTINUE clips (0.217s). This is
exactly the failure mode the assessment brief names as the hard case:
a pause that acoustically resembles an ending but linguistically isn't
(filler-preceded), which an audio-only, transcript-blind model has no way
to catch beyond acoustic residue.

## Step 10/11: Decision

**Question:** is the acoustic-only approach promising enough to continue,
per the three-option framework?

**Answer: Option B** β€” acoustic-only provides useful signal but is
insufficient. Evidence:

- EXP-002 (0.622 accuracy / 0.575 F1) clearly beats EXP-001 (0.567 / 0.400)
  and both clearly beat chance-level performance on a roughly balanced
  task, so acoustic features are picking up **real** signal β€” this isn't
  noise.
- But 0.622 accuracy on a near-50/50 binary task is not close to
  production-usable on its own, and the error analysis identifies a
  specific, well-understood, real ceiling: **filler-word-preceded pauses
  are systematically harder**, and no acoustic-only feature set can fully
  resolve that without something closer to lexical/semantic content β€” the
  exact thing Whisper Tiny's speech representations (not just its
  transcription output) are trained to encode implicitly, which is the
  actual empirical justification the brief asked for, now grounded in a
  real measured failure mode rather than "the brief suggested it."
- EXP-003b gives a concrete, real reason to believe **temporal/recency
  features help** (+0.112 F1), which is a design element worth carrying
  forward into whichever architecture comes next, Whisper-based or not.

This is 90 validation clips from one of ten shards β€” a first real signal,
not a final verdict. It's a directionally clear enough result to justify
moving toward EXP-003 (Whisper Tiny) as the next architecture to actually
try, while keeping EXP-002 as the benchmark it needs to beat.

## One recommended next experiment

**EXP-003: Whisper Tiny encoder representations + the same lightweight
classifier head, evaluated on the identical 210/90 dev/val split used
here**, so the comparison against EXP-002's 0.575 F1 / 0.400 false-END
rate / 0.343 false-CONTINUE rate is apples-to-apples. This directly tests
whether Whisper's representations close the filler-word gap identified in
error analysis, which is the specific, evidenced reason to expect it might
help β€” not just "the brief suggested Whisper Tiny."