Experiment registry
This is the source of truth for what has and has not been measured. A run is
QUALITY EVIDENCE only when its data, split, checkpoint, threshold, and metric
scope were registered before evaluation. Smoke tests remain smoke tests even if
their numbers look good.
Status legend:
COMPLETE β DATA: measured data/split evidence, no model-quality claim;COMPLETE β INTEGRATION: verifies code paths only;COMPLETE β DEVELOPMENT: bounded real-data development evidence, not test evidence;READY: implementation/config exists, experiment has not completed;EXTERNAL: requires data collection, compute, network, or governance work;SEALED: deliberately inaccessible until a freeze gate;RELEASE GATE: must pass before public/production claims.
Registry
| ID | Status | Question | Data and split | Configuration/artifact | Readout and decision |
|---|---|---|---|---|---|
| D-001 | COMPLETE β DATA | Can the local upstream shard be decoded, normalized, hashed, and validated? | 1/83 train shards; 3,265 rows | reports/partial_shard_audit.json |
3,265 valid, 0 invalid, 6.857 h; no exact-audio/ID/conflicting-label duplicates. Scope is one shard only. |
| D-002 | COMPLETE β DATA | Can a balanced group-safe development split be produced? | Same shard, 90/10 | reports/partial_iid_split.json; seed 42 |
2,939/326 rows; no group/audio/key crossings. Suitable for development integration, not a final corpus claim. |
| D-003 | COMPLETE β DATA | What does full source holdout expose? | Same shard, source-grouped 80/20 | reports/partial_source_holdout_split.json; seed 42 |
2,617/648, no crossings. Severe language and human/synthetic confounding; retain only as joint stress test. |
| E-000 | COMPLETE β INTEGRATION | Is training deterministic and end to end without corpus access? | Generated separable features; 32 validation examples | configs/smoke.json; 6,660 parameters; 2 epochs |
Checkpoint SHA-256 was identical across two runs. Do not use its perfect ranking as quality evidence. |
| E-001 | COMPLETE β INTEGRATION | Can an audited manifest lazily resolve its exact Parquet rows? | 8 train / 8 validation rows | artifacts/resolver-smoke/ |
Resolver, decode, frontend, collator, forward/backward, and checkpoint path completed. No statistical interpretation. |
| E-002 | COMPLETE β INTEGRATION | Does real audio train/evaluate through the debug path? | First capped 32 train / 32 validation records from partial IID manifest | 6,660 parameters; 2 s suffix; 1 epoch; artifacts/real-shard-smoke/ |
AUROC 0.246, FPR 0.111 and recall 0.043 at the selected loose 20% FPR-budget threshold. Expected poor debug result; reject as model-quality evidence. |
| E-003 | COMPLETE β DEVELOPMENT | How much endpoint signal exists in simple inspectable acoustics? | All 2,939 train / 326 validation rows in partial IID split; observed-key overlap zero, but all groups are singleton rows | Ten 4 s suffix waveform statistics + class-balanced logistic regression; artifacts/partial-baseline/metrics.json |
AUROC 0.7522, AP 0.7310; at validation-selected FPR β€ 0.02, threshold 0.7639 gives FPR 0.0181 and END recall 0.1313. Best-effort row-disjoint development floor only. |
| E-004 | COMPLETE β DEVELOPMENT | Can a 151k-parameter four-second TinyTCN learn the one-shard task? | Same 2,939/326 partial IID split | configs/partial_shard_preview.json; 3 epochs, LR 3e-4; artifacts/partial-shard-preview-tinytcn-4s/ |
Best AP 0.6874, AUROC 0.7291; recall 0.0938 at FPR 0.0181. Continue only as bounded development. |
| E-005 | COMPLETE β DEVELOPMENT | Does a higher learning rate improve the same TinyTCN? | Same partial IID split | configs/partial_shard_preview_lr1e3_8ep.json; LR 1e-3 |
Deliberately stopped after epoch 4 for poor trajectory; best AP 0.6190, AUROC 0.6829, recall 0.0313 at FPR 0.0181. Reject. |
| E-006 | COMPLETE β DEVELOPMENT | Does a fresh-optimizer continuation improve the selected TinyTCN? | Same partial IID split, already used adaptively | configs/partial_shard_warmstart_lr3e4_5ep.json; 5 continuation epochs; artifacts/partial-shard-warmstart-lr3e4-5ep/ |
Selected continuation epoch 1 by AP: AP 0.6881, AUROC 0.7390, Brier 0.2133; 10 TP/3 FP/163 TN/150 FN at threshold 0.7410. Baseline remains stronger. Packaged preview only. |
| E-007 | COMPLETE β DEVELOPMENT | Do baseline and TinyTCN survive held-out source groups? | 2,617/648 source-held-out split; 9 validation source groups | configs/partial_source_holdout_preview.json; source baseline and TinyTCN reports |
AP 0.5368 baseline vs 0.5500 TinyTCN; both AUROC about 0.55. Wide paired intervals and severe confounding establish no winner. |
| E-008 | COMPLETE β DEVELOPMENT | Does audio framing double-count trailing silence? | Same 326 IID validation rows; paired 0/200/400/800 ms valid-zero perturbations | reports/partial_tinytcn_silence_sensitivity.json; fixed E-006 threshold |
At 800 ms, 26/326 decisions flip and FPR rises 0.0181β0.0361; mean absolute score shift 0.1058. Add silence robustness and recalibrate controller. |
| E-009 | COMPLETE β INTEGRATION | Does the serialized controller latch END and emit only one response edge? | 8 hand-authored checkpoints across 3 synthetic turns; no ONNX scores or human labels | data/collection/controller_replay_fixture.jsonl; reports/controller_replay_integration.summary.json; E-006 metadata policy |
Three response edges, zero duplicates; model endpoint, relaxation, timeout, early interruption, and post-END callbacks exercised. Plumbing evidence only. |
| E-010 | READY | What is the TinyTCN full-data baseline? | All audited train shards; group-safe train/validation | configs/tiny_tcn.yaml; seed 17; 396,164 parameters |
Run after D-010. Primary selection: END recall under FPR β€ 0.02. |
| E-011 | READY | How much does source shift hurt? | Leave-one-source-out folds plus human-only slices | Same checkpoint family and frozen preprocessing | Report every viable source fold, macro distribution, raw counts, and confounding; do not select on one favorable fold. |
| E-012 | READY | Does a Whisper encoder improve constrained recall enough to justify cost? | Same rows and folds as E-010/E-011 | configs/whisper_teacher.yaml; frozen encoder first |
Compare paired predictions at identical FPR budgets plus size/latency/RSS. |
| E-013 | READY | Can teacher supervision improve the small student? | Same train/dev policy; cached teacher predictions | Distillation objective in src/turn_detection/training/losses.py |
Run only after E-012 validates the teacher. Compare hard-label student vs distilled student, same seed grid. |
| E-014 | READY | Which audio suffix is Pareto-optimal? | Same split | 2 s vs 4 s vs 8 s; student sizes 96Γ6, 128Γ8, 192Γ8 | Select on false interruptions, endpoint delay, CPU p95, bytesβnot accuracy alone. |
| E-015 | READY | Do filler auxiliaries help the intended failure cases? | Same split plus midfiller/endfiller slices | Endpoint-only vs masked multitask; all else fixed | Require paired grouped intervals, not aggregate delta only. |
| E-016 | READY | Does INT8 preserve decisions and latency? | Representative train-only calibration; frozen dev | FP32 ONNX vs static INT8 | Check numerical delta, recalibrate INT8 threshold, then compare constrained metrics and end-to-end latency. |
| H-001 | EXTERNAL | Does the model work on natural Shiprocket Hinglish holds? | 900 planned recordings, 30 speakers, 21/4/5 speaker split | data/collection/assignments.jsonl; docs/data_collection_protocol.md |
No recordings exist. Collect with consent; lock speaker-disjoint test; report agreement and all exclusions. |
| T-001 | SEALED | What is final generalization on the official test? | Pinned 10-shard Smart Turn test snapshot | scripts/download_test_dataset.sh |
Open once only after model/frontend/controller/threshold hashes are frozen. No post-test tuning. |
| R-001 | RELEASE GATE | Is the artifact publishable? | Final checkpoint, official-test metrics, Hinglish evidence | ONNX metadata/parity, benchmark, model/data cards, build_release.py |
FP32 development preview export/benchmark is complete; final empirical and rights gates fail. A preview may be packaged only with conspicuous development overrides. |
Completed-run details
D-001 β partial shard audit
- Generated: 2026-08-22T19:58:38Z.
- File:
train-00010-of-00083.parquet. - Duration range: 0.36β29.56 seconds.
- Endpoint balance: 1,667 HOLD / 1,598 END.
- Synthetic balance: 2,712 synthetic / 553 human-tagged.
- Filler labels:
midfiller1,339 true / 1,284 false / 642 null;endfiller905 true / 1,718 false / 642 null. - Known interpretation risk: this shard can be non-random and source-clustered.
Decision: the parser/schema are compatible with this shard. Audit the other 82 before training or making corpus-level distribution claims.
E-000 β synthetic deterministic smoke
- Model: 32 channels, 2 blocks, 6,660 parameters.
- Data: generated feature patterns, not audio.
- Training: 2 epochs on CPU, seed 17.
- Best validation loss: 0.6720276.
- Stored threshold: 0.4810289 under a deliberately loose smoke FPR budget 0.2.
- Checkpoint bytes: 111,159.
- Repeated-run checkpoint SHA-256:
fc519b0b868c31d3429316babe6a844086c58516381e04ac29c550d86df7a644. - Neural-only microbenchmark: PyTorch 2.11, Apple arm, one thread, batch 1, 100 frames, 3 warmups + 10 measurements; warm mean 0.247 ms.
Decision: deterministic training/checkpointing and benchmark instrumentation work. All classification numbers are invalid as real-audio evidence because the features were constructed to be separable.
E-002 β 32-example-per-split real-audio smoke
- Source: partial IID manifest resolved back to the local Parquet shard.
- Cap: 32 examples per split; validation contains 23 END / 9 HOLD.
- Model: 32 channels, 2 blocks, 6,660 parameters.
- Audio suffix: 2 seconds.
- Training: 1 epoch, batch size 4, CPU path.
- Best validation loss: 1.1148809.
- At threshold 0.5: recall 0, FPR 0, balanced accuracy 0.5.
- Stored operating threshold: 0.3017663 under the debug FPR budget 0.2.
- At stored threshold: 1 TP, 1 FP, 8 TN, 22 FN; recall 0.0435, FPR 0.1111, AUROC 0.2464.
- Checkpoint bytes: 111,159.
Decision: the data-addressing fix is validated. The run is intentionally too small and short to learn; do not compare it with any baseline.
E-003 β full partial-shard acoustic baseline
- Source: all rows in the partial IID split, with zero observed-key crossing. All 3,265 base IID linkage groups are singletons because speaker/conversation/voice IDs are absent, so this is not speaker-disjoint evidence. The source stress split separately groups by dataset source.
- Model: class-balanced full-batch logistic regression with ten interpretable waveform features; 1,000 iterations, learning rate 0.05, L2 0.001.
- Context: most recent four seconds.
- Train / validation: 2,939 / 326.
- Validation discrimination: AUROC 0.7521837, AP 0.7310146.
- Calibration: Brier 0.2038291, ECE 0.0867866.
- Threshold selected on validation under FPR β€ 0.02: 0.7638752.
- At that threshold: 21 TP, 3 FP, 163 TN, 139 FN; END recall 0.13125, FPR 0.0180723, precision 0.875, macro-F1 0.4624210.
Decision: simple acoustics contain useful endpoint signal. Treat this as a development floor for neural experiments, not an unbiased generalization estimate: the threshold and metrics use the same 326-item validation set, the source is one of 83 shards, and there is no Hinglish or official-test evidence.
E-004 to E-006 β partial-shard TinyTCN selection
- Architecture: 151,812 parameters, six 128-channel blocks, four-second suffix.
- E-004 best: AP 0.6873543, AUROC 0.7290663, recall 0.09375 at FPR 0.0180723.
- E-005: LR
1e-3; stopped after epoch 4; best AP 0.6189746 and AUROC 0.6828690. Rejected. - E-006: weights-only continuation of E-004 with fresh optimizer, LR
3e-4. Five epochs completed; continuation epoch 1 selected by AP. It has AP 0.6880690, AUROC 0.7389684, Brier 0.2133037, and threshold 0.7410008. Confusion is 10 TP / 3 FP / 163 TN / 150 FN. - Against E-003, E-006 minus baseline AP is β0.04295 (95% paired bootstrap CI [β0.11945, 0.03524]); AUROC β0.01322 [β0.07823, 0.05121]; constrained recall β0.06875 [β0.13298, β0.01220].
Decision: package E-006 as a fast development artifact, while retaining E-003 as the evidence winner. The same development split informed prior runs, checkpoint selection, and thresholds; intervals are conditional and optimistic.
E-007 β source-held-out stress
The baseline/TinyTCN respectively reach AP 0.53684/0.54996, AUROC 0.55387/0.55227, and recall 0.02786/0.04025 at FPR 0.01846. The paired AP delta is +0.01312 with nine-source 95% CI [β0.21492, 0.12613]. Source, language, and human/synthetic status shift together; no causal source claim or winner is supported.
E-008 β silence sensitivity
At the frozen E-006 threshold, appending 200/400/800 ms valid zero-valued silence changes FPR from 0.0181 to 0.0301/0.0301/0.0361 and flips 5/13/26 of 326 decisions. At 800 ms mean absolute probability shift is 0.1058. Because the controller also consumes silence duration, add trailing-silence jitter or speech-end normalization before controller recalibration.
E-009 β synthetic controller integration
The metadata-bound replay covers eight checkpoints across three fabricated
turns. It produces three response edges and zero duplicate emissions while
exercising a model endpoint, threshold relaxation, maximum timeout, premature
endpoint, and callbacks after a latched END. The input fixture SHA-256 is
a96626ce8e93bff0a88f00f652932c5ac806377e67fc4fec3b03ae34658f9a04;
the decision JSONL SHA-256 is
171ecc5129292f0b6fabdad4d3f0e6ed013817755183f648e7b8f6dc7788542c.
Decision: this proves controller plumbing and artifact-policy binding only. Probabilities and targets are hand-authored, latency values are synthetic, and none of its rates or delays are model, conversation, latency, or product-quality evidence. Real annotated multi-checkpoint replay remains external work.
Pre-registration template
Copy this block into the registry before starting any quality experiment:
Experiment ID:
Question/hypothesis:
Code revision:
Dataset IDs and immutable revisions:
Manifest SHA-256:
Grouping policy version:
Train/development/test split hashes:
Model config and parameter count:
Frontend and suffix duration:
Random seeds and number of repeats:
Primary metric and FPR budget:
Secondary metrics/slices:
Threshold-selection split and rule:
Early-stopping rule:
Maximum examples/epochs/compute:
Promotion criterion:
Known confounders:
After the run, append checkpoint/ONNX hashes, environment versions, training history path, prediction-file path, threshold, metric JSON path, latency JSON path, failures reviewed, and the decision. Never overwrite a completed run to make a rerun look like the same evidence.
Full-corpus experiment sequence
D-010: download and audit all 83 train shards; retain per-shard reports.D-011: freeze manifest, grouping policy, and IID/source-LOO split hashes.E-010: run the default student on three seeds.E-014: test duration/model-size ablations; identify Pareto candidates.E-015: isolate the filler-auxiliary contribution.E-012: run Whisper frozen-head; selectively unfreeze only if warranted.E-013: distill only if the teacher has a meaningful paired advantage.H-001: collect/annotate/lock the Hinglish benchmark independently.E-016: export and quantize the selected student; recalibrate threshold.- Freeze hashes and controller policy; execute
T-001once. - Run
R-001, update model card with measured evidence, then publish.
Metric reporting contract
Every quality table must state:
- whether
1means END (it does in this repository); - exact split and unique speakers/conversations/groups;
- confusion counts before rates;
- threshold and where/how it was selected;
- FPR denominator
FP + TNand FNR denominatorFN + TP; - slice numerator, denominator, and suppression threshold;
- whether latency is neural-only or end to end;
- hardware, threads, batch, warmup, repetitions, runtime, and quantization;
- grouped uncertainty interval and grouping unit;
- whether the official test was used at any point.
This prevents a debug metric, favorable slice, or neural-only microbenchmark from silently becoming a product claim.