tiny-hinglish-turn-detector / docs /experiment_registry.md
suvradeepp's picture
Publish Tiny Hinglish Turn Detector development preview
35d483e verified
|
Raw
History Blame Contribute Delete
16.6 kB

Experiment registry

This is the source of truth for what has and has not been measured. A run is QUALITY EVIDENCE only when its data, split, checkpoint, threshold, and metric scope were registered before evaluation. Smoke tests remain smoke tests even if their numbers look good.

Status legend:

  • COMPLETE β€” DATA: measured data/split evidence, no model-quality claim;
  • COMPLETE β€” INTEGRATION: verifies code paths only;
  • COMPLETE β€” DEVELOPMENT: bounded real-data development evidence, not test evidence;
  • READY: implementation/config exists, experiment has not completed;
  • EXTERNAL: requires data collection, compute, network, or governance work;
  • SEALED: deliberately inaccessible until a freeze gate;
  • RELEASE GATE: must pass before public/production claims.

Registry

ID Status Question Data and split Configuration/artifact Readout and decision
D-001 COMPLETE β€” DATA Can the local upstream shard be decoded, normalized, hashed, and validated? 1/83 train shards; 3,265 rows reports/partial_shard_audit.json 3,265 valid, 0 invalid, 6.857 h; no exact-audio/ID/conflicting-label duplicates. Scope is one shard only.
D-002 COMPLETE β€” DATA Can a balanced group-safe development split be produced? Same shard, 90/10 reports/partial_iid_split.json; seed 42 2,939/326 rows; no group/audio/key crossings. Suitable for development integration, not a final corpus claim.
D-003 COMPLETE β€” DATA What does full source holdout expose? Same shard, source-grouped 80/20 reports/partial_source_holdout_split.json; seed 42 2,617/648, no crossings. Severe language and human/synthetic confounding; retain only as joint stress test.
E-000 COMPLETE β€” INTEGRATION Is training deterministic and end to end without corpus access? Generated separable features; 32 validation examples configs/smoke.json; 6,660 parameters; 2 epochs Checkpoint SHA-256 was identical across two runs. Do not use its perfect ranking as quality evidence.
E-001 COMPLETE β€” INTEGRATION Can an audited manifest lazily resolve its exact Parquet rows? 8 train / 8 validation rows artifacts/resolver-smoke/ Resolver, decode, frontend, collator, forward/backward, and checkpoint path completed. No statistical interpretation.
E-002 COMPLETE β€” INTEGRATION Does real audio train/evaluate through the debug path? First capped 32 train / 32 validation records from partial IID manifest 6,660 parameters; 2 s suffix; 1 epoch; artifacts/real-shard-smoke/ AUROC 0.246, FPR 0.111 and recall 0.043 at the selected loose 20% FPR-budget threshold. Expected poor debug result; reject as model-quality evidence.
E-003 COMPLETE β€” DEVELOPMENT How much endpoint signal exists in simple inspectable acoustics? All 2,939 train / 326 validation rows in partial IID split; observed-key overlap zero, but all groups are singleton rows Ten 4 s suffix waveform statistics + class-balanced logistic regression; artifacts/partial-baseline/metrics.json AUROC 0.7522, AP 0.7310; at validation-selected FPR ≀ 0.02, threshold 0.7639 gives FPR 0.0181 and END recall 0.1313. Best-effort row-disjoint development floor only.
E-004 COMPLETE β€” DEVELOPMENT Can a 151k-parameter four-second TinyTCN learn the one-shard task? Same 2,939/326 partial IID split configs/partial_shard_preview.json; 3 epochs, LR 3e-4; artifacts/partial-shard-preview-tinytcn-4s/ Best AP 0.6874, AUROC 0.7291; recall 0.0938 at FPR 0.0181. Continue only as bounded development.
E-005 COMPLETE β€” DEVELOPMENT Does a higher learning rate improve the same TinyTCN? Same partial IID split configs/partial_shard_preview_lr1e3_8ep.json; LR 1e-3 Deliberately stopped after epoch 4 for poor trajectory; best AP 0.6190, AUROC 0.6829, recall 0.0313 at FPR 0.0181. Reject.
E-006 COMPLETE β€” DEVELOPMENT Does a fresh-optimizer continuation improve the selected TinyTCN? Same partial IID split, already used adaptively configs/partial_shard_warmstart_lr3e4_5ep.json; 5 continuation epochs; artifacts/partial-shard-warmstart-lr3e4-5ep/ Selected continuation epoch 1 by AP: AP 0.6881, AUROC 0.7390, Brier 0.2133; 10 TP/3 FP/163 TN/150 FN at threshold 0.7410. Baseline remains stronger. Packaged preview only.
E-007 COMPLETE β€” DEVELOPMENT Do baseline and TinyTCN survive held-out source groups? 2,617/648 source-held-out split; 9 validation source groups configs/partial_source_holdout_preview.json; source baseline and TinyTCN reports AP 0.5368 baseline vs 0.5500 TinyTCN; both AUROC about 0.55. Wide paired intervals and severe confounding establish no winner.
E-008 COMPLETE β€” DEVELOPMENT Does audio framing double-count trailing silence? Same 326 IID validation rows; paired 0/200/400/800 ms valid-zero perturbations reports/partial_tinytcn_silence_sensitivity.json; fixed E-006 threshold At 800 ms, 26/326 decisions flip and FPR rises 0.0181β†’0.0361; mean absolute score shift 0.1058. Add silence robustness and recalibrate controller.
E-009 COMPLETE β€” INTEGRATION Does the serialized controller latch END and emit only one response edge? 8 hand-authored checkpoints across 3 synthetic turns; no ONNX scores or human labels data/collection/controller_replay_fixture.jsonl; reports/controller_replay_integration.summary.json; E-006 metadata policy Three response edges, zero duplicates; model endpoint, relaxation, timeout, early interruption, and post-END callbacks exercised. Plumbing evidence only.
E-010 READY What is the TinyTCN full-data baseline? All audited train shards; group-safe train/validation configs/tiny_tcn.yaml; seed 17; 396,164 parameters Run after D-010. Primary selection: END recall under FPR ≀ 0.02.
E-011 READY How much does source shift hurt? Leave-one-source-out folds plus human-only slices Same checkpoint family and frozen preprocessing Report every viable source fold, macro distribution, raw counts, and confounding; do not select on one favorable fold.
E-012 READY Does a Whisper encoder improve constrained recall enough to justify cost? Same rows and folds as E-010/E-011 configs/whisper_teacher.yaml; frozen encoder first Compare paired predictions at identical FPR budgets plus size/latency/RSS.
E-013 READY Can teacher supervision improve the small student? Same train/dev policy; cached teacher predictions Distillation objective in src/turn_detection/training/losses.py Run only after E-012 validates the teacher. Compare hard-label student vs distilled student, same seed grid.
E-014 READY Which audio suffix is Pareto-optimal? Same split 2 s vs 4 s vs 8 s; student sizes 96Γ—6, 128Γ—8, 192Γ—8 Select on false interruptions, endpoint delay, CPU p95, bytesβ€”not accuracy alone.
E-015 READY Do filler auxiliaries help the intended failure cases? Same split plus midfiller/endfiller slices Endpoint-only vs masked multitask; all else fixed Require paired grouped intervals, not aggregate delta only.
E-016 READY Does INT8 preserve decisions and latency? Representative train-only calibration; frozen dev FP32 ONNX vs static INT8 Check numerical delta, recalibrate INT8 threshold, then compare constrained metrics and end-to-end latency.
H-001 EXTERNAL Does the model work on natural Shiprocket Hinglish holds? 900 planned recordings, 30 speakers, 21/4/5 speaker split data/collection/assignments.jsonl; docs/data_collection_protocol.md No recordings exist. Collect with consent; lock speaker-disjoint test; report agreement and all exclusions.
T-001 SEALED What is final generalization on the official test? Pinned 10-shard Smart Turn test snapshot scripts/download_test_dataset.sh Open once only after model/frontend/controller/threshold hashes are frozen. No post-test tuning.
R-001 RELEASE GATE Is the artifact publishable? Final checkpoint, official-test metrics, Hinglish evidence ONNX metadata/parity, benchmark, model/data cards, build_release.py FP32 development preview export/benchmark is complete; final empirical and rights gates fail. A preview may be packaged only with conspicuous development overrides.

Completed-run details

D-001 β€” partial shard audit

  • Generated: 2026-08-22T19:58:38Z.
  • File: train-00010-of-00083.parquet.
  • Duration range: 0.36–29.56 seconds.
  • Endpoint balance: 1,667 HOLD / 1,598 END.
  • Synthetic balance: 2,712 synthetic / 553 human-tagged.
  • Filler labels: midfiller 1,339 true / 1,284 false / 642 null; endfiller 905 true / 1,718 false / 642 null.
  • Known interpretation risk: this shard can be non-random and source-clustered.

Decision: the parser/schema are compatible with this shard. Audit the other 82 before training or making corpus-level distribution claims.

E-000 β€” synthetic deterministic smoke

  • Model: 32 channels, 2 blocks, 6,660 parameters.
  • Data: generated feature patterns, not audio.
  • Training: 2 epochs on CPU, seed 17.
  • Best validation loss: 0.6720276.
  • Stored threshold: 0.4810289 under a deliberately loose smoke FPR budget 0.2.
  • Checkpoint bytes: 111,159.
  • Repeated-run checkpoint SHA-256: fc519b0b868c31d3429316babe6a844086c58516381e04ac29c550d86df7a644.
  • Neural-only microbenchmark: PyTorch 2.11, Apple arm, one thread, batch 1, 100 frames, 3 warmups + 10 measurements; warm mean 0.247 ms.

Decision: deterministic training/checkpointing and benchmark instrumentation work. All classification numbers are invalid as real-audio evidence because the features were constructed to be separable.

E-002 β€” 32-example-per-split real-audio smoke

  • Source: partial IID manifest resolved back to the local Parquet shard.
  • Cap: 32 examples per split; validation contains 23 END / 9 HOLD.
  • Model: 32 channels, 2 blocks, 6,660 parameters.
  • Audio suffix: 2 seconds.
  • Training: 1 epoch, batch size 4, CPU path.
  • Best validation loss: 1.1148809.
  • At threshold 0.5: recall 0, FPR 0, balanced accuracy 0.5.
  • Stored operating threshold: 0.3017663 under the debug FPR budget 0.2.
  • At stored threshold: 1 TP, 1 FP, 8 TN, 22 FN; recall 0.0435, FPR 0.1111, AUROC 0.2464.
  • Checkpoint bytes: 111,159.

Decision: the data-addressing fix is validated. The run is intentionally too small and short to learn; do not compare it with any baseline.

E-003 β€” full partial-shard acoustic baseline

  • Source: all rows in the partial IID split, with zero observed-key crossing. All 3,265 base IID linkage groups are singletons because speaker/conversation/voice IDs are absent, so this is not speaker-disjoint evidence. The source stress split separately groups by dataset source.
  • Model: class-balanced full-batch logistic regression with ten interpretable waveform features; 1,000 iterations, learning rate 0.05, L2 0.001.
  • Context: most recent four seconds.
  • Train / validation: 2,939 / 326.
  • Validation discrimination: AUROC 0.7521837, AP 0.7310146.
  • Calibration: Brier 0.2038291, ECE 0.0867866.
  • Threshold selected on validation under FPR ≀ 0.02: 0.7638752.
  • At that threshold: 21 TP, 3 FP, 163 TN, 139 FN; END recall 0.13125, FPR 0.0180723, precision 0.875, macro-F1 0.4624210.

Decision: simple acoustics contain useful endpoint signal. Treat this as a development floor for neural experiments, not an unbiased generalization estimate: the threshold and metrics use the same 326-item validation set, the source is one of 83 shards, and there is no Hinglish or official-test evidence.

E-004 to E-006 β€” partial-shard TinyTCN selection

  • Architecture: 151,812 parameters, six 128-channel blocks, four-second suffix.
  • E-004 best: AP 0.6873543, AUROC 0.7290663, recall 0.09375 at FPR 0.0180723.
  • E-005: LR 1e-3; stopped after epoch 4; best AP 0.6189746 and AUROC 0.6828690. Rejected.
  • E-006: weights-only continuation of E-004 with fresh optimizer, LR 3e-4. Five epochs completed; continuation epoch 1 selected by AP. It has AP 0.6880690, AUROC 0.7389684, Brier 0.2133037, and threshold 0.7410008. Confusion is 10 TP / 3 FP / 163 TN / 150 FN.
  • Against E-003, E-006 minus baseline AP is βˆ’0.04295 (95% paired bootstrap CI [βˆ’0.11945, 0.03524]); AUROC βˆ’0.01322 [βˆ’0.07823, 0.05121]; constrained recall βˆ’0.06875 [βˆ’0.13298, βˆ’0.01220].

Decision: package E-006 as a fast development artifact, while retaining E-003 as the evidence winner. The same development split informed prior runs, checkpoint selection, and thresholds; intervals are conditional and optimistic.

E-007 β€” source-held-out stress

The baseline/TinyTCN respectively reach AP 0.53684/0.54996, AUROC 0.55387/0.55227, and recall 0.02786/0.04025 at FPR 0.01846. The paired AP delta is +0.01312 with nine-source 95% CI [βˆ’0.21492, 0.12613]. Source, language, and human/synthetic status shift together; no causal source claim or winner is supported.

E-008 β€” silence sensitivity

At the frozen E-006 threshold, appending 200/400/800 ms valid zero-valued silence changes FPR from 0.0181 to 0.0301/0.0301/0.0361 and flips 5/13/26 of 326 decisions. At 800 ms mean absolute probability shift is 0.1058. Because the controller also consumes silence duration, add trailing-silence jitter or speech-end normalization before controller recalibration.

E-009 β€” synthetic controller integration

The metadata-bound replay covers eight checkpoints across three fabricated turns. It produces three response edges and zero duplicate emissions while exercising a model endpoint, threshold relaxation, maximum timeout, premature endpoint, and callbacks after a latched END. The input fixture SHA-256 is a96626ce8e93bff0a88f00f652932c5ac806377e67fc4fec3b03ae34658f9a04; the decision JSONL SHA-256 is 171ecc5129292f0b6fabdad4d3f0e6ed013817755183f648e7b8f6dc7788542c.

Decision: this proves controller plumbing and artifact-policy binding only. Probabilities and targets are hand-authored, latency values are synthetic, and none of its rates or delays are model, conversation, latency, or product-quality evidence. Real annotated multi-checkpoint replay remains external work.

Pre-registration template

Copy this block into the registry before starting any quality experiment:

Experiment ID:
Question/hypothesis:
Code revision:
Dataset IDs and immutable revisions:
Manifest SHA-256:
Grouping policy version:
Train/development/test split hashes:
Model config and parameter count:
Frontend and suffix duration:
Random seeds and number of repeats:
Primary metric and FPR budget:
Secondary metrics/slices:
Threshold-selection split and rule:
Early-stopping rule:
Maximum examples/epochs/compute:
Promotion criterion:
Known confounders:

After the run, append checkpoint/ONNX hashes, environment versions, training history path, prediction-file path, threshold, metric JSON path, latency JSON path, failures reviewed, and the decision. Never overwrite a completed run to make a rerun look like the same evidence.

Full-corpus experiment sequence

  1. D-010: download and audit all 83 train shards; retain per-shard reports.
  2. D-011: freeze manifest, grouping policy, and IID/source-LOO split hashes.
  3. E-010: run the default student on three seeds.
  4. E-014: test duration/model-size ablations; identify Pareto candidates.
  5. E-015: isolate the filler-auxiliary contribution.
  6. E-012: run Whisper frozen-head; selectively unfreeze only if warranted.
  7. E-013: distill only if the teacher has a meaningful paired advantage.
  8. H-001: collect/annotate/lock the Hinglish benchmark independently.
  9. E-016: export and quantize the selected student; recalibrate threshold.
  10. Freeze hashes and controller policy; execute T-001 once.
  11. Run R-001, update model card with measured evidence, then publish.

Metric reporting contract

Every quality table must state:

  • whether 1 means END (it does in this repository);
  • exact split and unique speakers/conversations/groups;
  • confusion counts before rates;
  • threshold and where/how it was selected;
  • FPR denominator FP + TN and FNR denominator FN + TP;
  • slice numerator, denominator, and suppression threshold;
  • whether latency is neural-only or end to end;
  • hardware, threads, batch, warmup, repetitions, runtime, and quantization;
  • grouped uncertainty interval and grouping unit;
  • whether the official test was used at any point.

This prevents a debug metric, favorable slice, or neural-only microbenchmark from silently becoming a product claim.