Expanded v0.9 Chey follow-up training
This study inherits eligible training rows from v0.6 through the v0.9 pilot, adds the new expanded dataset, and fits one E4B decision-head candidate under the frozen protocol. The held-out authored test split is captured only after the head is locked. Live head and prompt settings are unchanged by these scripts.
C decision head โ 2026-09-23
Generic C11 CPU inference and a versioned GGUF sidecar for the existing trained
heads. Engine topic experiment/decision-head-c-20260923; current implementation
279df9f (initial implementation 9ccaff5). Public API/format documentation is
little-gemma/docs/decision-head.md in the engine repository.
models/decision-head-e4b.gguf,models/decision-head-12b.gguf: deterministic exports of the preserved conversational v0.1 selected heads. FP32 arrays are unchanged; full original contracts remain embedded as provenance.inputs/: 64 frozen feature vectors per model, spread across archived fit, evaluation and live/probe observations, plus original Python outputs.reports/nowhere-final.json: final numerical parity and head-call timing. Earliernowhere.jsonand unprefixed parity JSONL preserve the initial pass.reports/*-sources.json: source feature hashes and counts.scripts/validate.py: export reproduction, original-package preservation, full comparison to the existing Python runner, and portable comparison.scripts/check_gguf_reader.py: independent llama.cpp GGUF-reader validation.
Local full validation uses 1,046 E4B and 1,055 12B observations, including saved room/probe inputs. Both produce zero changed labels; maximum probability error is under 1.8e-7. These are numeric-equivalence checks, not new labeled accuracy tests. Historical labels, datasets and NPZ packages were not revised or trained.
Native validation uses engine 279df9f on all three machines. CTest passes:
nowhere 6/6, Cortex 6/6, Windows 5/5 (the Unix socket test is not registered on
Windows). Address/undefined-behavior/leak sanitizer checks also pass on nowhere.
The saved CTest logs include the analytic and malformed-GGUF checks. Cortex and
Windows each match all 128 frozen reference vectors, with unchanged exported
GGUF hashes and original NPZ/contract hashes. llama.cpp's GGUF reader recovers
all 14 tensors per model exactly; this does not test its model execution path.
Final head-only timing on nowhere uses 64 vectors per model, one excluded warmup and 50 calls per vector. Median per-vector mean: C E4B 76.51 us, Python 100.93 us; C 12B 115.46 us, Python 141.39 us. C excludes output JSON formatting; the Python method includes its normal result construction. Both exclude model loading, feature capture, IPC, ASR, GPU prefill and speech synthesis. This is not a live latency or pipeline speedup claim. Raw measurements and binary hashes are saved.
| CPU host | E4B C / Python (us) | 12B C / Python (us) |
|---|---|---|
| nowhere | 76.51 / 100.93 | 115.46 / 141.39 |
| Cortex | 143.41 / 269.40 | 205.12 / 339.43 |
| somewhere | 144.44 / 201.54 | 209.55 / 258.06 |
These are single sequential passes with the timing boundary above, not a
randomized performance study. CPU inventory is saved in reports/hardware-*.json.
The follow-up implementation fixes probability-rounding ties to select the first class, matching the reference's argmax after softmax. An analytic test exercises that case; none of the archived-vector choices changed between implementations.
The C library/CLI is integrated into the engine build as an independent module. The trio's live Python path and settings are still the deployed path. Switching the live runner and validating that scheduling/transport change are separate work. No live voice session, training, CUDA change, or public push was performed.
The tool is also built in the normal trio engine build directories:
voice-trio/little-gemma/build/decision-head on nowhere/Cortex, and
voice-trio/little-gemma/build/Release/decision-head.exe on Windows. The Linux
installed binaries are byte-identical to the validated isolated builds. The
Windows installed build has a different executable hash and passed a separate
128-vector check (reports/somewhere-installed.json); its later timings are
kept separately rather than substituted into the table above.
Engine source is integrated into authoritative nowhere research and both
remote trio engine branches. Cortex's clean research checkout is fast-forwarded.
Windows' original research checkout has unrelated modifications and is preserved;
the candidate is available in its fetched refs and a separate worktree at
voice-trio/state/decision-head-research-20260923. The GGUF packages there are
identical to the canonical protected exports on nowhere.
Example local reproduction (output host name must be new):
voice-trio/.venv/bin/python \
voice-trio/research/little-gemma/bench/decision-head-c-20260923/scripts/validate.py \
--engine voice-trio/little-gemma \
--binary voice-trio/little-gemma/build-decision-head/decision-head \
--python-runner voice-trio/little-gemma-tools/script/deviceui/turn_head.py \
--host reproduction --portable
- Downloads last month
- 11
We're not able to determine the quantization variants.