Expanded v0.9 Chey follow-up training

This study inherits eligible training rows from v0.6 through the v0.9 pilot, adds the new expanded dataset, and fits one E4B decision-head candidate under the frozen protocol. The held-out authored test split is captured only after the head is locked. Live head and prompt settings are unchanged by these scripts.

C decision head โ€” 2026-09-23

Generic C11 CPU inference and a versioned GGUF sidecar for the existing trained heads. Engine topic experiment/decision-head-c-20260923; current implementation 279df9f (initial implementation 9ccaff5). Public API/format documentation is little-gemma/docs/decision-head.md in the engine repository.

  • models/decision-head-e4b.gguf, models/decision-head-12b.gguf: deterministic exports of the preserved conversational v0.1 selected heads. FP32 arrays are unchanged; full original contracts remain embedded as provenance.
  • inputs/: 64 frozen feature vectors per model, spread across archived fit, evaluation and live/probe observations, plus original Python outputs.
  • reports/nowhere-final.json: final numerical parity and head-call timing. Earlier nowhere.json and unprefixed parity JSONL preserve the initial pass.
  • reports/*-sources.json: source feature hashes and counts.
  • scripts/validate.py: export reproduction, original-package preservation, full comparison to the existing Python runner, and portable comparison.
  • scripts/check_gguf_reader.py: independent llama.cpp GGUF-reader validation.

Local full validation uses 1,046 E4B and 1,055 12B observations, including saved room/probe inputs. Both produce zero changed labels; maximum probability error is under 1.8e-7. These are numeric-equivalence checks, not new labeled accuracy tests. Historical labels, datasets and NPZ packages were not revised or trained.

Native validation uses engine 279df9f on all three machines. CTest passes: nowhere 6/6, Cortex 6/6, Windows 5/5 (the Unix socket test is not registered on Windows). Address/undefined-behavior/leak sanitizer checks also pass on nowhere. The saved CTest logs include the analytic and malformed-GGUF checks. Cortex and Windows each match all 128 frozen reference vectors, with unchanged exported GGUF hashes and original NPZ/contract hashes. llama.cpp's GGUF reader recovers all 14 tensors per model exactly; this does not test its model execution path.

Final head-only timing on nowhere uses 64 vectors per model, one excluded warmup and 50 calls per vector. Median per-vector mean: C E4B 76.51 us, Python 100.93 us; C 12B 115.46 us, Python 141.39 us. C excludes output JSON formatting; the Python method includes its normal result construction. Both exclude model loading, feature capture, IPC, ASR, GPU prefill and speech synthesis. This is not a live latency or pipeline speedup claim. Raw measurements and binary hashes are saved.

CPU host E4B C / Python (us) 12B C / Python (us)
nowhere 76.51 / 100.93 115.46 / 141.39
Cortex 143.41 / 269.40 205.12 / 339.43
somewhere 144.44 / 201.54 209.55 / 258.06

These are single sequential passes with the timing boundary above, not a randomized performance study. CPU inventory is saved in reports/hardware-*.json.

The follow-up implementation fixes probability-rounding ties to select the first class, matching the reference's argmax after softmax. An analytic test exercises that case; none of the archived-vector choices changed between implementations.

The C library/CLI is integrated into the engine build as an independent module. The trio's live Python path and settings are still the deployed path. Switching the live runner and validating that scheduling/transport change are separate work. No live voice session, training, CUDA change, or public push was performed.

The tool is also built in the normal trio engine build directories: voice-trio/little-gemma/build/decision-head on nowhere/Cortex, and voice-trio/little-gemma/build/Release/decision-head.exe on Windows. The Linux installed binaries are byte-identical to the validated isolated builds. The Windows installed build has a different executable hash and passed a separate 128-vector check (reports/somewhere-installed.json); its later timings are kept separately rather than substituted into the table above.

Engine source is integrated into authoritative nowhere research and both remote trio engine branches. Cortex's clean research checkout is fast-forwarded. Windows' original research checkout has unrelated modifications and is preserved; the candidate is available in its fetched refs and a separate worktree at voice-trio/state/decision-head-research-20260923. The GGUF packages there are identical to the canonical protected exports on nowhere.

Example local reproduction (output host name must be new):

voice-trio/.venv/bin/python \
  voice-trio/research/little-gemma/bench/decision-head-c-20260923/scripts/validate.py \
  --engine voice-trio/little-gemma \
  --binary voice-trio/little-gemma/build-decision-head/decision-head \
  --python-runner voice-trio/little-gemma-tools/script/deviceui/turn_head.py \
  --host reproduction --portable
Downloads last month
11
GGUF
Model size
497k params
Architecture
decision_head
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support