- GoldWorm Coder β model bundle (gen0)
GoldWorm Coder β model bundle (gen0)
A self-contained local-first coding agent whose brain is this repository's model: the goldworm byte-level upcycled MoE. Small aux models are allowed ONLY as sub-jobs (the output layer that phrases answers is qwen2.5-0.5b-GGUF via llama.cpp, never the brain; the agent always ranks with goldworm and verifies by running the tests).
Honest scope: the brain is a 38M-param byte-LM (upcycled 4-expert MoE, ~113.6M params with experts). It is measured, not hyped. Everything below is number-first, run-through-the-project's-own-gate evidence.
What makes goldworm special (measured, not claimed)
- Byte-level MoE with REAL domain routing: 100% held-out tail router
accuracy, gate peak 0.856, bijective topicβexpert assignment. The mixture of
4 topic experts beats the best single expert (scaled
verify_moe.pymatrix: MoE mean tail 2.2385 vs best-single 2.2700, oracle 1.0734 for stories). - Gate-validated improvement, never wishful: every champion push goes through
rsi/gate.py(frozen thresholds, hash-chained audit log) β determinism proof, drift/KL guardrails, router identity, entropy collapse, contamination canary, Goodhart audit-set guard. Two strong candidates in Phase 5 were rejected by the gate even though they improved most axes (the ctx-1024 candidate still regressed TS-val; the code-heavy candidate dropped the router peak). The gate is the single door β that's the project's honesty, engineered. - Score-then-verify agent: deterministic candidate generation, goldworm BPC ranking, strict red-to-green test verification in a throwaway git worktree. The agent never trusts a patch until a test goes redβgreen.
Measured numbers (this exact bundle)
| metric | value | protocol |
|---|---|---|
| TinyStories val BPC | 0.8109 | train_scaled.eval_bpc, wins=128 |
| mean topic-tail BPC (stories/code/math/wiki) | 2.2385 | verified via verify_moe.bpc |
| mean gate peak | 0.856 (bijective) | verify_moe.gate_dist |
| router held-out accuracy | 4/4 (100%) | expert_router |
Agentic skills (product, 28/28 sub-checks, this session)
| skill | result |
|---|---|
| classify | 10/10 exact (fix/test/chat/help/generate/inspect) |
| fix (redβgreen) | 2/2 verified_by_tests |
| test | 2/2 |
| inspect/triage | 4/4 |
| chat (qwen2.5-0.5b output layer) | 5/5 (greet, 2+2, 7Γ8, capital, session memory) |
| remediation | 5/5 (clean/reset/confirm/401 gates) |
goldworm-hybrid vs qwen2.5:0.5b alone (same skills)
| skill | goldworm | qwen-alone |
|---|---|---|
| classify | 10/10 | 4/10 |
| fix | 2/2 | 1/2 (flaky) |
| inspect | 4/4 | 3/4 |
| chat | 5/5 | 5/5 (identical prompts) |
ARC-AGI probe (3 pinned games, disclosed selection)
0/3 for both arms (random β 1/16-17). ARC grids are NOT this model's strength β recorded, not hidden.
Known honest limits
The underlying model: 0/5 arithmetic, ARC chance level, confident-but-wrong on free-form generation. The product works because it never asks a weak model to write β it asks goldworm to rank, route, and verify.
How to use
- Load via the project's load path:
rsi/gate.load_moe("moe.pt", dev)β never trust the checkpoint's embedded cfg (it's stale). The loader derives the topic list from the checkpoint, so the 4-expert champion and any staged 5-expert candidate both load correctly. - The full app: the self-contained Linux bundle (start.sh) β see the
Gold_LabSpace for a showcase + the agent source + benchmarks. - The staged v8 candidate: copy or symlink the four champion
experts/*.ptintocand_v8/(onlyhtml.pt/proof.ptare stored there), thenrsi/gate.load_moe("cand_v8/cand_moe_v8.pt", dev, expert_dir="cand_v8"). Self-proof of any generated bytes:session_v8_scripts/selfprove.py(--model cand_v8/proof.pt --in file --report out.json --repaired fixed.html).
Contents
moe.ptβ champion MoE checkpointexperts/{stories,code,math,wiki}.ptβ topic experts (d512/L12)active_bundle.json,manifest.jsonβ atomic promotion pointer + hashescand_v8/β session v8 staged artifacts (see below)session_v8_scripts/β session v8 code (trainer, corpora builders, selfprove, verifier)
Session v8 artifacts (2026-10-09, staged β gate REJECTED, champion unchanged)
A 10h crash-safe training session added a 5th domain (standalone HTML) and a
self-proving denoising expert, then ran the full promotion gate on a 5-expert
candidate. Verdict: REJECT (drift kl 0.370, router binding/peak,
contamination 4.1M html template self-overlap hits, efficiency slowdown) β
per the single-door rule the champion moe.pt is untouched; all v8 artifacts
are staged under cand_v8/ for review/re-entry.
| artifact | value |
|---|---|
cand_v8/html.pt |
html expert, val BPC 2.6567 β 0.4234 (64k steps/240 min) |
cand_v8/proof.pt |
denoising x0 expert, val BPC 2.3682 β 0.5732 |
cand_v8/cand_moe_v8.pt |
5-expert candidate, balanced BPC 6.08 β 2.79 |
cand_v8/iter_v8_gate.json |
full gate verdict + metrics + rejections |
cand_v8/v8_corpus.bin |
exact gate-training bytes (contamination canary input) |
cand_v8/*_history.json, *.manifest.json, session_summary.txt |
training provenance |
session_v8_scripts/ |
the session code (crash-safe trainer, corpus builders, selfprove.py, verifier, supervisor) |
Measured honesty notes: html corpus was synthetic-only (the-stack-smol is gated; fallback recorded in the manifest); the proofer localizes damage at 78% with 9/10 clean-input hygiene; its repair effect is small at this scale.
Session v9 artifacts (2026-10-10, staged β gate REJECTED, champion unchanged)
Context scaled 256β1024 with mixed-length training (keeps short-ctx strength), the
router was fixed (routeAcc 0.726, peak 0.892 β html now binds its own expert), and
reasoning traces were mixed into code/html training. Gate verdict on
cand_v9/cand_moe_v9.pt: REJECT (frozen drift walls + contamination + 5th-expert
slowdown; audit rows 33β35) β artifacts staged under cand_v9/, champion untouched.
| artifact | value |
|---|---|
cand_v9/{stories,code,math,wiki,html}.pt |
5 ctx-1024 experts (code 1.53β1.12, html 1.47β0.33 val BPC) |
cand_v9/cand_moe_v9.pt |
5-expert candidate (balanced BPC 6.72 β 1.53) |
cand_v9/iter_v9_gate.json |
full gate verdict + metrics |
cand_v9/benchmark_hf_v9.md |
real HF test-set benchmark (below) |
cand_v9/*_history.json, html.manifest.json |
training provenance |
session_v9_scripts/ |
session code (mixed-ctx trainer, corpora, verifier, bench) |
Benchmark on real HF test sets (byte-BPC, lower=better; TinyStories val + WikiText-2 test)
| model | params | TinyStories val | WikiText-2 test |
|---|---|---|---|
| TinyStories-33M (matched) | 33M | 0.6287 | 3.62 |
| SmolLM-135M (0-shot) | 135M | 0.9568 | 1.4457 |
| gpt2-medium (0-shot) | 355M | 1.0430 | 1.4321 |
| gpt2 (0-shot) | 124M | 1.1203 | 1.5317 |
| pythia-70m (0-shot) | 70M | 1.2600 | 1.7224 |
| moe_v9 (staged) | ~190M/5x | 0.7963 | 1.9053 |
| moe_v7 (champion) | ~114M/4x | 0.8125 | 2.2305 |
moe_v9 beats the champion on every held-out corpus (incl. the champion's own).
Specialist rows (held-out audit tails): html_v9 0.3607 on HTML (html_v8:
1.9095), code_v9 1.1524 on code β within 2.4% of SmolLM-135M zero-shot.
Full protocol + raw log: cand_v9/benchmark_hf_v9.md.