GoldWorm Coder β€” model bundle (gen0)

A self-contained local-first coding agent whose brain is this repository's model: the goldworm byte-level upcycled MoE. Small aux models are allowed ONLY as sub-jobs (the output layer that phrases answers is qwen2.5-0.5b-GGUF via llama.cpp, never the brain; the agent always ranks with goldworm and verifies by running the tests).

Honest scope: the brain is a 38M-param byte-LM (upcycled 4-expert MoE, ~113.6M params with experts). It is measured, not hyped. Everything below is number-first, run-through-the-project's-own-gate evidence.

What makes goldworm special (measured, not claimed)

  • Byte-level MoE with REAL domain routing: 100% held-out tail router accuracy, gate peak 0.856, bijective topicβ†’expert assignment. The mixture of 4 topic experts beats the best single expert (scaled verify_moe.py matrix: MoE mean tail 2.2385 vs best-single 2.2700, oracle 1.0734 for stories).
  • Gate-validated improvement, never wishful: every champion push goes through rsi/gate.py (frozen thresholds, hash-chained audit log) β€” determinism proof, drift/KL guardrails, router identity, entropy collapse, contamination canary, Goodhart audit-set guard. Two strong candidates in Phase 5 were rejected by the gate even though they improved most axes (the ctx-1024 candidate still regressed TS-val; the code-heavy candidate dropped the router peak). The gate is the single door β€” that's the project's honesty, engineered.
  • Score-then-verify agent: deterministic candidate generation, goldworm BPC ranking, strict red-to-green test verification in a throwaway git worktree. The agent never trusts a patch until a test goes redβ†’green.

Measured numbers (this exact bundle)

metric value protocol
TinyStories val BPC 0.8109 train_scaled.eval_bpc, wins=128
mean topic-tail BPC (stories/code/math/wiki) 2.2385 verified via verify_moe.bpc
mean gate peak 0.856 (bijective) verify_moe.gate_dist
router held-out accuracy 4/4 (100%) expert_router

Agentic skills (product, 28/28 sub-checks, this session)

skill result
classify 10/10 exact (fix/test/chat/help/generate/inspect)
fix (red→green) 2/2 verified_by_tests
test 2/2
inspect/triage 4/4
chat (qwen2.5-0.5b output layer) 5/5 (greet, 2+2, 7Γ—8, capital, session memory)
remediation 5/5 (clean/reset/confirm/401 gates)

goldworm-hybrid vs qwen2.5:0.5b alone (same skills)

skill goldworm qwen-alone
classify 10/10 4/10
fix 2/2 1/2 (flaky)
inspect 4/4 3/4
chat 5/5 5/5 (identical prompts)

ARC-AGI probe (3 pinned games, disclosed selection)

0/3 for both arms (random β‰ˆ 1/16-17). ARC grids are NOT this model's strength β€” recorded, not hidden.

Known honest limits

The underlying model: 0/5 arithmetic, ARC chance level, confident-but-wrong on free-form generation. The product works because it never asks a weak model to write β€” it asks goldworm to rank, route, and verify.

How to use

  1. Load via the project's load path: rsi/gate.load_moe("moe.pt", dev) β€” never trust the checkpoint's embedded cfg (it's stale). The loader derives the topic list from the checkpoint, so the 4-expert champion and any staged 5-expert candidate both load correctly.
  2. The full app: the self-contained Linux bundle (start.sh) β€” see the Gold_Lab Space for a showcase + the agent source + benchmarks.
  3. The staged v8 candidate: copy or symlink the four champion experts/*.pt into cand_v8/ (only html.pt / proof.pt are stored there), then rsi/gate.load_moe("cand_v8/cand_moe_v8.pt", dev, expert_dir="cand_v8"). Self-proof of any generated bytes: session_v8_scripts/selfprove.py (--model cand_v8/proof.pt --in file --report out.json --repaired fixed.html).

Contents

  • moe.pt β€” champion MoE checkpoint
  • experts/{stories,code,math,wiki}.pt β€” topic experts (d512/L12)
  • active_bundle.json, manifest.json β€” atomic promotion pointer + hashes
  • cand_v8/ β€” session v8 staged artifacts (see below)
  • session_v8_scripts/ β€” session v8 code (trainer, corpora builders, selfprove, verifier)

Session v8 artifacts (2026-10-09, staged β€” gate REJECTED, champion unchanged)

A 10h crash-safe training session added a 5th domain (standalone HTML) and a self-proving denoising expert, then ran the full promotion gate on a 5-expert candidate. Verdict: REJECT (drift kl 0.370, router binding/peak, contamination 4.1M html template self-overlap hits, efficiency slowdown) β€” per the single-door rule the champion moe.pt is untouched; all v8 artifacts are staged under cand_v8/ for review/re-entry.

artifact value
cand_v8/html.pt html expert, val BPC 2.6567 β†’ 0.4234 (64k steps/240 min)
cand_v8/proof.pt denoising x0 expert, val BPC 2.3682 β†’ 0.5732
cand_v8/cand_moe_v8.pt 5-expert candidate, balanced BPC 6.08 β†’ 2.79
cand_v8/iter_v8_gate.json full gate verdict + metrics + rejections
cand_v8/v8_corpus.bin exact gate-training bytes (contamination canary input)
cand_v8/*_history.json, *.manifest.json, session_summary.txt training provenance
session_v8_scripts/ the session code (crash-safe trainer, corpus builders, selfprove.py, verifier, supervisor)

Measured honesty notes: html corpus was synthetic-only (the-stack-smol is gated; fallback recorded in the manifest); the proofer localizes damage at 78% with 9/10 clean-input hygiene; its repair effect is small at this scale.

Session v9 artifacts (2026-10-10, staged β€” gate REJECTED, champion unchanged)

Context scaled 256β†’1024 with mixed-length training (keeps short-ctx strength), the router was fixed (routeAcc 0.726, peak 0.892 β€” html now binds its own expert), and reasoning traces were mixed into code/html training. Gate verdict on cand_v9/cand_moe_v9.pt: REJECT (frozen drift walls + contamination + 5th-expert slowdown; audit rows 33β†’35) β€” artifacts staged under cand_v9/, champion untouched.

artifact value
cand_v9/{stories,code,math,wiki,html}.pt 5 ctx-1024 experts (code 1.53β†’1.12, html 1.47β†’0.33 val BPC)
cand_v9/cand_moe_v9.pt 5-expert candidate (balanced BPC 6.72 β†’ 1.53)
cand_v9/iter_v9_gate.json full gate verdict + metrics
cand_v9/benchmark_hf_v9.md real HF test-set benchmark (below)
cand_v9/*_history.json, html.manifest.json training provenance
session_v9_scripts/ session code (mixed-ctx trainer, corpora, verifier, bench)

Benchmark on real HF test sets (byte-BPC, lower=better; TinyStories val + WikiText-2 test)

model params TinyStories val WikiText-2 test
TinyStories-33M (matched) 33M 0.6287 3.62
SmolLM-135M (0-shot) 135M 0.9568 1.4457
gpt2-medium (0-shot) 355M 1.0430 1.4321
gpt2 (0-shot) 124M 1.1203 1.5317
pythia-70m (0-shot) 70M 1.2600 1.7224
moe_v9 (staged) ~190M/5x 0.7963 1.9053
moe_v7 (champion) ~114M/4x 0.8125 2.2305

moe_v9 beats the champion on every held-out corpus (incl. the champion's own). Specialist rows (held-out audit tails): html_v9 0.3607 on HTML (html_v8: 1.9095), code_v9 1.1524 on code β€” within 2.4% of SmolLM-135M zero-shot. Full protocol + raw log: cand_v9/benchmark_hf_v9.md.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support