# Local Evals The compiler (`claude/compile_results_local.py`) parses ONLY the `## Models to Evaluate` section below. Everything else in this file is prose for humans. Rules: - `### Baselines` entries are the delta reference. One baseline per model size; the size token (`0.6B`, `1.7B`, `4B`, `8B`, `32B`) is read out of the name. A `## Qwen3 ` section is emitted only if a baseline for that size exists. - Every other `### ...` heading is an experiment group. The size comes from the heading text, falling back to the experiment name. - Each `- ` under an experiment heading must be the EXACT experiment name, which is also the eval output directory prefix: `-step/`. Do not truncate, do not add a `-step` suffix here. - Row order inside a suite table follows this file. The `#### Best Checkpoints` table re-sorts by sampling compute (`nofilter` first, then `_n` asc, `_vr` asc, `allvalid` last). - Suite routing is a substring match on the name: `science` -> Science; `code` / `coding` / `rstarcoder` -> Code; otherwise Math. The suite decides which benchmarks are read AND which table the row lands in. - Section order across sizes is `int(re.sub(r"\D", "", size))`, so 0.6B sorts as 6 and 1.7B as 17: 4B, 0.6B, 8B, 1.7B, 32B. Known quirk, kept for continuity with the existing results.md. ## Models to Evaluate ### Baselines - Qwen/Qwen3-4B - Qwen/Qwen3-8B ### SFT 4B Experiments - exp_sft_qwen3_4b_selfinstill_ot3_math53k_n8_vr5_round1 ### SFT 8B Experiments - exp_sft_qwen3_8b_selfinstill_science_depth4v3_n8_vr5_2k_lr5e6_wd01 - exp_sft_qwen3_8b_selfinstill_coding_depth4v3_n8_vr5_2k_lr5e6_wd01