|
Download gpu-sft/claude/evals.md from fzzhang/svd-code: direct link, hf CLI and curl.
- Browser
- Download file 1.66 kB
-
https://huggingface.co/fzzhang/svd-code/resolve/main/gpu-sft/claude/evals.md
- Command line
-
hf download hf://fzzhang/svd-code/gpu-sft/claude/evals.md
-
curl -L -o evals.md https://huggingface.co/fzzhang/svd-code/resolve/main/gpu-sft/claude/evals.md
1.66 kB
Local Evals
The compiler (claude/compile_results_local.py) parses ONLY the
## Models to Evaluate section below. Everything else in this file is prose for
humans.
Rules:
### Baselinesentries are the delta reference. One baseline per model size; the size token (0.6B,1.7B,4B,8B,32B) is read out of the name. A## Qwen3 <SIZE>section is emitted only if a baseline for that size exists.- Every other
### ...heading is an experiment group. The size comes from the heading text, falling back to the experiment name. - Each
- <name>under an experiment heading must be the EXACT experiment name, which is also the eval output directory prefix:<name>-step<N>/. Do not truncate, do not add a-step<N>suffix here. - Row order inside a suite table follows this file. The
#### Best Checkpointstable re-sorts by sampling compute (nofilterfirst, then_n<N>asc,_vr<N>asc,allvalidlast). - Suite routing is a substring match on the name:
science-> Science;code/coding/rstarcoder-> Code; otherwise Math. The suite decides which benchmarks are read AND which table the row lands in. - Section order across sizes is
int(re.sub(r"\D", "", size)), so 0.6B sorts as 6 and 1.7B as 17: 4B, 0.6B, 8B, 1.7B, 32B. Known quirk, kept for continuity with the existing results.md.
Models to Evaluate
Baselines
- Qwen/Qwen3-4B
- Qwen/Qwen3-8B
SFT 4B Experiments
- exp_sft_qwen3_4b_selfinstill_ot3_math53k_n8_vr5_round1
SFT 8B Experiments
- exp_sft_qwen3_8b_selfinstill_science_depth4v3_n8_vr5_2k_lr5e6_wd01
- exp_sft_qwen3_8b_selfinstill_coding_depth4v3_n8_vr5_2k_lr5e6_wd01