Download docs/05-eval-plan.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 11.1 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/05-eval-plan.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/docs/05-eval-plan.md
-
curl -L -o 05-eval-plan.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/05-eval-plan.md
STATUS 2026-09-21T14:30Z: NOT EXECUTED. The owner stopped Phase 6 (D-020) before any task produced a number, so no score exists anywhere in this project. Everything below stays as the frozen, pre-registered protocol — it is a specification for a future run, and must not be read as a result.
05 — Evaluation plan, pinned before any score exists
§5 Phase 6 says "pin the harness and per-task settings". The reason to write this file now, while no model has been trained, is §3.3: shot counts, prompts and parsers become a scoring lever the moment results exist. Everything here is a frozen choice with a reason, not a default I will tune later.
1. Harness pin (verified against PyPI today, not remembered)
| Item | Pin | Why |
|---|---|---|
| Harness | lm-eval 0.4.13 (EleutherAI lm-evaluation-harness) |
the reference implementation behind the published numbers of every ~100M-1B base model used as a target anchor in 01-plan.md §7; and the only one of the three candidate stacks with canonical prompts for all eight tasks (incl. PIQA and ARC-Easy) |
| Alternate, as a cross-check only | lighteval 0.13.0 |
if the two disagree, the discrepancy is reported, not averaged |
| Backend | model=hf (transformers), not vLLM |
vLLM's support for Turing (sm 7.5) at ≥0.18 is unverified on this hardware, and p0c already showed what an unverified stack assumption costs. A 106M model needs no serving optimisation: the whole eval set is minutes of fp16 forward passes on one card |
| Devices | one T4, no data parallelism | determinism beats speed here; two ranks add a gather path that can reorder few-shot examples |
| dtype | float16 |
matches training precision's storage; bf16 does not exist on sm 7.5 |
| Loading | AutoModelForCausalLM + AutoTokenizer, no --trust-remote-code |
§8 forbids custom code; if it needs trust-remote-code the export is wrong and Phase 5 has failed first |
| Chat template | none | base model, §3.9. A chat template would be an unearned capability |
| Model source | the public Hub repo, fresh instance, empty cache | the same clean-room condition as Gate 5 |
Command template (the exact string goes into the model repo with the results):
lm_eval --model hf \
--model_args pretrained=Cion-lab/ounce100m-<n>m,dtype=float16,trust_remote_code=False \
--tasks <task-ids> --batch_size 8 --seed 42 --verbosity INFO \
--output_path results/ --log_samples
2. Tasks, metrics and shot policy
Eight tasks, one row each, metric names as the harness reports them:
| Task | lm-eval task id (confirmed at run time through TaskManager().task_index; --tasks list is
not a listing command in 0.4.13 — E-048, measured on free CPU 2026-09-20 12:23Z) | Metric | Shots |
|---|---|---|---|
| ARC-Challenge | arc_challenge | acc | task default (25) |
| ARC-Easy | arc_easy | acc | task default (25) |
| HellaSwag | hellaswag | acc | task default |
| MMLU | mmlu | acc (macro over subjects) | task default |
| TruthfulQA | truthfulqa_mc2 (mc1 reported alongside) | acc,none | task default (0) |
| Winogrande | winogrande | acc | task default |
| PIQA | piqa | acc | task default |
| GSM8K | gsm8k (main, flexible-extract) | exact_match,flexible-extract | task default |
Shot policy, frozen. PRIMARY results use each task's harness default num_fewshot with no
--num_fewshot override anywhere in the run — the point of using the file in the repo rather than a
number I typed is that nobody chose it per task. A SECONDARY uniform 5-shot column is reported for every
task in the same run. Both columns ship; the headline is PRIMARY. Deciding this now means "re-run with
more shots until MMLU moves" is not available as an option later.
Seeds: --seed 42 fixed for the whole table, one invocation per task, all eight in a single job so a
partial run cannot be presented as a complete one. Incremental result files are pushed after each task
completes (§5 Phase 6: an interruption must not lose finished work).
3. Expected outcome, and how it will be reported honestly
The bands in 01-plan.md §7 / D-005 were set from base models of comparable size trained on 300 B
tokens; this model sees ~1 B. Several tasks are therefore expected to sit at or near chance:
MMLU (25 %), WinoGrande (50 %), ARC-C (25 %), GSM8K (~0 %), TruthfulQA mc2 low single digits. That is
the pre-registered prediction, and a table that lands there is a success, not a disappointment to be
engineered around. Anything materially above band gets a contamination re-check before it is published,
not a celebration — an unexpectedly high score is more likely to be a leak than a breakthrough.
Reporting rules, also frozen now:
- All eight, always, with metric name + shot count + the command above.
- No subset, no "best of two runs", no per-task parser swap.
- Chance level stated next to every score so a reader can see the signal, if any.
docs/05-eval-report.mdrecords deviations from this file verbatim, including any that were forced.
4. Contamination posture, restated at the point of use
Test splits are untouched until this phase (§3.3) and are read only by the harness, never by me. The mix
was audited mechanically against train/validation/dev material of these same tasks (build/audit_contamination.py,
13-token windows, counts only, no item text ever surfaced). The final report must carry that audit's
numbers next to these scores, because the two together are the only defensible statement about the
results.
5. Sources checked today
lm-eval0.4.13 andlighteval0.13.0 read from the PyPI JSON API at 18:17Z on 2026-09-19, not from memory (§3.6).lm-eval'shfextra declarestransformers>=4.1,accelerate>=0.26.0; the Kaggle image carries transformers 5.0.0 (measured in preflightpargs), so compatibility is asserted at run time by a--limit 5smoke test on all eight tasks before the full table is produced.
6. Amendment after Gate 3 (D-011): the model's context is 1024, and that is reported
The sequence length was changed from 2048 to 1024 on measured throughput (D-011), before the run was frozen. Consequence for this plan, recorded now so it cannot become an excuse later:
arc_challengeandarc_easyat their default 25 shots produce prompts longer than 1024 tokens. lm-eval truncates from the left for causal models. The score is still the harness's own metric on the harness's own configuration; it is simply a score for a 1024-context model.- No shot count is being changed in response. Every other task's prompt fits.
- If ARC's number looks odd next to a 2048-context anchor, the explanation is the context window, and it
belongs in the model card and
final-report.md, not in a re-tuned prompt.
Measured on free CPU 2026-09-20 (E-049, ounce100m-p6-cli-probe): the nine flags this file's command lines use all exist in lm-eval 0.4.13's CLI, and the metric names below are what the pinned YAMLs declare — but the results key is <name>,<aggregation> and the aggregation is not settled by reading configs. So each row's key is verified by the --limit 5 smoke pass before the full columns run, and a mismatch there is fixed in the constant, never by choosing a different metric after seeing a score.
7. Which split each task is actually scored on — read out of the pinned config, not assumed
Added 2026-09-19T20:25Z, before any score exists, while Gate 2's contamination finding was being
diagnosed. Every line below was read from lm_eval/tasks/** at tag v0.4.13 through the GitHub API
(file names listed from the tag, then the YAML itself) — not from memory and not from a README.
| task config | scored on | few-shot exemplars from | num_fewshot declared |
auditable by me? |
|---|---|---|---|---|
winogrande/default.yaml |
validation (test_split absent) |
train | none | yes — and it is in my reference set |
hellaswag/hellaswag.yaml |
validation (test_split: null) |
train | none | yes — in my reference set |
piqa/piqa.yaml |
validation (test_split: null) |
train | none | yes — was unreadable in Gate 2, fixed for 2c |
truthfulqa/truthfulqa_mc1/mc2 |
validation (test_split: null) |
none (0-shot) | 0 | yes — in my reference set |
arc/arc_easy.yaml, arc_challenge |
test | validation | none | partially: val is in my reference set, test is not |
mmlu/default/_default_template_yaml |
test | fewshot_split: dev |
none | partially: dev is in my reference set |
gsm8k/gsm8k.yaml |
test | train | 5 | partially: train is in my reference set |
Three consequences, recorded now so they cannot be argued later:
- Four of the eight tasks are scored on
validation— Winogrande, HellaSwag, PIQA and TruthfulQA. §3.3 lets me read validation but never test, so those four are the ones a contamination audit can fully cover, and they are also the ones where overlap directly inflates a headline number. ARC, MMLU and GSM8K are scored ontest, which stays unread until Phase 6: for them the audit covers exemplar material, not the scoring items. That asymmetry is a limitation of the method, not a choice about which results to check. - "Each task's own default shot count" (D-010) needs its definition pinned, and here it is: the
num_fewshotvalue in the task's own config at v0.4.13 if present, otherwise the harness default of 0. On the record before results: that makes TruthfulQA 0-shot by its own declaration, GSM8K 5-shot by its own declaration, and the remaining six 0-shot in the headline column — which is exactly why the pre-registered uniform 5-shot secondary column exists. The primary metric stays the harness's own (accuracy for the multiple-choice tasks,exact_matchfor GSM8K,mc2for TruthfulQA). - The upstream decontamination that the mix sources document is against test sets. FineMath's card
states: "Following Qwen2.5-Math's approach, we removed samples with 13-gram overlaps against test
sets from GSM8k, MATH, MMLU and ARC" (with logs published as
HuggingFaceTB/finemath_contamination_report), and Cosmopedia's card describes a 10-gram +difflibpass "against the test benchmarks". OpenWebMath's card documents SimHash deduplication only, with no benchmark decontamination claim. So the sources were cleaned for the splits I am forbidden to read, and left untouched for the four splits that Winogrande, HellaSwag, PIQA and TruthfulQA are actually scored on. That is consistent with Gate 2's measured hits and is the reason the audit found something real rather than imaginary: an overlap onvalidationis precisely what their filters were not looking for. Sources: FineMath card, Cosmopedia card, OpenWebMath card.