ounce100m-code / docs /05-eval-plan.md
Cion-lab's picture
publish docs/: plan, mix rationale, preflight report, run log, frozen eval protocol, final report
f345921 verified
|
Raw History Blame Contribute Delete
11.1 kB

STATUS 2026-09-21T14:30Z: NOT EXECUTED. The owner stopped Phase 6 (D-020) before any task produced a number, so no score exists anywhere in this project. Everything below stays as the frozen, pre-registered protocol — it is a specification for a future run, and must not be read as a result.

05 — Evaluation plan, pinned before any score exists

§5 Phase 6 says "pin the harness and per-task settings". The reason to write this file now, while no model has been trained, is §3.3: shot counts, prompts and parsers become a scoring lever the moment results exist. Everything here is a frozen choice with a reason, not a default I will tune later.

1. Harness pin (verified against PyPI today, not remembered)

Item Pin Why
Harness lm-eval 0.4.13 (EleutherAI lm-evaluation-harness) the reference implementation behind the published numbers of every ~100M-1B base model used as a target anchor in 01-plan.md §7; and the only one of the three candidate stacks with canonical prompts for all eight tasks (incl. PIQA and ARC-Easy)
Alternate, as a cross-check only lighteval 0.13.0 if the two disagree, the discrepancy is reported, not averaged
Backend model=hf (transformers), not vLLM vLLM's support for Turing (sm 7.5) at ≥0.18 is unverified on this hardware, and p0c already showed what an unverified stack assumption costs. A 106M model needs no serving optimisation: the whole eval set is minutes of fp16 forward passes on one card
Devices one T4, no data parallelism determinism beats speed here; two ranks add a gather path that can reorder few-shot examples
dtype float16 matches training precision's storage; bf16 does not exist on sm 7.5
Loading AutoModelForCausalLM + AutoTokenizer, no --trust-remote-code §8 forbids custom code; if it needs trust-remote-code the export is wrong and Phase 5 has failed first
Chat template none base model, §3.9. A chat template would be an unearned capability
Model source the public Hub repo, fresh instance, empty cache the same clean-room condition as Gate 5

Command template (the exact string goes into the model repo with the results):

lm_eval --model hf \
  --model_args pretrained=Cion-lab/ounce100m-<n>m,dtype=float16,trust_remote_code=False \
  --tasks <task-ids> --batch_size 8 --seed 42 --verbosity INFO \
  --output_path results/ --log_samples

2. Tasks, metrics and shot policy

Eight tasks, one row each, metric names as the harness reports them:

| Task | lm-eval task id (confirmed at run time through TaskManager().task_index; --tasks list is

not a listing command in 0.4.13 — E-048, measured on free CPU 2026-09-20 12:23Z) | Metric | Shots | |---|---|---|---| | ARC-Challenge | arc_challenge | acc | task default (25) | | ARC-Easy | arc_easy | acc | task default (25) | | HellaSwag | hellaswag | acc | task default | | MMLU | mmlu | acc (macro over subjects) | task default | | TruthfulQA | truthfulqa_mc2 (mc1 reported alongside) | acc,none | task default (0) | | Winogrande | winogrande | acc | task default | | PIQA | piqa | acc | task default | | GSM8K | gsm8k (main, flexible-extract) | exact_match,flexible-extract | task default |

Shot policy, frozen. PRIMARY results use each task's harness default num_fewshot with no --num_fewshot override anywhere in the run — the point of using the file in the repo rather than a number I typed is that nobody chose it per task. A SECONDARY uniform 5-shot column is reported for every task in the same run. Both columns ship; the headline is PRIMARY. Deciding this now means "re-run with more shots until MMLU moves" is not available as an option later.

Seeds: --seed 42 fixed for the whole table, one invocation per task, all eight in a single job so a partial run cannot be presented as a complete one. Incremental result files are pushed after each task completes (§5 Phase 6: an interruption must not lose finished work).

3. Expected outcome, and how it will be reported honestly

The bands in 01-plan.md §7 / D-005 were set from base models of comparable size trained on 300 B tokens; this model sees ~1 B. Several tasks are therefore expected to sit at or near chance: MMLU (25 %), WinoGrande (50 %), ARC-C (25 %), GSM8K (~0 %), TruthfulQA mc2 low single digits. That is the pre-registered prediction, and a table that lands there is a success, not a disappointment to be engineered around. Anything materially above band gets a contamination re-check before it is published, not a celebration — an unexpectedly high score is more likely to be a leak than a breakthrough.

Reporting rules, also frozen now:

  1. All eight, always, with metric name + shot count + the command above.
  2. No subset, no "best of two runs", no per-task parser swap.
  3. Chance level stated next to every score so a reader can see the signal, if any.
  4. docs/05-eval-report.md records deviations from this file verbatim, including any that were forced.

4. Contamination posture, restated at the point of use

Test splits are untouched until this phase (§3.3) and are read only by the harness, never by me. The mix was audited mechanically against train/validation/dev material of these same tasks (build/audit_contamination.py, 13-token windows, counts only, no item text ever surfaced). The final report must carry that audit's numbers next to these scores, because the two together are the only defensible statement about the results.

5. Sources checked today

  • lm-eval 0.4.13 and lighteval 0.13.0 read from the PyPI JSON API at 18:17Z on 2026-09-19, not from memory (§3.6). lm-eval's hf extra declares transformers>=4.1, accelerate>=0.26.0; the Kaggle image carries transformers 5.0.0 (measured in preflight pargs), so compatibility is asserted at run time by a --limit 5 smoke test on all eight tasks before the full table is produced.

6. Amendment after Gate 3 (D-011): the model's context is 1024, and that is reported

The sequence length was changed from 2048 to 1024 on measured throughput (D-011), before the run was frozen. Consequence for this plan, recorded now so it cannot become an excuse later:

  • arc_challenge and arc_easy at their default 25 shots produce prompts longer than 1024 tokens. lm-eval truncates from the left for causal models. The score is still the harness's own metric on the harness's own configuration; it is simply a score for a 1024-context model.
  • No shot count is being changed in response. Every other task's prompt fits.
  • If ARC's number looks odd next to a 2048-context anchor, the explanation is the context window, and it belongs in the model card and final-report.md, not in a re-tuned prompt.

Measured on free CPU 2026-09-20 (E-049, ounce100m-p6-cli-probe): the nine flags this file's command lines use all exist in lm-eval 0.4.13's CLI, and the metric names below are what the pinned YAMLs declare — but the results key is <name>,<aggregation> and the aggregation is not settled by reading configs. So each row's key is verified by the --limit 5 smoke pass before the full columns run, and a mismatch there is fixed in the constant, never by choosing a different metric after seeing a score.

7. Which split each task is actually scored on — read out of the pinned config, not assumed

Added 2026-09-19T20:25Z, before any score exists, while Gate 2's contamination finding was being diagnosed. Every line below was read from lm_eval/tasks/** at tag v0.4.13 through the GitHub API (file names listed from the tag, then the YAML itself) — not from memory and not from a README.

task config scored on few-shot exemplars from num_fewshot declared auditable by me?
winogrande/default.yaml validation (test_split absent) train none yes — and it is in my reference set
hellaswag/hellaswag.yaml validation (test_split: null) train none yes — in my reference set
piqa/piqa.yaml validation (test_split: null) train none yes — was unreadable in Gate 2, fixed for 2c
truthfulqa/truthfulqa_mc1/mc2 validation (test_split: null) none (0-shot) 0 yes — in my reference set
arc/arc_easy.yaml, arc_challenge test validation none partially: val is in my reference set, test is not
mmlu/default/_default_template_yaml test fewshot_split: dev none partially: dev is in my reference set
gsm8k/gsm8k.yaml test train 5 partially: train is in my reference set

Three consequences, recorded now so they cannot be argued later:

  1. Four of the eight tasks are scored on validation — Winogrande, HellaSwag, PIQA and TruthfulQA. §3.3 lets me read validation but never test, so those four are the ones a contamination audit can fully cover, and they are also the ones where overlap directly inflates a headline number. ARC, MMLU and GSM8K are scored on test, which stays unread until Phase 6: for them the audit covers exemplar material, not the scoring items. That asymmetry is a limitation of the method, not a choice about which results to check.
  2. "Each task's own default shot count" (D-010) needs its definition pinned, and here it is: the num_fewshot value in the task's own config at v0.4.13 if present, otherwise the harness default of 0. On the record before results: that makes TruthfulQA 0-shot by its own declaration, GSM8K 5-shot by its own declaration, and the remaining six 0-shot in the headline column — which is exactly why the pre-registered uniform 5-shot secondary column exists. The primary metric stays the harness's own (accuracy for the multiple-choice tasks, exact_match for GSM8K, mc2 for TruthfulQA).
  3. The upstream decontamination that the mix sources document is against test sets. FineMath's card states: "Following Qwen2.5-Math's approach, we removed samples with 13-gram overlaps against test sets from GSM8k, MATH, MMLU and ARC" (with logs published as HuggingFaceTB/finemath_contamination_report), and Cosmopedia's card describes a 10-gram + difflib pass "against the test benchmarks". OpenWebMath's card documents SimHash deduplication only, with no benchmark decontamination claim. So the sources were cleaned for the splits I am forbidden to read, and left untouched for the four splits that Winogrande, HellaSwag, PIQA and TruthfulQA are actually scored on. That is consistent with Gate 2's measured hits and is the reason the audit found something real rather than imaginary: an overlap on validation is precisely what their filters were not looking for. Sources: FineMath card, Cosmopedia card, OpenWebMath card.