Commit History

run_benchmarks: publish the effective shot count with its provenance, so a 0-shot primary row can never be read as few-shot
f6d90c6
verified

Cion-lab commited on

run_benchmarks: the 5-row smoke also verifies PRIMARY_METRIC against real results keys, so a wrong metric name cannot survive to the full columns
fbe96e3
verified

Cion-lab commited on

run_benchmarks: install with deps, ask the task registry via the API, refuse a row whose primary metric key is absent (E-047/E-048)
ad9e5a2
verified

Cion-lab commited on

run_benchmarks: real sample row counts, no abort on legitimate null shots, verdict honours status (E-046)
972b64f
verified

Cion-lab commited on

run_benchmarks: drop the unused binding from the transformers assertion
c66bf84
verified

Cion-lab commited on

run_benchmarks: task-list hard fail, per-task shot/denominator provenance, primary metric pre-registered, failed cells not zeros, push verified anonymously (E-045)
1c6dfb9
verified

Cion-lab commited on

review round 3 (probe/model-card/eval): rank-visible eval failure + cross-rank peak memory + complete run_summary; card derived-and-asserted from config; harness install --no-deps and task ids confirmed; torchrun redirects 3 + tee 3
f45605d
verified

Cion-lab commited on

Phase 6 tooling: the eight benchmarks exactly as docs/05-eval-plan.md froze them (lm-eval 0.4.13, hf backend, fp16, one card, task-default shots + uniform 5-shot column, --limit 5 smoke gate first, incremental push after every task)
8b75198
verified

Cion-lab commited on