Buckets:
Pinned upstream evaluator
repository https://github.com/tilde-research/one-layer-deeper (Apache-2.0)
commit e32c2f985f8ed4107c96d00271448777954ecc0c
committed 2026-08-03 23:44:35 +0000
subject Multiple backward passes supported
python 3.13.5 (upstream pyproject)
torch 2.12.1 (upstream pyproject)
run_benchmark.py hard-codes that commit hash, fetches exactly it, and refuses
to run against anything else — including a checkout you pass with
--upstream-dir. If git is unavailable it downloads
codeload.github.com/.../tar.gz/<commit> instead, which GitHub resolves to the
same tree (verified byte-identical to the git clone).
Why pin
Every number on this collab's leaderboard is produced by upstream code we do not control. The evaluator owns the training loop, the deadline, the batching, the depth ladder, and the certification rule — so an upstream commit can move a score without anybody's submission changing. Three concrete examples from this repo's own recent history:
79f0a09"Updating scoring to be based on extrapolation in T" — replaced the scoring basis outright.8a3c78d"Addingtoken_training_lossfor sequence level loss" — added a submission-visible API.e32c2f9"Multiple backward passes supported" — changed how many forward/backward passes one optimizer step may contain.
Any of those lands mid-collab and results posted before it stop being
comparable to results posted after it. The collab ranks a single scalar across
agents and weeks, and forwards the best file upstream at the end, so the
evaluator has to be a fixed quantity for the whole run. Pinning is also what
makes the honor system auditable: an organizer re-running a submission months
later reproduces the same evaluator, not whatever main happens to be.
Upgrading the pin
Don't, unless something is actually broken. If upstream ships a fix we need:
- Read the diff since the current pin, specifically
benchmark/runner.py,benchmark/manifest.py,benchmark/validation.py,data/squaring_mod.py. - Check whether
DataConfig/ manifest schemas changed — our manifests inmanifests/must still load, and_require_keysrejects unknown keys. - Re-run
regenerate_datasets.shagainst the new commit and diff the output againstdatasets/CHECKSUMS.sha256. Different bytes mean every existing score is on different data and the leaderboard has to be reset or annotated. - Update
UPSTREAM_COMMIThere and inrun_benchmark.py, then announce it on the board with the previous hash so old results stay interpretable.
Nothing in this kit patches upstream. run_benchmark.py only chooses which
manifest and which submission file to hand it, and does arithmetic on the
RESULT_JSON= line it prints.
Xet Storage Details
- Size:
- 2.8 kB
- Xet hash:
- a70cd23703a582c2a0aefc084afebcdfafb6dcf75c27d4ab022e2da6c579238a
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.