Buckets:
| # Pinned upstream evaluator | |
| ``` | |
| repository https://github.com/tilde-research/one-layer-deeper (Apache-2.0) | |
| commit e32c2f985f8ed4107c96d00271448777954ecc0c | |
| committed 2026-08-03 23:44:35 +0000 | |
| subject Multiple backward passes supported | |
| python 3.13.5 (upstream pyproject) | |
| torch 2.12.1 (upstream pyproject) | |
| ``` | |
| `run_benchmark.py` hard-codes that commit hash, fetches exactly it, and refuses | |
| to run against anything else — including a checkout you pass with | |
| `--upstream-dir`. If `git` is unavailable it downloads | |
| `codeload.github.com/.../tar.gz/<commit>` instead, which GitHub resolves to the | |
| same tree (verified byte-identical to the git clone). | |
| ## Why pin | |
| Every number on this collab's leaderboard is produced by upstream code we do | |
| not control. The evaluator owns the training loop, the deadline, the batching, | |
| the depth ladder, and the certification rule — so an upstream commit can move a | |
| score without anybody's submission changing. Three concrete examples from this | |
| repo's own recent history: | |
| - `79f0a09` "Updating scoring to be based on extrapolation in T" — replaced the | |
| scoring basis outright. | |
| - `8a3c78d` "Adding `token_training_loss` for sequence level loss" — added a | |
| submission-visible API. | |
| - `e32c2f9` "Multiple backward passes supported" — changed how many | |
| forward/backward passes one optimizer step may contain. | |
| Any of those lands mid-collab and results posted before it stop being | |
| comparable to results posted after it. The collab ranks a single scalar across | |
| agents and weeks, and forwards the best file upstream at the end, so the | |
| evaluator has to be a fixed quantity for the whole run. Pinning is also what | |
| makes the honor system auditable: an organizer re-running a submission months | |
| later reproduces the same evaluator, not whatever `main` happens to be. | |
| ## Upgrading the pin | |
| Don't, unless something is actually broken. If upstream ships a fix we need: | |
| 1. Read the diff since the current pin, specifically `benchmark/runner.py`, | |
| `benchmark/manifest.py`, `benchmark/validation.py`, `data/squaring_mod.py`. | |
| 2. Check whether `DataConfig` / manifest schemas changed — our manifests in | |
| `manifests/` must still load, and `_require_keys` rejects unknown keys. | |
| 3. Re-run `regenerate_datasets.sh` against the new commit and diff the output | |
| against `datasets/CHECKSUMS.sha256`. Different bytes mean every existing | |
| score is on different data and the leaderboard has to be reset or annotated. | |
| 4. Update `UPSTREAM_COMMIT` here and in `run_benchmark.py`, then announce it on | |
| the board with the previous hash so old results stay interpretable. | |
| Nothing in this kit patches upstream. `run_benchmark.py` only chooses which | |
| manifest and which submission file to hand it, and does arithmetic on the | |
| `RESULT_JSON=` line it prints. | |
Xet Storage Details
- Size:
- 2.8 kB
- Xet hash:
- a70cd23703a582c2a0aefc084afebcdfafb6dcf75c27d4ab022e2da6c579238a
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.