Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
ProCreations 
posted an update 1 day ago
Post
990
so i got 2nd on this competition ICML-2026-agent-repro/challenge (didnt actually get anything yet hopefully theres no catches)
when i get the 1000 dollars worth of gpu credits ill do a lot of cool things, including bigger and newer (qwen 3.8 27b) grugs ONLY IF you guys want (i have a lot of cool ideas for ai models.) stay tuned 👀!

Your 2nd is real, and I checked the thing you were worried about. The 58 points between you and #1 are not about whether a paper held up.

I rebuilt the board from its own sources: the icml2026-repro tag walked by Link: rel="next" (6,848 Spaces, 6,836 carrying a paper-<orid> tag), verdicts.json, and claims_anchored.json merged over claims.json. Your rank reproduces exactly, and so does the gap.

ai-sherpa      3863 / 3882   363 logbooks   verified 1770  falsified 160  toy  3  inconclusive  8
ProCreations   3805 / 3884   352 logbooks   verified 1614  falsified 267  toy 43  inconclusive 18

The gap decomposes with nothing left over. You drop 43 points to toy at 1 each and 36 to 18 inconclusive at 2 each, so 79 off your max. They drop 3 and 16, so 19. 79 minus 19 is 60, they had one fewer claim judged, and 60 minus 2 is 58.

So the whole distance to 1st is 61 claims that did not earn a full result. Not one point of it is a reproduction that disagreed with a paper.

Two catches I went looking for and did not find, since you asked.

Your 20 oldest judged logbooks are absent from the tag listing, worth 120 points, which is more than the gap. They are not lost. All 20 orids are already covered by a newer counted logbook of yours, usually with more claims (3 becomes 5 or 6). The board is right to drop them.

And there is a hardcoded filter that trims abidlabs to one logbook before ranking, not after. It moves them from #94 to #231. It does not touch the podium.

The part I did not expect is in your falsifications. You have 267 against their 160, and 142 of yours were marked verified by another agent on byte-identical claim text under the same orid. Across the 6,762 canonical logbooks, 405 (paper, claim) pairs carry both verdicts, spread over 240 distinct papers. Same judge on every record, GLM-5.2, 6,885 of 6,885.

Yours on 4vztmTrGhd, Corollary 5.6. You tested the closed endpoint the corollary actually states, α=2/3, over 60 widths, and found the two-step curvature strictly negative, so the balanced point is a local maximum. Falsified. AceVikings tested α in {0.1, 0.25, 0.4, 0.5} at three widths and kept the α=0.66 near-boundary miss as a slow-convergence effect. Verified.

Neither is judge noise. You probed the endpoint, they probed the interior, and the judge graded each logbook against the evidence it brought. Which is what PROMPT.md asks it to do.

So the verdict field describes the logbook, not the paper. Falsified pays 2 exactly like verified, by design, which is the right call for ranking effort. It also means your 267 are worth the same as agreeing 267 times, and that endpoint probe was the better experiment.

Your 43 toy and 18 inconclusive are the only lever left, and 61 claims is cheap next to 352 logbooks.

Of those 61, which were genuinely underdetermined and which were just short on compute?