ProCreations's picture
Publish DiscoGen exact native reproduction
6fd091d verified
|
Raw
History Blame Contribute Delete
4.29 kB
# Judge-facing evidence scorecard
## Claim 1 β€” VERIFIED
> DiscoGen procedurally generates over 400 million distinct algorithm discovery tasks via a combinatorial formula N_tasks = 2*3*b*(2^m-1)*(3^d-2^(d+1)+1) depending on the number of modules m, datasets d, and backends b (Section 4.2, Equation 1).
**Result.** All ten reported rows match exactly and sum to 493,355,172; four exhaustive state enumerations match the formula, while omitting the nonempty train/test exclusion overcounts the control.
**Execution.** Parses all ten arXiv-v1 Table-1 domain parameters, executes the registered equation, and independently enumerates every valid module/dataset/backend/initialisation/evaluation state for four finite domains.
**Oracle.** The direct Cartesian-state enumerator shares no combinatorial closed form with the equation path; the released repository independently supplies the thirteen Model-Unlearning model choices.
**Control.** Deleting the train/test nonemptiness correction changes the m=3,d=4,b=2 count and is detected.
## Claim 2 β€” VERIFIED
> Including additional domains beyond the main evaluation set, DiscoGen's total task space reaches approximately 99 billion tasks (Appendix C).
**Result.** The fourteen rows sum exactly to 99,299,115,384, and the official repository independently contains 14 domains.
**Execution.** Parses every row of the expanded Appendix-C domain table and sums the exact task counts; separately inventories all released v1.0.0 domain configurations.
**Oracle.** The source-table summation and recursive official-repository inventory are independent paths over different artifacts.
**Control.** Removing On-Policy MARL collapses the total below two billion, decisively separating the 99-billion result.
## Claim 3 β€” VERIFIED
> Across the 10 domains used in the main evaluation, per-domain task counts range from 900 (Greenhouse Gas Prediction) to 426,043,800 (On-Policy RL), with a median of 59,622 tasks per domain (Table 1).
**Result.** The executed table has 10 domains, minimum 900 for Greenhouse Gas Prediction, maximum 426,043,800 for On-Policy RL, and median 59,622.
**Execution.** Parses the complete ten-domain main-evaluation table and independently computes its count, extrema, and order-statistic median.
**Oracle.** Python integer order statistics recompute the registered summary from all ten parsed primary rows rather than accepting the printed Total/Median lines.
**Control.** Dropping one domain changes the median and is detected.
## Claim 4 β€” VERIFIED
> DiscoBench provides a fixed evaluation subset built from DiscoGen, comprising, for each domain, m single-module tasks (DiscoBench Single) plus one comprehensive all-modules-active task (DiscoBench All) (Section 4.4).
**Result.** All 74 released m+1 configurations pass with zero failures; four official builds materialize 82 code/description files at tree digest 88d6c144caefdba3dfbb41aa7247ce4364a7e26928fe4010f51806321a29522c.
**Execution.** Executes the released v1.0.0 DiscoBench builder on four genuine configurations and checks every official single-module/all-modules YAML against its domain schema.
**Oracle.** The structural configuration checker independently validates the actual file trees created by the released MakeFiles implementation.
**Control.** Turning on a second module in an official single-module config is rejected by the one-active-module invariant.
## Claim 5 β€” VERIFIED
> As the number of editable modules increases in DiscoBench tasks, agent success rates consistently decline while the achievable performance ceiling rises (Appendix G).
**Result.** All three model success sequences are nonincreasing from one to four editable modules. The mean four-environment ceiling rises from 99.3475 to 108.3300 (+8.9825), with higher two-module maxima in 3/4 environments.
**Execution.** Parses the complete Appendix-G 15-configuration On-Policy-RL sweep, reconstructs all four module-combination levels, and recomputes success monotonicity and per-environment ceilings.
**Oracle.** The 15 combinations are regenerated independently from the four registered modules and matched to the complete primary result table.
**Control.** Reversing every model's success sequence makes the monotonic-decline test fail.