Buckets:
Experiment Spec: ICML 2026 Chebyshev Policies reproduction
Identity
- Experiment id:
icml2026_chebyshev_v1 - Spec version:
1.0.0 - OpenReview id:
aNWIVNjocB - arXiv input:
2605.22305v4(2026-07-08) - Challenge input: five anchored claims supplied on 2026-07-16
- Paper code commit:
fad6ebb46f976db29eb5ba509622dc3c24b7ac5f - Companion commit:
ed945331ce989592d75af9e8410823ee007e13a8
Research questions and claims
- Does the argument of Theorem 2.4 imply a loss-minimizing action field
alpha(t) = C * x_dot(t), independent of position? - On 100 evenly spaced deterministic Mountain Car starts in
[-0.6, -0.4], does the two-phase analytical policy achieve the reported return range and outperform the bundled ARS, PPO, and SAC policies by the reported margins? - Under the released 20-seed, 80,000-step protocol, does degree-3 CH-ARS attain the current v4 Table 1 performance? Are the paper's parameter count and reduction factor arithmetically and implementation-wise correct?
- Under the released protocols, do CH-3-PPO and CH-3-REINFORCE attain the current v4 Table 1 performance and regret improvements?
- Does the released Pendulum protocol reproduce CH-6-ARS
-150.8versus neural ARS-218.3on the 50 x 50 deterministic grid? The Aero 2 physical result is audited separately because the required hardware is unavailable.
Version-drift contract
The generated challenge Claim 3 quotes an older manuscript (regret 0.65,
4.18x), while arXiv v4 reports 0.44 and 6.18x. A generic judge claim also
quotes 268x, while the anchored claim and v4 say 277x. No value is silently
substituted: raw and derived outputs carry claim_source, paper_version, and
the reported comparator value. The logbook presents both the stale challenge
wording and the current paper result.
Experimental units
- Claims 1-2: one deterministic episode per policy and start position.
- Claims 3-4: one independently initialized training run per algorithm and seed, followed by one deterministic episode per trained policy and start.
- REINFORCE selection: one training run per seed, 50 seeded random-start evaluation episodes per trained policy, select exactly the highest mean, then evaluate that selected policy on the fixed 100-start grid.
- Claim 5: one training run per CH-6-ARS seed and one episode per policy and point on the fixed 2,500-state Pendulum grid; select the highest grid mean.
Systems and conditions
- Analytical two-phase Mountain Car policy with
C1=4.3346,C2=4.8358. - Released RL Baselines3 Zoo checkpoints: ARS, PPO, SAC.
- CH-3-ARS: degree 3, learning rate
0.018,delta_std=0.1,n_delta=4,n_top=1, 80,000 steps, seeds0, 123, ..., 2337. - CH-3-PPO: degree 3, learning rate
0.001,n_steps=2048, batch size1048, 2 epochs, clip ranges0.4, 70,000 steps, the same 20 seeds. - CH-3-REINFORCE: AdamW, degree 3, 100 episodes, discount
0.9, learning rates0.0003for mean and0.00003for sigma, 20 explicit seeds, then 50 evaluation episodes per policy. - Pendulum CH-6-ARS: degree 6, learning rate
0.018,delta_std=0.3,n_delta=8, all deltas retained, 2,000,000 steps, 30 released-protocol seeds, evaluated on 50 angles x 50 angular velocities.
Primary metrics
- Episode return, with mean, standard deviation, minimum, and maximum.
- Regret: analytical-policy mean return minus evaluated-policy mean return.
- Regret improvement: neural baseline regret divided by Chebyshev regret.
- Trainable scalar count derived from executable model state and from the architecture stated in the paper; both sources remain separate.
Secondary metrics
- Episode length, terminal velocity, wall time, peak resident memory.
- Per-start return and per-seed outcome distributions.
- Training checkpoints/evaluation curves when emitted by the algorithm.
- Policy action values on a fixed position-velocity grid for visualization.
Panels and selection rules
mountaincar_primary_v1: exactly 100 evenly spaced starts in[-0.6, -0.4]; all successful rows included.training_all_seeds_v1: every planned seed, including failures.best_by_fixed_grid_mean_v1: select the trained policy with the largest mean on the predeclared fixed evaluation grid. Never select by a single episode or by the paper target.pendulum_primary_v1: Cartesian product of 50 angles in[-pi, pi]and 50 angular velocities in[-1, 1].
Success, failure, and exclusions
- A full empirical reproduction requires complete planned row coverage and a Hugging Face GPU Job. Any missing seed/start is a failed panel, not silently dropped.
- A run is successful only if the process exits zero, all metrics are finite, and the expected output/checkpoint count is exact.
- No retries are automatic. A retry receives a new
batch_idandattempt_id; the original failure remains raw evidence. - Local smoke profiles are explicitly
toyand cannot support a full claim. - Performance is considered numerically reproduced when the deterministic
mean is within
0.05return of the current paper value and the reported ordering holds. This tolerance was fixed before reportable collection. - The historical priority phrase "first" and the physical Aero 2 result are not inferred from simulation. The former is a literature claim; the latter remains unverified without the hardware.
Raw evidence
Each unique runs/raw/<batch-id>/ contains records.jsonl, manifest.json,
environment metadata, stdout/stderr captured by Trackio/HF Jobs, and newly
trained coefficients. Raw directories are never overwritten.
Canonical tables
runs/tables/episodes.csv: one row per evaluated episode/start.runs/tables/training_runs.csv: one row per training seed/attempt.runs/tables/summary.csv: one row per method and declared panel.runs/tables/parameter_audit.csv: one row per parameter-count source.runs/tables/MANIFEST.json: hashes, row counts, source batches, command.
Planned analysis outputs
- Interactive position-velocity policy microscope with trajectories.
- Per-start return curves for analytical, neural, and Chebyshev policies.
- All-seed strip/violin distributions with the selected seed identified.
- Parameter-count versus regret Pareto chart with count-source annotations.
- Pendulum initial-state return heatmaps and paired method differences.
- Claim-evidence map linking every conclusion to tables, scripts, and raw rows.
Hardware and budget
Local execution is only a smoke test. Full empirical runs use Hugging Face
t4-medium Jobs (8 vCPU, 30 GB system RAM, one 16 GB T4) to stay below the
32 GB requirement. The released implementations are CPU-bound; GPU presence
is recorded honestly and no CUDA speedup is claimed. Each Job has a fixed
timeout no greater than 24 hours.
Known limitations
- The two author repositories have no declared license, so they are fetched at pinned commits rather than republished.
- The released notebooks rely on an uncommitted PickleShare database. This reproduction executes source APIs from clean configs and does not treat notebook outputs as measurements.
- The physical Quanser Aero 2 experiment cannot be recreated without access to that system. Simulation is not substituted for the real-hardware claim.
- Theorem 2.4 is addressed with a proof audit and numerical consequences; a numerical experiment alone cannot establish historical priority.
Revision log
1.0.0(2026-07-16): Initial pre-collection specification. No reportable reproduction results had been collected in this repository.
Xet Storage Details
- Size:
- 7.64 kB
- Xet hash:
- f4d039f00f2d10185c52ea02046895a0b3853b3e06d13e234e513e3bd1013f57
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.