JacobLinCool's picture
|
download
raw
7.64 kB

Experiment Spec: ICML 2026 Chebyshev Policies reproduction

Identity

  • Experiment id: icml2026_chebyshev_v1
  • Spec version: 1.0.0
  • OpenReview id: aNWIVNjocB
  • arXiv input: 2605.22305v4 (2026-07-08)
  • Challenge input: five anchored claims supplied on 2026-07-16
  • Paper code commit: fad6ebb46f976db29eb5ba509622dc3c24b7ac5f
  • Companion commit: ed945331ce989592d75af9e8410823ee007e13a8

Research questions and claims

  1. Does the argument of Theorem 2.4 imply a loss-minimizing action field alpha(t) = C * x_dot(t), independent of position?
  2. On 100 evenly spaced deterministic Mountain Car starts in [-0.6, -0.4], does the two-phase analytical policy achieve the reported return range and outperform the bundled ARS, PPO, and SAC policies by the reported margins?
  3. Under the released 20-seed, 80,000-step protocol, does degree-3 CH-ARS attain the current v4 Table 1 performance? Are the paper's parameter count and reduction factor arithmetically and implementation-wise correct?
  4. Under the released protocols, do CH-3-PPO and CH-3-REINFORCE attain the current v4 Table 1 performance and regret improvements?
  5. Does the released Pendulum protocol reproduce CH-6-ARS -150.8 versus neural ARS -218.3 on the 50 x 50 deterministic grid? The Aero 2 physical result is audited separately because the required hardware is unavailable.

Version-drift contract

The generated challenge Claim 3 quotes an older manuscript (regret 0.65, 4.18x), while arXiv v4 reports 0.44 and 6.18x. A generic judge claim also quotes 268x, while the anchored claim and v4 say 277x. No value is silently substituted: raw and derived outputs carry claim_source, paper_version, and the reported comparator value. The logbook presents both the stale challenge wording and the current paper result.

Experimental units

  • Claims 1-2: one deterministic episode per policy and start position.
  • Claims 3-4: one independently initialized training run per algorithm and seed, followed by one deterministic episode per trained policy and start.
  • REINFORCE selection: one training run per seed, 50 seeded random-start evaluation episodes per trained policy, select exactly the highest mean, then evaluate that selected policy on the fixed 100-start grid.
  • Claim 5: one training run per CH-6-ARS seed and one episode per policy and point on the fixed 2,500-state Pendulum grid; select the highest grid mean.

Systems and conditions

  • Analytical two-phase Mountain Car policy with C1=4.3346, C2=4.8358.
  • Released RL Baselines3 Zoo checkpoints: ARS, PPO, SAC.
  • CH-3-ARS: degree 3, learning rate 0.018, delta_std=0.1, n_delta=4, n_top=1, 80,000 steps, seeds 0, 123, ..., 2337.
  • CH-3-PPO: degree 3, learning rate 0.001, n_steps=2048, batch size 1048, 2 epochs, clip ranges 0.4, 70,000 steps, the same 20 seeds.
  • CH-3-REINFORCE: AdamW, degree 3, 100 episodes, discount 0.9, learning rates 0.0003 for mean and 0.00003 for sigma, 20 explicit seeds, then 50 evaluation episodes per policy.
  • Pendulum CH-6-ARS: degree 6, learning rate 0.018, delta_std=0.3, n_delta=8, all deltas retained, 2,000,000 steps, 30 released-protocol seeds, evaluated on 50 angles x 50 angular velocities.

Primary metrics

  • Episode return, with mean, standard deviation, minimum, and maximum.
  • Regret: analytical-policy mean return minus evaluated-policy mean return.
  • Regret improvement: neural baseline regret divided by Chebyshev regret.
  • Trainable scalar count derived from executable model state and from the architecture stated in the paper; both sources remain separate.

Secondary metrics

  • Episode length, terminal velocity, wall time, peak resident memory.
  • Per-start return and per-seed outcome distributions.
  • Training checkpoints/evaluation curves when emitted by the algorithm.
  • Policy action values on a fixed position-velocity grid for visualization.

Panels and selection rules

  • mountaincar_primary_v1: exactly 100 evenly spaced starts in [-0.6, -0.4]; all successful rows included.
  • training_all_seeds_v1: every planned seed, including failures.
  • best_by_fixed_grid_mean_v1: select the trained policy with the largest mean on the predeclared fixed evaluation grid. Never select by a single episode or by the paper target.
  • pendulum_primary_v1: Cartesian product of 50 angles in [-pi, pi] and 50 angular velocities in [-1, 1].

Success, failure, and exclusions

  • A full empirical reproduction requires complete planned row coverage and a Hugging Face GPU Job. Any missing seed/start is a failed panel, not silently dropped.
  • A run is successful only if the process exits zero, all metrics are finite, and the expected output/checkpoint count is exact.
  • No retries are automatic. A retry receives a new batch_id and attempt_id; the original failure remains raw evidence.
  • Local smoke profiles are explicitly toy and cannot support a full claim.
  • Performance is considered numerically reproduced when the deterministic mean is within 0.05 return of the current paper value and the reported ordering holds. This tolerance was fixed before reportable collection.
  • The historical priority phrase "first" and the physical Aero 2 result are not inferred from simulation. The former is a literature claim; the latter remains unverified without the hardware.

Raw evidence

Each unique runs/raw/<batch-id>/ contains records.jsonl, manifest.json, environment metadata, stdout/stderr captured by Trackio/HF Jobs, and newly trained coefficients. Raw directories are never overwritten.

Canonical tables

  • runs/tables/episodes.csv: one row per evaluated episode/start.
  • runs/tables/training_runs.csv: one row per training seed/attempt.
  • runs/tables/summary.csv: one row per method and declared panel.
  • runs/tables/parameter_audit.csv: one row per parameter-count source.
  • runs/tables/MANIFEST.json: hashes, row counts, source batches, command.

Planned analysis outputs

  • Interactive position-velocity policy microscope with trajectories.
  • Per-start return curves for analytical, neural, and Chebyshev policies.
  • All-seed strip/violin distributions with the selected seed identified.
  • Parameter-count versus regret Pareto chart with count-source annotations.
  • Pendulum initial-state return heatmaps and paired method differences.
  • Claim-evidence map linking every conclusion to tables, scripts, and raw rows.

Hardware and budget

Local execution is only a smoke test. Full empirical runs use Hugging Face t4-medium Jobs (8 vCPU, 30 GB system RAM, one 16 GB T4) to stay below the 32 GB requirement. The released implementations are CPU-bound; GPU presence is recorded honestly and no CUDA speedup is claimed. Each Job has a fixed timeout no greater than 24 hours.

Known limitations

  • The two author repositories have no declared license, so they are fetched at pinned commits rather than republished.
  • The released notebooks rely on an uncommitted PickleShare database. This reproduction executes source APIs from clean configs and does not treat notebook outputs as measurements.
  • The physical Quanser Aero 2 experiment cannot be recreated without access to that system. Simulation is not substituted for the real-hardware claim.
  • Theorem 2.4 is addressed with a proof audit and numerical consequences; a numerical experiment alone cannot establish historical priority.

Revision log

  • 1.0.0 (2026-07-16): Initial pre-collection specification. No reportable reproduction results had been collected in this repository.

Xet Storage Details

Size:
7.64 kB
·
Xet hash:
f4d039f00f2d10185c52ea02046895a0b3853b3e06d13e234e513e3bd1013f57

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.