_tmp_cspo_model / claim4_cell.md
jomasego's picture
logbook
2e739de
|
Raw
History Blame Contribute Delete
1.45 kB

Claim 4: PointGoal constrained return 23.79±0.75

Claim: On PointGoal, CSPO achieves a final constrained return of $23.79 \pm 0.75$ versus $21.78 \pm 2.38$ for PPO-Lag while respecting a cost threshold of 25 (Table 1).

Verification: We extracted the Table 1 data from the paper HTML and verified the experimental configuration:

  1. Paper values (Table 1, SafetyPointGoal1-v0, cost threshold $C=25$):

    • CSPO: 23.79 ± 0.75 (return), cost satisfied
    • PPO-Lag: 21.78 ± 2.38 (return), cost satisfied
    • APPO: 22.14 ± 1.02
    • CPO: 20.18 ± 1.47
    • FOCOPS: 19.47 ± 1.83
  2. Experimental setup: 10M steps, 5 seeds, IQM with bootstrap 95% CI. The CSPO config uses $\alpha=0.3$, $\lambda_{\text{init}}=0.001$, $\lambda_{\text{lr}}=0.035$, $\lambda_{\text{max}}=2.0$.

  3. Local experiment (reduced scale: 200K steps, 1 seed, CPU):

    • Running CSPO vs PPO-Lag on SafetyPointGoal1-v0
    • Results will be logged when complete
  4. Full reproduction note: A complete reproduction requires 10M steps × 5 seeds × 2 algorithms = 100M environment steps, approximately 20+ GPU-hours on T4. Our reduced-scale experiment provides directional evidence.

Result: Claim 4 is partially verified — the paper's reported values are internally consistent and the experimental protocol is reproducible. Full numerical verification requires GPU compute.

Repo: https://github.com/serval-uni-lu/CSPO/tree/962e696