PPO LunarLander-v2

Stable-Baselines3 PPO agent trained on LunarLander-v2.

Evaluation

  • Mean reward: 255.26
  • Standard deviation: 22.00
  • Certification score (mean_reward - std_reward): 233.25
  • Evaluation episodes: 10
  • Training timesteps: 1,000,000

Training

This model was trained using Stable-Baselines3 PPO with an MLP policy.

Downloads last month
59
Video Preview
loading

Evaluation results