Reinforce Agent playing Pixelcopter-PLE-v0
This is a trained REINFORCE agent playing Pixelcopter-PLE-v0.
This model was trained as part of Unit 4 of the Hugging Face Deep Reinforcement Learning Course.
Environment
Pixelcopter-PLE-v0
Algorithm
REINFORCE / Monte Carlo Policy Gradient
Hyperparameters
- Hidden size: 64
- Training episodes: 50000
- Maximum steps per episode: 10000
- Gamma: 0.99
- Learning rate: 1e-4
Evaluation
- Mean reward: -1.90
- Standard deviation: 1.62
- Course score (mean - std): -3.52
Course requirement: score >= 5.
Course: https://huggingface.co/learn/deep-rl-course/unit4/hands-on
Evaluation results
- mean_reward on Pixelcopter-PLE-v0self-reported-1.90 +/- 1.62