--- library_name: ml-agents tags: - reinforcement-learning - deep-reinforcement-learning - Pyramids - ML-Agents-Pyramids model-index: - name: ppo-Pyramids results: - task: type: reinforcement-learning name: Reinforcement Learning dataset: name: ML-Agents-Pyramids type: ML-Agents-Pyramids metrics: - name: mean_reward type: mean_reward value: -0.9999999310821295 +/- 0.0 --- # PPO agent playing Pyramids Trained from scratch for 63,991 steps in Google Colab using Unity ML-Agents, with coding and execution assistance from Codex. This is an introductory course model, not a fully converged policy. ## Evaluation Evaluated using the exported ONNX policy with deterministic actions over 20 completed agent episodes, seed 12345. Mean reward: -0.9999999310821295; standard deviation: 0.0. Soccer rewards include the team reward. Full episode returns are in evaluation.json. ## Files The ONNX file is the evaluated policy. configuration.yaml and config.json contain training settings. checkpoint.pt permits continued training; training.log records the actual run. ## References - https://huggingface.co/learn/deep-rl-course/en/unit5/hands-on - https://huggingface.co/learn/deep-rl-course/en/unit7/hands-on - https://github.com/Unity-Technologies/ml-agents