ppo-Pyramids / README.md
rondahahda's picture
Upload 7 files
9e3ffe2 verified
|
Raw History Blame Contribute Delete
1.31 kB
metadata
library_name: ml-agents
tags:
  - reinforcement-learning
  - deep-reinforcement-learning
  - Pyramids
  - ML-Agents-Pyramids
model-index:
  - name: ppo-Pyramids
    results:
      - task:
          type: reinforcement-learning
          name: Reinforcement Learning
        dataset:
          name: ML-Agents-Pyramids
          type: ML-Agents-Pyramids
        metrics:
          - name: mean_reward
            type: mean_reward
            value: '-0.9999999310821295 +/- 0.0'

PPO agent playing Pyramids

Trained from scratch for 63,991 steps in Google Colab using Unity ML-Agents, with coding and execution assistance from Codex. This is an introductory course model, not a fully converged policy.

Evaluation

Evaluated using the exported ONNX policy with deterministic actions over 20 completed agent episodes, seed 12345. Mean reward: -0.9999999310821295; standard deviation: 0.0. Soccer rewards include the team reward. Full episode returns are in evaluation.json.

Files

The ONNX file is the evaluated policy. configuration.yaml and config.json contain training settings. checkpoint.pt permits continued training; training.log records the actual run.

References