PPO on ALE/Breakout-v5 (from scratch, no SB3)
Custom PyTorch implementation of REINFORCE → A2C → A2C+GAE → PPO, built up incrementally for a coursework assignment. This checkpoint is the final PPO stage.
Environment
ALE/Breakout-v5,AtariPreprocessing(frameskip=4, grayscale, 84x84),FrameStackObservation(stack_size=4)
Architecture
Nature-DQN-style CNN encoder (3 conv layers → 512-d FC) shared between actor and critic heads.
Hyperparameters
- gamma=0.99, lambda=0.95, clip_eps=0.2
- lr=2.5e-4 (Adam), rollout_len=128, n_epochs=4, minibatch_size=32
- value loss coef=0.5, entropy coef=0.01, grad clip=0.5
Evaluation
20-episode greedy evaluation: mean=2.35, std=0.48
Training progression
| Stage | Eval mean reward |
|---|---|
| REINFORCE | 0.00 |
| A2C | 2.35 |
| A2C+GAE | 2.35 |
| PPO | 2.35 |
Usage
model = ActorCritic(n_actions)
model.load_state_dict(torch.load("ppo.pt"))
Limitations
Trained for a small number of environment steps relative to typical Atari PPO budgets (millions of frames); this demonstrates correct algorithm implementation and relative improvement across stages rather than a solved/high-scoring agent.