PPO on ALE/Breakout-v5 (from scratch, no SB3)

Custom PyTorch implementation of REINFORCE → A2C → A2C+GAE → PPO, built up incrementally for a coursework assignment. This checkpoint is the final PPO stage.

Environment

  • ALE/Breakout-v5, AtariPreprocessing (frameskip=4, grayscale, 84x84), FrameStackObservation(stack_size=4)

Architecture

Nature-DQN-style CNN encoder (3 conv layers → 512-d FC) shared between actor and critic heads.

Hyperparameters

  • gamma=0.99, lambda=0.95, clip_eps=0.2
  • lr=2.5e-4 (Adam), rollout_len=128, n_epochs=4, minibatch_size=32
  • value loss coef=0.5, entropy coef=0.01, grad clip=0.5

Evaluation

20-episode greedy evaluation: mean=2.35, std=0.48

Training progression

Stage Eval mean reward
REINFORCE 0.00
A2C 2.35
A2C+GAE 2.35
PPO 2.35

Usage

model = ActorCritic(n_actions)
model.load_state_dict(torch.load("ppo.pt"))

Limitations

Trained for a small number of environment steps relative to typical Atari PPO budgets (millions of frames); this demonstrates correct algorithm implementation and relative improvement across stages rather than a solved/high-scoring agent.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading