File size: 3,370 Bytes
829d2c6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
---
library_name: reinforce
tags:
- Pixelcopter-PLE-v0
- deep-reinforcement-learning
- reinforcement-learning
- policy-gradient
- reinforce
model-index:
- name: REINFORCE
  results:
  - task:
      type: reinforcement-learning
      name: reinforcement-learning
    dataset:
      name: Pixelcopter-PLE-v0
      type: Pixelcopter-PLE-v0
    metrics:
    - type: mean_reward
      value: 31.85 +/- 26.03
      name: mean_reward
      verified: false
---
# **REINFORCE** Agent playing **Pixelcopter-PLE-v0**

This is a trained model of a **REINFORCE** agent playing **Pixelcopter-PLE-v0**
using PyTorch and the [Deep Reinforcement Learning Course](https://huggingface.co/deep-rl-course/unit4).

## Algorithm
REINFORCE is a policy gradient method that:
- Directly optimizes the policy π(a|s)
- Uses Monte Carlo sampling to estimate returns
- Updates parameters in the direction of higher expected returns
- Belongs to the family of Policy Gradient methods


## Something To Say
- 😤Reach PLE (0.0.1) through trial and error with SSH Key on (https://github.com/ntasfi/PyGame-Learning-Environment)

- 😭Evaluate 100 turns to get a relatively low score

- 💡PixelCopter is wrapped with ```gymnasium.spaces``` in ```Unit 4_2.py``` 

- 🙂Continue training 20k steps with ```Unit 4_2_continue.py``` after 40k steps in ```Unit 4_2.py``` 

- Running time Reference: **3h15min** (40k steps)

- ☀️Wish you a good time~~~
## Evaluation Results
| Metric | Value |
|--------|-------|
| Mean Reward | 31.85 |
| Std Reward | 26.03 |
| Min Reward | 2.00 |
| Max Reward | 118.00 |
| Mean Episode Length | 220.25 |
| Score (mean - std) | 5.82 |
| Evaluation Episodes | 100 |
## Usage
```python
import torch
import torch.nn as nn
import torch.nn.functional as F
import gymnasium as gym
import numpy as np
class Policy(nn.Module):
    def __init__(self, s_size, a_size, h_size=128):
        super(Policy, self).__init__()
        self.fc1 = nn.Linear(s_size, h_size)
        self.fc2 = nn.Linear(h_size, h_size * 2)
        self.fc3 = nn.Linear(h_size * 2, a_size)
    def forward(self, x):
        x = F.relu(self.fc1(x))
        x = F.relu(self.fc2(x))
        return F.softmax(self.fc3(x), dim=1)
device = torch.device("cuda:0" if torch.cuda.is_available() else "cpu")
checkpoint = torch.load("reinforce_pixelcopter.pth", map_location=device)
policy = Policy(checkpoint['s_size'], checkpoint['a_size'], checkpoint['hidden_size'])
policy.load_state_dict(checkpoint['policy_state_dict'])
policy.eval()
env = gym.make("Pixelcopter-PLE-v0")
state, _ = env.reset()
for step in range(1000):
    state_tensor = torch.from_numpy(state).float().unsqueeze(0)
    with torch.no_grad():
        probs = policy(state_tensor)
        action = torch.argmax(probs, dim=1).item()
    
    state, reward, terminated, truncated, _ = env.step(action)
    
    if terminated or truncated:
        state, _ = env.reset()
## Training Configuration
- **Algorithm**: REINFORCE (Policy Gradient)
- **Policy Network**: 3-layer MLP (256-512 hidden units)
- **Optimizer**: Adam
- **Learning Rate**: 0.00003+0.00002
- **Discount Factor**: 0.99+0.995
- **Training Episodes**: 40000+20000
- **Device**: cuda:0
## Training Hyperparameters
- Episodes: 40000+20000
- Max steps per episode: 1000
- Learning rate: 0.00003+0.00002
- Gamma (discount factor): 0.99+0.995
- Hidden layer size: 256-512
- Optimizer: Adam