File size: 4,265 Bytes
629c7a2
 
 
10177c7
629c7a2
 
 
 
10177c7
629c7a2
 
 
10177c7
629c7a2
 
10177c7
629c7a2
10177c7
 
 
 
 
 
 
 
 
 
 
d91d6e0
 
 
 
10177c7
 
629c7a2
 
 
d91d6e0
 
 
 
 
 
10177c7
 
629c7a2
 
10177c7
0b9ea84
 
629c7a2
 
10177c7
629c7a2
d91d6e0
629c7a2
10177c7
 
d91d6e0
629c7a2
10177c7
629c7a2
10177c7
629c7a2
10177c7
d91d6e0
 
10177c7
629c7a2
 
 
 
 
10177c7
 
 
 
d91d6e0
10177c7
 
 
 
 
 
 
d91d6e0
 
 
10177c7
 
 
 
 
 
 
 
86521e2
 
 
10177c7
 
 
0b9ea84
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
---
library_name: stable-baselines3
pipeline_tag: reinforcement-learning
license: mit
tags:
  - reinforcement-learning
  - stable-baselines3
  - ppo
  - supermariobrosnes-turbo
  - rlab
  - SuperMarioBros-Nes-v0
metrics:
  - success-rate
---

# SuperMarioBros-Nes-v0 — Level4-1 — PPO

Stable-Baselines3 PPO policy for `SuperMarioBros-Nes-v0` `Level4-1`, trained and evaluated with
[`rlab`](https://github.com/tsilva/rlab).

## At a Glance

| Item | Value |
|---|---|
| Task | Complete `SuperMarioBros-Nes-v0` `Level4-1` |
| Provider | `supermariobrosnes-turbo` |
| Algorithm | `ppo` |
| Checkpoint | Step `7500000` |
| Evaluation | `stochastic` full evaluation, `100` episodes |
| Success | minimum `97.0%`, mean `97.0%` |
| Mean return | `3529.275` |
| Release | `v1` |
| Preview | Root `replay.mp4` |
| YouTube | [Watch on YouTube](https://www.youtube.com/watch?v=HqTb6AT5M_8) |

## Quick Start

```bash
git clone https://github.com/tsilva/rlab
cd rlab
git checkout 044dcc3d2a978fb1a77524822238c4c7814c002d
uv sync --frozen
```

Import the ROM, then play or evaluate the immutable checkpoint:

```bash
uv run rlab import-roms ~/roms --game SuperMarioBros-Nes-v0
uv run rlab play https://huggingface.co/tsilva/Level4-1_stable-baselines3-ppo_fe8a02dc/resolve/v1/model.zip
uv run rlab eval https://huggingface.co/tsilva/Level4-1_stable-baselines3-ppo_fe8a02dc/resolve/v1/model.zip
```

## Evaluation

Action selection was `stochastic` under the published evaluation environment contract.

| Start | Episodes | Successes | Success rate | Mean return |
|---|---:|---:|---:|---:|
| Level4-1 | 100 | 97 | 97.0% | 3529.275 |

## Environment and Policy Contract

| Item | Value |
|---|---|
| Environment | `supermariobrosnes-turbo:SuperMarioBros-Nes-v0` |
| Environment hash | `sha256:fa0f3dd473fac528938a2cc2130dfe2d516ef164946b6545f2262ff6777c4bef` |
| Preprocessing | `{"frame_skip":4,"frame_stack":4,"max_pool_frames":false,"obs_copy":"safe_view","obs_crop":[32,0,0,0],"obs_crop_fill":0,"obs_crop_mode":"remove","obs_grayscale":true,"obs_resize":[84,84],"obs_resize_algorithm":"area","pipeline":"supermariobrosnes_turbo_native_vec_env","policy_observation_layout":"channel_first","sticky_action_prob":0.0}` |
| Action contract | `{"set":"simple"}` |

## Provenance

| Item | Value |
|---|---|
| Source | [rlab](https://github.com/tsilva/rlab) |
| Run | [Level4-1_base_s1_20260704T182045Z](https://wandb.ai/tsilva/SuperMarioBros-Nes-v0/runs/4xrsz2v1) |
| Recipe | `base` |
| Seed | `1` |
| Source commit | `044dcc3d2a978fb1a77524822238c4c7814c002d` |
| Evaluated artifact | `tsilva/SuperMarioBros-Nes-v0/Level4-1_base_s1_20260704T182045Z-checkpoint:step-7500000` |

## Files

| File | Purpose |
|---|---|
| `model.zip` | Stable-Baselines3 policy checkpoint |
| `model.json` | Versioned checkpoint identity, policy type, provenance, and recipe binding |
| `recipe.json` | Versioned execution and evaluation contract |
| `release_manifest.json` | Release identity, evaluation evidence, and artifact hashes |
| `replay.mp4` | Browser-safe representative episode |
| `LICENSE` | License for rlab-authored policy weights and publication material |

## Limitations

Evaluation establishes performance only for the published environment hash, start distribution,
policy preprocessing, and action-selection protocol. It does not establish generalization to
other levels, environments, ROM revisions, or contracts.

## Licensing

The rlab-authored policy weights and publication material are licensed under the MIT License in
`LICENSE`. Emulator/runtime software and game assets remain governed by their own licenses and
terms. This repository does not redistribute a game ROM.

## Policy Lineage

This is a legacy `rlab` policy trained by Stable-Baselines3 PPO. It is not a GradLab-trained checkpoint and no GradLab compatibility is claimed.

- Trainer: `Stable-Baselines3`
- Algorithm: `PPO`
- Model class: `stable_baselines3.ppo.ppo.PPO`
- Full lineage digest: `fe8a02dca01803a336dfecd8803be53541b192c0b0d05a525b4a6afcbae4c58a`
- Immutable release: `hf://tsilva/Level4-1_stable-baselines3-ppo_fe8a02dc@v1`
- Exact checkpoint tag: `checkpoint-7500000`

The mutable main branch contains this corrected card; the original v1 release files, commit, and tag remain unchanged.