File size: 4,265 Bytes
a0606eb
 
 
b2098e0
a0606eb
 
 
 
b2098e0
a0606eb
 
 
b2098e0
a0606eb
 
b2098e0
a0606eb
b2098e0
 
 
 
 
 
 
 
 
 
 
76560ae
 
 
 
b2098e0
 
a0606eb
 
 
76560ae
 
 
 
 
 
b2098e0
 
a0606eb
 
b2098e0
97fc4d2
 
a0606eb
 
b2098e0
a0606eb
76560ae
a0606eb
b2098e0
 
76560ae
a0606eb
b2098e0
a0606eb
b2098e0
a0606eb
b2098e0
76560ae
 
b2098e0
a0606eb
 
 
 
 
b2098e0
 
 
 
76560ae
b2098e0
 
 
 
 
 
 
76560ae
 
 
b2098e0
 
 
 
 
 
 
 
a1eb9fb
 
 
b2098e0
 
 
97fc4d2
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
---
library_name: stable-baselines3
pipeline_tag: reinforcement-learning
license: mit
tags:
  - reinforcement-learning
  - stable-baselines3
  - ppo
  - supermariobrosnes-turbo
  - rlab
  - SuperMarioBros-Nes-v0
metrics:
  - success-rate
---

# SuperMarioBros-Nes-v0 — Level1-2 — PPO

Stable-Baselines3 PPO policy for `SuperMarioBros-Nes-v0` `Level1-2`, trained and evaluated with
[`rlab`](https://github.com/tsilva/rlab).

## At a Glance

| Item | Value |
|---|---|
| Task | Complete `SuperMarioBros-Nes-v0` `Level1-2` |
| Provider | `supermariobrosnes-turbo` |
| Algorithm | `ppo` |
| Checkpoint | Step `7500000` |
| Evaluation | `stochastic` full evaluation, `100` episodes |
| Success | minimum `97.0%`, mean `97.0%` |
| Mean return | `3077.086` |
| Release | `v1` |
| Preview | Root `replay.mp4` |
| YouTube | [Watch on YouTube](https://www.youtube.com/watch?v=JtgtSLy4hnI) |

## Quick Start

```bash
git clone https://github.com/tsilva/rlab
cd rlab
git checkout 40b8613743e41d36489701888cafcee24306604e
uv sync --frozen
```

Import the ROM, then play or evaluate the immutable checkpoint:

```bash
uv run rlab import-roms ~/roms --game SuperMarioBros-Nes-v0
uv run rlab play https://huggingface.co/tsilva/Level1-2_stable-baselines3-ppo_91df110f/resolve/v1/model.zip
uv run rlab eval https://huggingface.co/tsilva/Level1-2_stable-baselines3-ppo_91df110f/resolve/v1/model.zip
```

## Evaluation

Action selection was `stochastic` under the published evaluation environment contract.

| Start | Episodes | Successes | Success rate | Mean return |
|---|---:|---:|---:|---:|
| Level1-2 | 100 | 97 | 97.0% | 3077.086 |

## Environment and Policy Contract

| Item | Value |
|---|---|
| Environment | `supermariobrosnes-turbo:SuperMarioBros-Nes-v0` |
| Environment hash | `sha256:3ed482a01669b9e31884d67157ab5724d1bd45f97a1db19ae2515331e4387f73` |
| Preprocessing | `{"frame_skip":4,"frame_stack":4,"max_pool_frames":false,"obs_copy":"safe_view","obs_crop":[32,0,0,0],"obs_crop_fill":0,"obs_crop_mode":"remove","obs_grayscale":true,"obs_resize":[84,84],"obs_resize_algorithm":"area","pipeline":"supermariobrosnes_turbo_native_vec_env","policy_observation_layout":"channel_first","sticky_action_prob":0.0}` |
| Action contract | `{"set":"simple"}` |

## Provenance

| Item | Value |
|---|---|
| Source | [rlab](https://github.com/tsilva/rlab) |
| Run | [Level1-2_base_s1_20260704T072752Z](https://wandb.ai/tsilva/SuperMarioBros-Nes-v0/runs/dtpvaqq2) |
| Recipe | `base` |
| Seed | `1` |
| Source commit | `40b8613743e41d36489701888cafcee24306604e` |
| Evaluated artifact | `tsilva/SuperMarioBros-Nes-v0/Level1-2_base_s1_20260704T072752Z-checkpoint:step-7500000` |

## Files

| File | Purpose |
|---|---|
| `model.zip` | Stable-Baselines3 policy checkpoint |
| `model.json` | Versioned checkpoint identity, policy type, provenance, and recipe binding |
| `recipe.json` | Versioned execution and evaluation contract |
| `release_manifest.json` | Release identity, evaluation evidence, and artifact hashes |
| `replay.mp4` | Browser-safe representative episode |
| `LICENSE` | License for rlab-authored policy weights and publication material |

## Limitations

Evaluation establishes performance only for the published environment hash, start distribution,
policy preprocessing, and action-selection protocol. It does not establish generalization to
other levels, environments, ROM revisions, or contracts.

## Licensing

The rlab-authored policy weights and publication material are licensed under the MIT License in
`LICENSE`. Emulator/runtime software and game assets remain governed by their own licenses and
terms. This repository does not redistribute a game ROM.

## Policy Lineage

This is a legacy `rlab` policy trained by Stable-Baselines3 PPO. It is not a GradLab-trained checkpoint and no GradLab compatibility is claimed.

- Trainer: `Stable-Baselines3`
- Algorithm: `PPO`
- Model class: `stable_baselines3.ppo.ppo.PPO`
- Full lineage digest: `91df110f1c98e4200d38c20a204f67a8fbaf4e3b6ec1a0fc2fc382a65c7814ea`
- Immutable release: `hf://tsilva/Level1-2_stable-baselines3-ppo_91df110f@v1`
- Exact checkpoint tag: `checkpoint-7500000`

The mutable main branch contains this corrected card; the original v1 release files, commit, and tag remain unchanged.