Instructions to use tsilva/Level4-1_stable-baselines3-ppo_fe8a02dc with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- stable-baselines3
How to use tsilva/Level4-1_stable-baselines3-ppo_fe8a02dc with stable-baselines3:
from huggingface_sb3 import load_from_hub checkpoint = load_from_hub( repo_id="tsilva/Level4-1_stable-baselines3-ppo_fe8a02dc", filename="{MODEL FILENAME}.zip", ) - Notebooks
- Google Colab
- Kaggle
File size: 4,265 Bytes
629c7a2 10177c7 629c7a2 10177c7 629c7a2 10177c7 629c7a2 10177c7 629c7a2 10177c7 d91d6e0 10177c7 629c7a2 d91d6e0 10177c7 629c7a2 10177c7 0b9ea84 629c7a2 10177c7 629c7a2 d91d6e0 629c7a2 10177c7 d91d6e0 629c7a2 10177c7 629c7a2 10177c7 629c7a2 10177c7 d91d6e0 10177c7 629c7a2 10177c7 d91d6e0 10177c7 d91d6e0 10177c7 86521e2 10177c7 0b9ea84 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 | ---
library_name: stable-baselines3
pipeline_tag: reinforcement-learning
license: mit
tags:
- reinforcement-learning
- stable-baselines3
- ppo
- supermariobrosnes-turbo
- rlab
- SuperMarioBros-Nes-v0
metrics:
- success-rate
---
# SuperMarioBros-Nes-v0 — Level4-1 — PPO
Stable-Baselines3 PPO policy for `SuperMarioBros-Nes-v0` `Level4-1`, trained and evaluated with
[`rlab`](https://github.com/tsilva/rlab).
## At a Glance
| Item | Value |
|---|---|
| Task | Complete `SuperMarioBros-Nes-v0` `Level4-1` |
| Provider | `supermariobrosnes-turbo` |
| Algorithm | `ppo` |
| Checkpoint | Step `7500000` |
| Evaluation | `stochastic` full evaluation, `100` episodes |
| Success | minimum `97.0%`, mean `97.0%` |
| Mean return | `3529.275` |
| Release | `v1` |
| Preview | Root `replay.mp4` |
| YouTube | [Watch on YouTube](https://www.youtube.com/watch?v=HqTb6AT5M_8) |
## Quick Start
```bash
git clone https://github.com/tsilva/rlab
cd rlab
git checkout 044dcc3d2a978fb1a77524822238c4c7814c002d
uv sync --frozen
```
Import the ROM, then play or evaluate the immutable checkpoint:
```bash
uv run rlab import-roms ~/roms --game SuperMarioBros-Nes-v0
uv run rlab play https://huggingface.co/tsilva/Level4-1_stable-baselines3-ppo_fe8a02dc/resolve/v1/model.zip
uv run rlab eval https://huggingface.co/tsilva/Level4-1_stable-baselines3-ppo_fe8a02dc/resolve/v1/model.zip
```
## Evaluation
Action selection was `stochastic` under the published evaluation environment contract.
| Start | Episodes | Successes | Success rate | Mean return |
|---|---:|---:|---:|---:|
| Level4-1 | 100 | 97 | 97.0% | 3529.275 |
## Environment and Policy Contract
| Item | Value |
|---|---|
| Environment | `supermariobrosnes-turbo:SuperMarioBros-Nes-v0` |
| Environment hash | `sha256:fa0f3dd473fac528938a2cc2130dfe2d516ef164946b6545f2262ff6777c4bef` |
| Preprocessing | `{"frame_skip":4,"frame_stack":4,"max_pool_frames":false,"obs_copy":"safe_view","obs_crop":[32,0,0,0],"obs_crop_fill":0,"obs_crop_mode":"remove","obs_grayscale":true,"obs_resize":[84,84],"obs_resize_algorithm":"area","pipeline":"supermariobrosnes_turbo_native_vec_env","policy_observation_layout":"channel_first","sticky_action_prob":0.0}` |
| Action contract | `{"set":"simple"}` |
## Provenance
| Item | Value |
|---|---|
| Source | [rlab](https://github.com/tsilva/rlab) |
| Run | [Level4-1_base_s1_20260704T182045Z](https://wandb.ai/tsilva/SuperMarioBros-Nes-v0/runs/4xrsz2v1) |
| Recipe | `base` |
| Seed | `1` |
| Source commit | `044dcc3d2a978fb1a77524822238c4c7814c002d` |
| Evaluated artifact | `tsilva/SuperMarioBros-Nes-v0/Level4-1_base_s1_20260704T182045Z-checkpoint:step-7500000` |
## Files
| File | Purpose |
|---|---|
| `model.zip` | Stable-Baselines3 policy checkpoint |
| `model.json` | Versioned checkpoint identity, policy type, provenance, and recipe binding |
| `recipe.json` | Versioned execution and evaluation contract |
| `release_manifest.json` | Release identity, evaluation evidence, and artifact hashes |
| `replay.mp4` | Browser-safe representative episode |
| `LICENSE` | License for rlab-authored policy weights and publication material |
## Limitations
Evaluation establishes performance only for the published environment hash, start distribution,
policy preprocessing, and action-selection protocol. It does not establish generalization to
other levels, environments, ROM revisions, or contracts.
## Licensing
The rlab-authored policy weights and publication material are licensed under the MIT License in
`LICENSE`. Emulator/runtime software and game assets remain governed by their own licenses and
terms. This repository does not redistribute a game ROM.
## Policy Lineage
This is a legacy `rlab` policy trained by Stable-Baselines3 PPO. It is not a GradLab-trained checkpoint and no GradLab compatibility is claimed.
- Trainer: `Stable-Baselines3`
- Algorithm: `PPO`
- Model class: `stable_baselines3.ppo.ppo.PPO`
- Full lineage digest: `fe8a02dca01803a336dfecd8803be53541b192c0b0d05a525b4a6afcbae4c58a`
- Immutable release: `hf://tsilva/Level4-1_stable-baselines3-ppo_fe8a02dc@v1`
- Exact checkpoint tag: `checkpoint-7500000`
The mutable main branch contains this corrected card; the original v1 release files, commit, and tag remain unchanged.
|