tsilva's picture
Correct legacy trainer provenance and lineage links
c3b4efc verified
|
Raw
History Blame Contribute Delete
4.27 kB
---
library_name: stable-baselines3
pipeline_tag: reinforcement-learning
license: mit
tags:
- reinforcement-learning
- stable-baselines3
- ppo
- supermariobrosnes-turbo
- rlab
- SuperMarioBros-Nes-v0
metrics:
- success-rate
---
# SuperMarioBros-Nes-v0 — Level1-4 — PPO
Stable-Baselines3 PPO policy for `SuperMarioBros-Nes-v0` `Level1-4`, trained and evaluated with
[`rlab`](https://github.com/tsilva/rlab).
## At a Glance
| Item | Value |
|---|---|
| Task | Complete `SuperMarioBros-Nes-v0` `Level1-4` |
| Provider | `supermariobrosnes-turbo` |
| Algorithm | `ppo` |
| Checkpoint | Step `4500000` |
| Evaluation | `stochastic` full evaluation, `100` episodes |
| Success | minimum `97.0%`, mean `97.0%` |
| Mean return | `2294.958` |
| Release | `v1` |
| Preview | Root `replay.mp4` |
| YouTube | [Watch on YouTube](https://www.youtube.com/watch?v=fvxFCTK9is8) |
## Quick Start
```bash
git clone https://github.com/tsilva/rlab
cd rlab
git checkout 62e08d739b44f950cc7762b7319a75c2eba527ce
uv sync --frozen
```
Import the ROM, then play or evaluate the immutable checkpoint:
```bash
uv run rlab import-roms ~/roms --game SuperMarioBros-Nes-v0
uv run rlab play https://huggingface.co/tsilva/Level1-4_stable-baselines3-ppo_be9abbc8/resolve/v1/model.zip
uv run rlab eval https://huggingface.co/tsilva/Level1-4_stable-baselines3-ppo_be9abbc8/resolve/v1/model.zip
```
## Evaluation
Action selection was `stochastic` under the published evaluation environment contract.
| Start | Episodes | Successes | Success rate | Mean return |
|---|---:|---:|---:|---:|
| Level1-4 | 100 | 97 | 97.0% | 2294.958 |
## Environment and Policy Contract
| Item | Value |
|---|---|
| Environment | `supermariobrosnes-turbo:SuperMarioBros-Nes-v0` |
| Environment hash | `sha256:603e8f5a693bb5c7e56fbf316879e5858b1534e56cbbd11be79cfeec219b02eb` |
| Preprocessing | `{"frame_skip":4,"frame_stack":4,"max_pool_frames":false,"obs_copy":"safe_view","obs_crop":[32,0,0,0],"obs_crop_fill":0,"obs_crop_mode":"remove","obs_grayscale":true,"obs_resize":[84,84],"obs_resize_algorithm":"area","pipeline":"supermariobrosnes_turbo_native_vec_env","policy_observation_layout":"channel_first","sticky_action_prob":0.0}` |
| Action contract | `{"set":"simple"}` |
## Provenance
| Item | Value |
|---|---|
| Source | [rlab](https://github.com/tsilva/rlab) |
| Run | [Level1-4_base_s_20260704T081304Z](https://wandb.ai/tsilva/SuperMarioBros-Nes-v0/runs/7gjw67kl) |
| Recipe | `base` |
| Seed | `123` |
| Source commit | `62e08d739b44f950cc7762b7319a75c2eba527ce` |
| Evaluated artifact | `tsilva/SuperMarioBros-Nes-v0/Level1-4_base_s_20260704T081304Z-checkpoint:step-4500000` |
## Files
| File | Purpose |
|---|---|
| `model.zip` | Stable-Baselines3 policy checkpoint |
| `model.json` | Versioned checkpoint identity, policy type, provenance, and recipe binding |
| `recipe.json` | Versioned execution and evaluation contract |
| `release_manifest.json` | Release identity, evaluation evidence, and artifact hashes |
| `replay.mp4` | Browser-safe representative episode |
| `LICENSE` | License for rlab-authored policy weights and publication material |
## Limitations
Evaluation establishes performance only for the published environment hash, start distribution,
policy preprocessing, and action-selection protocol. It does not establish generalization to
other levels, environments, ROM revisions, or contracts.
## Licensing
The rlab-authored policy weights and publication material are licensed under the MIT License in
`LICENSE`. Emulator/runtime software and game assets remain governed by their own licenses and
terms. This repository does not redistribute a game ROM.
## Policy Lineage
This is a legacy `rlab` policy trained by Stable-Baselines3 PPO. It is not a GradLab-trained checkpoint and no GradLab compatibility is claimed.
- Trainer: `Stable-Baselines3`
- Algorithm: `PPO`
- Model class: `stable_baselines3.ppo.ppo.PPO`
- Full lineage digest: `be9abbc80fcfab25e8d08cd1bdd432e7c61199f0aed99c54454ac7568b670336`
- Immutable release: `hf://tsilva/Level1-4_stable-baselines3-ppo_be9abbc8@v1`
- Exact checkpoint tag: `checkpoint-4500000`
The mutable main branch contains this corrected card; the original v1 release files, commit, and tag remain unchanged.