Instructions to use tsilva/Level1-4_stable-baselines3-ppo_be9abbc8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- stable-baselines3
How to use tsilva/Level1-4_stable-baselines3-ppo_be9abbc8 with stable-baselines3:
from huggingface_sb3 import load_from_hub checkpoint = load_from_hub( repo_id="tsilva/Level1-4_stable-baselines3-ppo_be9abbc8", filename="{MODEL FILENAME}.zip", ) - Notebooks
- Google Colab
- Kaggle
| library_name: stable-baselines3 | |
| pipeline_tag: reinforcement-learning | |
| license: mit | |
| tags: | |
| - reinforcement-learning | |
| - stable-baselines3 | |
| - ppo | |
| - supermariobrosnes-turbo | |
| - rlab | |
| - SuperMarioBros-Nes-v0 | |
| metrics: | |
| - success-rate | |
| # SuperMarioBros-Nes-v0 — Level1-4 — PPO | |
| Stable-Baselines3 PPO policy for `SuperMarioBros-Nes-v0` `Level1-4`, trained and evaluated with | |
| [`rlab`](https://github.com/tsilva/rlab). | |
| ## At a Glance | |
| | Item | Value | | |
| |---|---| | |
| | Task | Complete `SuperMarioBros-Nes-v0` `Level1-4` | | |
| | Provider | `supermariobrosnes-turbo` | | |
| | Algorithm | `ppo` | | |
| | Checkpoint | Step `4500000` | | |
| | Evaluation | `stochastic` full evaluation, `100` episodes | | |
| | Success | minimum `97.0%`, mean `97.0%` | | |
| | Mean return | `2294.958` | | |
| | Release | `v1` | | |
| | Preview | Root `replay.mp4` | | |
| | YouTube | [Watch on YouTube](https://www.youtube.com/watch?v=fvxFCTK9is8) | | |
| ## Quick Start | |
| ```bash | |
| git clone https://github.com/tsilva/rlab | |
| cd rlab | |
| git checkout 62e08d739b44f950cc7762b7319a75c2eba527ce | |
| uv sync --frozen | |
| ``` | |
| Import the ROM, then play or evaluate the immutable checkpoint: | |
| ```bash | |
| uv run rlab import-roms ~/roms --game SuperMarioBros-Nes-v0 | |
| uv run rlab play https://huggingface.co/tsilva/Level1-4_stable-baselines3-ppo_be9abbc8/resolve/v1/model.zip | |
| uv run rlab eval https://huggingface.co/tsilva/Level1-4_stable-baselines3-ppo_be9abbc8/resolve/v1/model.zip | |
| ``` | |
| ## Evaluation | |
| Action selection was `stochastic` under the published evaluation environment contract. | |
| | Start | Episodes | Successes | Success rate | Mean return | | |
| |---|---:|---:|---:|---:| | |
| | Level1-4 | 100 | 97 | 97.0% | 2294.958 | | |
| ## Environment and Policy Contract | |
| | Item | Value | | |
| |---|---| | |
| | Environment | `supermariobrosnes-turbo:SuperMarioBros-Nes-v0` | | |
| | Environment hash | `sha256:603e8f5a693bb5c7e56fbf316879e5858b1534e56cbbd11be79cfeec219b02eb` | | |
| | Preprocessing | `{"frame_skip":4,"frame_stack":4,"max_pool_frames":false,"obs_copy":"safe_view","obs_crop":[32,0,0,0],"obs_crop_fill":0,"obs_crop_mode":"remove","obs_grayscale":true,"obs_resize":[84,84],"obs_resize_algorithm":"area","pipeline":"supermariobrosnes_turbo_native_vec_env","policy_observation_layout":"channel_first","sticky_action_prob":0.0}` | | |
| | Action contract | `{"set":"simple"}` | | |
| ## Provenance | |
| | Item | Value | | |
| |---|---| | |
| | Source | [rlab](https://github.com/tsilva/rlab) | | |
| | Run | [Level1-4_base_s_20260704T081304Z](https://wandb.ai/tsilva/SuperMarioBros-Nes-v0/runs/7gjw67kl) | | |
| | Recipe | `base` | | |
| | Seed | `123` | | |
| | Source commit | `62e08d739b44f950cc7762b7319a75c2eba527ce` | | |
| | Evaluated artifact | `tsilva/SuperMarioBros-Nes-v0/Level1-4_base_s_20260704T081304Z-checkpoint:step-4500000` | | |
| ## Files | |
| | File | Purpose | | |
| |---|---| | |
| | `model.zip` | Stable-Baselines3 policy checkpoint | | |
| | `model.json` | Versioned checkpoint identity, policy type, provenance, and recipe binding | | |
| | `recipe.json` | Versioned execution and evaluation contract | | |
| | `release_manifest.json` | Release identity, evaluation evidence, and artifact hashes | | |
| | `replay.mp4` | Browser-safe representative episode | | |
| | `LICENSE` | License for rlab-authored policy weights and publication material | | |
| ## Limitations | |
| Evaluation establishes performance only for the published environment hash, start distribution, | |
| policy preprocessing, and action-selection protocol. It does not establish generalization to | |
| other levels, environments, ROM revisions, or contracts. | |
| ## Licensing | |
| The rlab-authored policy weights and publication material are licensed under the MIT License in | |
| `LICENSE`. Emulator/runtime software and game assets remain governed by their own licenses and | |
| terms. This repository does not redistribute a game ROM. | |
| ## Policy Lineage | |
| This is a legacy `rlab` policy trained by Stable-Baselines3 PPO. It is not a GradLab-trained checkpoint and no GradLab compatibility is claimed. | |
| - Trainer: `Stable-Baselines3` | |
| - Algorithm: `PPO` | |
| - Model class: `stable_baselines3.ppo.ppo.PPO` | |
| - Full lineage digest: `be9abbc80fcfab25e8d08cd1bdd432e7c61199f0aed99c54454ac7568b670336` | |
| - Immutable release: `hf://tsilva/Level1-4_stable-baselines3-ppo_be9abbc8@v1` | |
| - Exact checkpoint tag: `checkpoint-4500000` | |
| The mutable main branch contains this corrected card; the original v1 release files, commit, and tag remain unchanged. | |