Instructions to use limmin33/opd-rl with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use limmin33/opd-rl with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("limmin33/opd-rl", device_map="auto") - Notebooks
- Google Colab
- Kaggle
opd-rl checkpoints
Checkpoints are stored four levels deep:
<environment>/<model>/<training-method>/<checkpoint>/
Each checkpoint directory holds the verl FSDP actor state:
<checkpoint>/actor/model_world_size_<N>_rank_<r>.pt sharded policy weights
<checkpoint>/actor/optim_world_size_<N>_rank_<r>.pt optimizer state (full uploads only)
<checkpoint>/actor/extra_state_world_size_<N>_rank_<r>.pt
<checkpoint>/actor/config.json, tokenizer files
<checkpoint>/data.pt dataloader state
<checkpoint>/hf/ merged HuggingFace weights (when present)
Merge the shards back into a loadable HuggingFace model with verl:
python scripts/model_merger.py merge --backend fsdp \
--local_dir <checkpoint>/actor --target_dir ./merged
Contents
| environment | model | training method | checkpoint |
|---|---|---|---|
alfworld |
qwen2.5-3b |
grpo |
global_step_50 |
alfworld |
qwen2.5-3b |
rlsd_warmdown_100 |
global_step_150 |
alfworld |
qwen2.5-3b |
sdar_coef0.01_beta5.0_skillallfalse |
global_step_150 |
alfworld |
qwen3-1.7b |
grpo_2gpu |
global_step_100 |
alfworld |
qwen3-1.7b |
grpo_2gpu |
global_step_150 |
alfworld |
qwen3-1.7b |
grpo_2gpu |
global_step_50 |
alfworld |
qwen3-1.7b |
rlsd_warmdown_100 |
global_step_100 |
alfworld |
qwen3-1.7b |
rlsd_warmdown_100 |
global_step_150 |
alfworld |
qwen3-1.7b |
rlsd_warmdown_100 |
global_step_50 |
alfworld |
qwen3-1.7b |
sdar_coef0.01_beta5.0_skillallfalse |
global_step_100 |
alfworld |
qwen3-1.7b |
sdar_coef0.01_beta5.0_skillallfalse |
global_step_150 |
alfworld |
qwen3-1.7b |
sdar_coef0.01_beta5.0_skillallfalse |
global_step_50 |
Uploaded with scripts/upload_checkpoints_to_hf.py (--content full).