Instructions to use efficient-moe-agent-project/scienceworld with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use efficient-moe-agent-project/scienceworld with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
ScienceWorld — single-expert LoRA (GiGPO RL, gold-SFT warm start)
LoRA adapters for Qwen/Qwen2.5-1.5B-Instruct, trained as a single-expert agent on the ScienceWorld text environment (30 elementary-science tasks) with GiGPO reinforcement learning, warm-started from behaviour cloning on gold paths.
Run: scienceworld_single_warmstart_ms100_1k_seed0 (seed 0).
Checkpoints
| folder | step | val test_score | val success_rate | note |
|---|---|---|---|---|
single_expert_warmstart/final_step1000/ |
1000 | 5.11 | 0.271 | final policy (1000 RL steps) |
single_expert_warmstart/best_step140/ |
140 | 5.50 | 0.314 | best checkpoint by val/success_rate |
val/text/test_score is the mean ScienceWorld raw score / 10 over a fixed, stratified dev
validation set (140 episodes, all 30 task types, greedy decode). So test_score ≈ 5.1 means an
average raw score of ~51/100; success_rate is the fraction of episodes fully solved (raw
score = 100).
Note: intermediate checkpoints (e.g. the test_score peaks at steps 210 / 510, both ~5.6–5.7) were not retained; only the final and the best-by-success-rate checkpoints are available.
Training setup
- Base model: Qwen/Qwen2.5-1.5B-Instruct
- Adapter: LoRA, r = 64, α = 64, target = all-linear (q,k,v,o,gate,up,down)
- Warm start: gold-path behaviour-cloning SFT adapter (replay of
env.getGoldActionSequence(); ~6k (prompt,<think>…</think><action>…</action>) pairs; 3 epochs). Cold-start RL never solves a task (val success ≈ 0), so the warm start is required to give GiGPO a learning signal. - RL: GiGPO (
adv_estimator=gigpo, γ = 0.95), lr = 3e-6, 1000 steps.train_batch_size = 8, GiGPO grouprollout.n = 8,ppo_mini_batch_size = 64- invalid-action penalty coef = 0.1
- Env: ScienceWorld,
max_steps = 100,history_length = 2(last 2 obs+action pairs in the prompt). Reward = per-step ScienceWorld score delta / 10 (dense). - Validation: fixed stratified dev set (140 episodes, 30 tasks × ~5 variations), greedy (temperature 0), every 10 steps.
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = "Qwen/Qwen2.5-1.5B-Instruct"
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="bfloat16", device_map="auto")
tok = AutoTokenizer.from_pretrained(base)
model = PeftModel.from_pretrained(
model,
"efficient-moe-agent-project/scienceworld",
subfolder="single_expert_warmstart/best_step140", # or final_step1000
)
- Downloads last month
- -
Task type is invalid.