SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

arXiv PDF Code NeurIPS 2026 Project page SkillForge collection License: MIT


SkillForge pipeline: seed skills are pre-retired under the base model, the survivors seed retirement-aware cold-start, then skills and the policy co-evolve through trial, active, stable and retired states.

Introduction

This is the final SkillForge policy for WebShop: a full fine-tune of Qwen2.5-7B-Instruct trained with GRPO while its skill library was simultaneously forged — retired, stabilized, demoted and mutated — under the fitness-driven lifecycle described in the paper.

Most memory-augmented agents keep the library append-only. A skill that was correct at step 20 encodes a procedure the policy has outgrown by step 120, and it is still being retrieved into the context. SkillForge instead scores each skill against the policy's own rollouts and moves it between four states — trial, active, stable, retired — so the library and the model co-evolve.

On WebShop this reaches 78.4% task success and 87.8 partial-credit score, against 72.7% / 85.2 for SkillRL, and 7.8% / 26.4 for the untuned base model.

What is in this repository. The policy weights and tokenizer only. The evolved skill library is not shipped here: the agent is the policy plus the retrieved skills injected into its context. Skill contents, fitness trajectories and retirement events are documented in SkillFurnace (Appendix C of the paper) and in the paper's case studies.

The same method, two more environments: ALFWorld and Search-Augmented QA. All three sit in the SkillForge collection, alongside the paper.

Results

WebShop. Score is the partial-credit reward; Succ. is binary task success (%).

Method Score Succ.
Vanilla (base model) 26.4 7.8
ReAct 46.2 19.5
SimpleMem+GRPO 67.8 46.9
RLOO 80.3 65.7
GRPO (no library) 79.3 66.1
SkillRL 85.2 72.7
SkillForge 87.8 78.4

Two failure modes the lifecycle targets show up mostly here: hallucinated actions (13 of 121 retirement events, all in WebShop) and overly rigid ordering, where a skill prescribes a checkout sequence the site no longer accepts.

The library this model ended up with

Seed library Final library
Total skills 66 95
General 15 —
Task-specific 39 —
Common-mistake 12 —

The run saturates its skill cap (S_max = 95). The w/o retirement ablation grows to 121 skills and drops to 75.9% / 2.5 below SkillForge.

How it was trained

  1. Pre-retirement. The seed library, inherited from the SkillRL release, is scored under 450 rollout episodes of the base model with skills injected. Anything whose proto-fitness falls below delta_pre = 0.3 is retired before training starts.
  2. Retirement-aware cold start. The base model is fine-tuned with cross-entropy on the rollouts that survived, with trajectories that leaned on a since-retired skill filtered out.
  3. Skill-policy co-evolution. GRPO takes over and, every 10 steps, the forging cycle reads each skill's runtime fitness and promotes, demotes, retires or mutates it. Mutation is LLM-guided (Kimi-K2.5 as teacher); at most 5 mutations and 3 retirements per cycle; retrieval is top-6 by task-type match.
Hyperparameter Value
Base model Qwen2.5-7B-Instruct
Optimizer GRPO, via verl
Learning rate 1e-6
Batch size / group size 16 / 8
Clip epsilon / KL beta 0.2 / 0.001
Training steps 200
Sampling temperature (train and eval) 1.0
Hardware 64 x NVIDIA H200

Quick Start

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "YuyaoGe/SkillForge_Webshop_7B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")

# Skills are retrieved (top-6, by task-type match) and injected into the context.
# The exact prompt template and the action space are in the paper.
messages = [
    {"role": "system", "content": "<retrieved skills for this task type>"},
    {"role": "user", "content": "<observation>\n\n> "},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True,
                              return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][ids.shape[-1]:], skip_special_tokens=True))

The training loop, the skill lifecycle and the environment harnesses are in YuyaoGe/SkillForge — but that is the code, not the runtime library: this policy still needs its skills retrieved and injected at call time.

This is a research artifact: a 7B text policy for a text environment. It exposes no tool interface of its own and will emit WebShop-style search and click actions for anything that resembles a WebShop prompt.

Citation

@misc{ge2026skillforgecoevolvingskillsagents,
      title={SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles},
      author={Yuyao Ge and Yiwei Wang and Yuchen He and Baolong Bi and Lingrui Mei and Jiayu Yao and Lizhe Chen and Shenghua Liu},
      year={2026},
      eprint={2610.09832},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2610.09832},
}

Acknowledgments

Training runs on verl for the GRPO loop, with seed skill libraries inherited from the SkillRL release.

Downloads last month
84
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YuyaoGe/SkillForge_Webshop_7B

Base model

Qwen/Qwen2.5-7B
Finetuned
(3155)
this model

Collection including YuyaoGe/SkillForge_Webshop_7B

Paper for YuyaoGe/SkillForge_Webshop_7B