SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

arXiv PDF Code NeurIPS 2026 Project page SkillForge collection License: MIT


SkillForge pipeline: seed skills are pre-retired under the base model, the survivors seed retirement-aware cold-start, then skills and the policy co-evolve through trial, active, stable and retired states.

Introduction

This is the final SkillForge policy for Search-Augmented QA: a full fine-tune of Qwen2.5-7B-Instruct trained with GRPO while its skill library was simultaneously forged — retired, stabilized, demoted and mutated — under the fitness-driven lifecycle described in the paper.

Most memory-augmented agents keep the library append-only. A skill that was correct at step 20 encodes a procedure the policy has outgrown by step 120, and it is still being retrieved into the context. SkillForge instead scores each skill against the policy's own rollouts and moves it between four states — trial, active, stable, retired — so the library and the model co-evolve.

On Search-Augmented QA this reaches 48.7% overall accuracy across 51,713 test samples, the best of any method compared, against 46.8% for SkillRL and 45.2% for ZeroSearch.

What is in this repository. The policy weights and tokenizer only. The evolved skill library is not shipped here: the agent is the policy plus the retrieved skills injected into its context. Skill contents, fitness trajectories and retirement events are documented in SkillFurnace (Appendix C of the paper) and in the paper's case studies.

The same method, two more environments: ALFWorld and WebShop. All three sit in the SkillForge collection, alongside the paper.

Results

Search-Augmented QA, per dataset (%). * = in-domain, ** = out-of-domain. The Overall column is the sample-weighted micro-average over the same 51,713-sample union for every method.

Method NQ* TriviaQA** PopQA** HotpotQA* 2Wiki** MuSiQue** Bamboogle** Overall
RAG 27.4 58.2 17.8 25.8 23.2 9.4 16.8 29.4
Search-R1 39.3 61.0 39.7 37.0 40.1 14.6 36.8 42.9
ZeroSearch 43.6 61.8 51.5 34.6 35.2 18.4 27.8 45.2
EvolveR 43.5 63.4 44.6 38.2 42.0 15.6 54.4 45.8
SkillRL 45.9 63.3 45.9 43.2 40.3 20.2 73.8 46.8
SkillForge 48.2 65.0 50.1 43.9 40.4 20.3 77.2 48.7

The largest single gains are out-of-domain: +3.4 on Bamboogle and +5.9 over ZeroSearch on MuSiQue. ZeroSearch keeps PopQA, where it leads by 1.4.

The library this model ended up with

Seed library Final library
Total skills 41 85
General 10 —
Task-specific 20 —
Common-mistake 11 —

The run saturates its skill cap (S_max = 85). This environment retires the fewest skills of the three (59 events against ALFWorld's 138), and the two categories that stand out here — redundant with the policy and contradicting the environment — account for almost a quarter of them.

How it was trained

  1. Pre-retirement. The seed library, inherited from the SkillRL release, is scored under 400 rollout episodes of the base model with skills injected. Anything whose proto-fitness falls below delta_pre = 0.3 is retired before training starts.
  2. Retirement-aware cold start. The base model is fine-tuned with cross-entropy on the rollouts that survived, with trajectories that leaned on a since-retired skill filtered out.
  3. Skill-policy co-evolution. GRPO takes over and, every 10 steps, the forging cycle reads each skill's runtime fitness and promotes, demotes, retires or mutates it. Mutation is LLM-guided (Kimi-K2.5 as teacher); at most 5 mutations and 3 retirements per cycle; retrieval is top-6 by task-type match.
Hyperparameter Value
Base model Qwen2.5-7B-Instruct
Optimizer GRPO, via verl
Learning rate 1e-6
Batch size / group size 16 / 8
Clip epsilon / KL beta 0.2 / 0.001
Training steps 200
Sampling temperature (train and eval) 1.0
Hardware 64 x NVIDIA H200

Quick Start

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "YuyaoGe/SkillForge_Search_7B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")

# Skills are retrieved (top-6, by task-type match) and injected into the context.
# The exact prompt template and the action space are in the paper.
messages = [
    {"role": "system", "content": "<retrieved skills for this task type>"},
    {"role": "user", "content": "<observation>\n\n> "},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True,
                              return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][ids.shape[-1]:], skip_special_tokens=True))

The training loop, the skill lifecycle and the environment harnesses are in YuyaoGe/SkillForge — but that is the code, not the runtime library: this policy still needs its skills retrieved and injected at call time.

This is a research artifact: a 7B text policy for a text environment. It emits search queries and answers in the Search-R1 action format; the retrieval backend is not part of this repository.

Citation

@misc{ge2026skillforgecoevolvingskillsagents,
      title={SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles},
      author={Yuyao Ge and Yiwei Wang and Yuchen He and Baolong Bi and Lingrui Mei and Jiayu Yao and Lizhe Chen and Shenghua Liu},
      year={2026},
      eprint={2610.09832},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2610.09832},
}

Acknowledgments

Training runs on verl for the GRPO loop, with seed skill libraries inherited from the SkillRL release.

Downloads last month
85
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YuyaoGe/SkillForge_Search_7B

Base model

Qwen/Qwen2.5-7B
Finetuned
(3155)
this model

Collection including YuyaoGe/SkillForge_Search_7B

Paper for YuyaoGe/SkillForge_Search_7B