--- license: mit library_name: transformers pipeline_tag: text-generation base_model: Qwen/Qwen2.5-7B-Instruct tags: - agentic-rl - llm-agent - skill-library - grpo - search-qa --- # SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles [![arXiv](https://img.shields.io/badge/arXiv-2610.09832-b31b1b)](https://arxiv.org/abs/2610.09832) [![PDF](https://img.shields.io/badge/PDF-full%20paper-white)](https://arxiv.org/pdf/2610.09832) [![Code](https://img.shields.io/badge/Code-SkillForge-181717?logo=github)](https://github.com/YuyaoGe/SkillForge) [![NeurIPS 2026](https://img.shields.io/badge/NeurIPS-2026-1783ff)](https://geyuyao.com/skillforge/) [![Project page](https://img.shields.io/badge/project-SkillForge-b5613c)](https://geyuyao.com/skillforge/) [![SkillForge collection](https://img.shields.io/badge/%F0%9F%A4%97-collection-ffc107)](https://huggingface.co/collections/YuyaoGe/skillforge) [![License: MIT](https://img.shields.io/badge/license-MIT-f5de53)](https://huggingface.co/YuyaoGe/SkillForge_Search_7B/blob/main/README.md) --- ![SkillForge pipeline: seed skills are pre-retired under the base model, the survivors seed retirement-aware cold-start, then skills and the policy co-evolve through trial, active, stable and retired states.](assets/overview.png) ## Introduction This is the **final SkillForge policy for Search-Augmented QA**: a full fine-tune of `Qwen2.5-7B-Instruct` trained with GRPO while its skill library was simultaneously forged — retired, stabilized, demoted and mutated — under the fitness-driven lifecycle described in the paper. Most memory-augmented agents keep the library **append-only**. A skill that was correct at step 20 encodes a procedure the policy has outgrown by step 120, and it is still being retrieved into the context. SkillForge instead scores each skill against the policy's own rollouts and moves it between four states — `trial`, `active`, `stable`, `retired` — so the library and the model co-evolve. On Search-Augmented QA this reaches **48.7%** overall accuracy across 51,713 test samples, the best of any method compared, against **46.8%** for SkillRL and **45.2%** for ZeroSearch. > **What is in this repository.** The policy weights and tokenizer only. The evolved > skill library is *not* shipped here: the agent is the policy **plus** the retrieved > skills injected into its context. Skill contents, fitness trajectories and > retirement events are documented in SkillFurnace (Appendix C of the paper) and in > the paper's case studies. The same method, two more environments: [ALFWorld](https://huggingface.co/YuyaoGe/SkillForge_Alfworld_7B) and [WebShop](https://huggingface.co/YuyaoGe/SkillForge_Webshop_7B). All three sit in the [SkillForge collection](https://huggingface.co/collections/YuyaoGe/skillforge), alongside the paper. ## Results Search-Augmented QA, per dataset (%). \* = in-domain, \*\* = out-of-domain. The `Overall` column is the sample-weighted micro-average over the same 51,713-sample union for every method. | Method | NQ\* | TriviaQA\*\* | PopQA\*\* | HotpotQA\* | 2Wiki\*\* | MuSiQue\*\* | Bamboogle\*\* | **Overall** | |---|---|---|---|---|---|---|---|---| | RAG | 27.4 | 58.2 | 17.8 | 25.8 | 23.2 | 9.4 | 16.8 | 29.4 | | Search-R1 | 39.3 | 61.0 | 39.7 | 37.0 | 40.1 | 14.6 | 36.8 | 42.9 | | ZeroSearch | 43.6 | 61.8 | **51.5** | 34.6 | 35.2 | 18.4 | 27.8 | 45.2 | | EvolveR | 43.5 | 63.4 | 44.6 | 38.2 | 42.0 | 15.6 | 54.4 | 45.8 | | SkillRL | 45.9 | 63.3 | 45.9 | 43.2 | 40.3 | 20.2 | 73.8 | 46.8 | | **SkillForge** | **48.2** | **65.0** | 50.1 | **43.9** | 40.4 | **20.3** | **77.2** | **48.7** | The largest single gains are out-of-domain: **+3.4** on Bamboogle and **+5.9** over ZeroSearch on MuSiQue. ZeroSearch keeps PopQA, where it leads by 1.4. ### The library this model ended up with | | Seed library | Final library | |---|---|---| | Total skills | 41 | **85** | | General | 10 | — | | Task-specific | 20 | — | | Common-mistake | 11 | — | The run saturates its skill cap (`S_max = 85`). This environment retires the fewest skills of the three (59 events against ALFWorld's 138), and the two categories that stand out here — redundant with the policy and contradicting the environment — account for almost a quarter of them. ## How it was trained 1. **Pre-retirement.** The seed library, inherited from the SkillRL release, is scored under 400 rollout episodes of the *base* model with skills injected. Anything whose proto-fitness falls below `delta_pre = 0.3` is retired before training starts. 2. **Retirement-aware cold start.** The base model is fine-tuned with cross-entropy on the rollouts that survived, with trajectories that leaned on a since-retired skill filtered out. 3. **Skill-policy co-evolution.** GRPO takes over and, every 10 steps, the forging cycle reads each skill's runtime fitness and promotes, demotes, retires or mutates it. Mutation is LLM-guided (Kimi-K2.5 as teacher); at most 5 mutations and 3 retirements per cycle; retrieval is top-6 by task-type match. | Hyperparameter | Value | |---|---| | Base model | `Qwen2.5-7B-Instruct` | | Optimizer | GRPO, via `verl` | | Learning rate | 1e-6 | | Batch size / group size | 16 / 8 | | Clip epsilon / KL beta | 0.2 / 0.001 | | Training steps | 200 | | Sampling temperature (train and eval) | 1.0 | | Hardware | 64 x NVIDIA H200 | ## Quick Start ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "YuyaoGe/SkillForge_Search_7B" tok = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto") # Skills are retrieved (top-6, by task-type match) and injected into the context. # The exact prompt template and the action space are in the paper. messages = [ {"role": "system", "content": ""}, {"role": "user", "content": "\n\n> "}, ] ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device) out = model.generate(ids, max_new_tokens=256, do_sample=False) print(tok.decode(out[0][ids.shape[-1]:], skip_special_tokens=True)) ``` The training loop, the skill lifecycle and the environment harnesses are in [YuyaoGe/SkillForge](https://github.com/YuyaoGe/SkillForge) — but that is the code, not the runtime library: this policy still needs its skills retrieved and injected at call time. This is a research artifact: a 7B text policy for a text environment. It emits search queries and answers in the Search-R1 action format; the retrieval backend is not part of this repository. ## Citation ```bibtex @misc{ge2026skillforgecoevolvingskillsagents, title={SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles}, author={Yuyao Ge and Yiwei Wang and Yuchen He and Baolong Bi and Lingrui Mei and Jiayu Yao and Lizhe Chen and Shenghua Liu}, year={2026}, eprint={2610.09832}, archivePrefix={arXiv}, primaryClass={cs.AI}, url={https://arxiv.org/abs/2610.09832}, } ``` ## Acknowledgments Training runs on [`verl`](https://github.com/volcengine/verl) for the GRPO loop, with seed skill libraries inherited from the SkillRL release.