ToolWeave_stage2 / README.md
muradil211's picture
docs: simplify model card links
0047085 verified
|
Raw
History Blame Contribute Delete
2.82 kB
---
library_name: transformers
pipeline_tag: text-generation
tags:
- qwen3
- tool-calling
- bfcl
- agentic-rl
- progress-reward
---
<div align="center">
<img src="assets/toolweave-mark.svg" alt="ToolWeave mark" width="128">
<h1>ToolWeave · Stage 2</h1>
<p><strong>🎯 Progress-Reward Learning</strong></p>
<p>Turning correct tool use into measurable multi-turn task progress.</p>
<p>
<a href="https://github.com/Muradil-mamat-211/ToolWeave">🧵 Project</a>
</p>
</div>
> 🎯 **Curriculum role:** build on Stage 1 tool competence and optimize actual progress through multi-turn environment interaction.
## 🧭 At a glance
| Field | Details |
|---|---|
| 🧠 Base family | Qwen3-4B-Instruct |
| 🪜 Curriculum stage | Stage 2 — Progress-Reward Learning |
| 🧱 Starting point | ToolWeave Stage 1 update 25 |
| 📍 Checkpoint | Selected Stage 2 update 25 |
| 🎛️ Training signal | Fixed-denominator multi-turn Progress Reward |
| ✅ Release status | Selected checkpoint; not the final ToolWeave Stage 3 model |
ToolWeave Stage 2 trains on multi-turn BFCL environment tasks with a fixed-denominator Progress Reward, moving from correct tool execution toward reliable task completion.
## 📊 Evaluation (eval_400)
This is an internal ToolWeave validation on `val_400_combined`: 400 examples, with 100 examples each from Base, Long Context, Missing Function, and Missing Parameter. Validation used deterministic decoding (`n=1`, `do_sample=false`). The validation split is not included in this model repository, and these results are not official BFCL leaderboard results.
For Stage 2, `score` is the fixed-denominator Progress Reward and is equal to `progress` in this evaluation. It is therefore not directly comparable to the Stage 1 format-gate score.
### Overall
| Samples | Score / Progress | Format reward | Tool-call reward | Tool-call rate | Terminal coverage | Incomplete trajectories |
|---:|---:|---:|---:|---:|---:|---:|
| 400 | **0.4567** | 0.8582 | 0.9174 | 0.9275 | 0.8739 | 0.1450 |
### By evaluation category
| Split | Samples | Progress / Score | Terminal coverage | Incomplete trajectories |
|---|---:|---:|---:|---:|
| Base | 100 | **0.6027** | 0.9575 | 0.0500 |
| Long Context | 100 | 0.3952 | 0.7730 | 0.2600 |
| Missing Function | 100 | 0.4515 | 0.8868 | 0.1300 |
| Missing Parameter | 100 | 0.3774 | 0.8781 | 0.1400 |
## 🚀 Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "muradil211/ToolWeave_stage2"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
```
Tool-use inference requires the model's function schemas and the Qwen3-compatible tool-call format.
## 🔗 Links
- [🧵 ToolWeave project](https://github.com/Muradil-mamat-211/ToolWeave)