How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="muradil211/ToolWeave_stage3")
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)
# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("muradil211/ToolWeave_stage3")
model = AutoModelForCausalLM.from_pretrained("muradil211/ToolWeave_stage3", device_map="auto")
messages = [
    {"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))
Quick Links
ToolWeave mark

ToolWeave · Stage 3

🧭 Boundary-Guided Online Reinforcement Learning

Verified online data synthesis for multi-turn tool-calling agents.

🧵 Project

🧭 At a glance

Field Details
🧠 Base family Qwen3-4B-Instruct
🪜 Curriculum stage Stage 3 — Boundary-Guided Online Reinforcement Learning
🧱 Starting point ToolWeave Stage 2 update 25
🎛️ Training signal Verified online data synthesis + multi-turn Progress Reward
✅ Release status Final ToolWeave Stage 3 model

ToolWeave Stage 3 expands multi-turn tool-use learning through capability-boundary detection, verified online data synthesis, strict execution and semantic validation, dynamic replay, and combined global/local tool-call credit.

📊 Stage 3 evaluation

The final ToolWeave Stage 3 checkpoint was evaluated on the canonical balanced 400-row held-in set: 100 entries each from Base, Missing Function, Missing Parameter, and Long Context. These values are complete-entry BFCL Multi-Turn accuracies, not the training-time Progress Reward (R_P).

Model Overall Base Missing Function Missing Parameter Long Context Correct entries
ToolWeave Stage 3 48.50 56.00 50.00 42.00 46.00 194 / 400

Because the four categories are balanced, the overall score is their unweighted mean and the complete-entry accuracy over all 400 entries:

(56.00 + 50.00 + 42.00 + 46.00) / 4 = 48.50

🚀 Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "muradil211/ToolWeave_stage3"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")

Tool-use inference requires the model's function schemas and the Qwen3-compatible tool-call format.

🔗 Links

Downloads last month
666
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support