RefinedNeuro's picture
RefinedToolCall-V5-3B model card
08058b1 verified
|
Raw
History Blame Contribute Delete
3.88 kB
---
license: apache-2.0
base_model: WeiboAI/VibeThinker-3B
datasets:
- lambda/hermes-agent-reasoning-traces
language:
- en
pipeline_tag: text-generation
tags:
- tool-calling
- function-calling
- multi-turn
- agentic
- hermes
- reasoning
- qwen2
---
# ๐Ÿ› ๏ธ๐Ÿง  RefinedToolCall-V5-3B
### A 3B model that *reasons* and *calls tools* โ€” and actually holds a multi-turn conversation.
**Math-grade reasoning ยท real function calling ยท multi-turn agentic ยท 2.5 GB ยท runs on your laptop.**
> `ollama run refinedneuro/refinedtoolcallv5-3b`
---
## Why it's different
Most 3B tool-callers nail a single function call and then fall apart the moment the task spans
several turns. **RefinedToolCall-V5** was built specifically to fix that โ€” and the numbers moved on
*every* axis at once, not just the one we were targeting.
- ๐Ÿ” **Multi-turn agentic that actually works** โ€” **~3.7ร— better** at stateful, multi-step
tool-use (Berkeley Function-Calling Leaderboard `multi_turn`) than where we started.
- ๐Ÿ› ๏ธ **Sharper single-turn calling** โ€” **70.7%** on BFCL single-turn (held-out), our best ever.
- ๐Ÿ’ช **Recovers from tool errors** โ€” **0.896** recovery rate; it diagnoses failures instead of
looping on them.
- ๐Ÿงฎ **Reasoning fully intact** โ€” **AIME-2024 pass@8 0.933**, unchanged by all the tool training.
- โšก **Tiny & local** โ€” 3B params, **2.5 GB** Q6_K, one command on Ollama, no GPU required.
- ๐Ÿ†“ **Apache-2.0** โ€” use it, ship it, fine-tune it.
---
## The receipts (all held-out, canary-gated)
| capability | this model |
|---|---|
| ๐Ÿ” Multi-turn agentic (BFCL `multi_turn`, k=3) | **0.220 avg / 0.298 pass@3** |
| ๐Ÿ› ๏ธ Single-turn function calling (BFCL, held-out) | **0.707** |
| ๐Ÿฉน Recovery from tool errors (n=250) | **0.896** |
| ๐Ÿงฎ Reasoning (AIME-2024 pass@8) | **0.933** |
Every number is the **best across five fine-tuning rounds** โ€” multi-turn, single-turn, recovery,
*and* reasoning all peaked together.
---
## How we got here (and why it generalizes)
We didn't just throw data at it. Five disciplined rounds, each one gated so it could **never**
regress reasoning or recovery:
1. **Grounding** โ€” stop inventing shell commands; call the actual functions.
2. **Plan + finish** โ€” think before calling, and know when the turn is done.
3. **Scale + long context** โ€” harder tasks, up to 24k tokens.
4. **On-policy self-improvement (the breakthrough)** โ€” the model learns from its *own* successful
multi-turn solutions (expert iteration), which broke past the imitation ceiling **and** sharpened
single-turn calling and error-recovery as a bonus.
---
## Quick start
**Ollama**
```bash
ollama run refinedneuro/refinedtoolcallv5-3b # latest = Q6_K, 2.5 GB
```
> ๐Ÿ’ก Use **Q6_K or higher** for tool-calling โ€” lower quants corrupt the call tokens.
**Format:** ChatML + Hermes tools. Each turn the model emits a `<think>` plan โ†’ one or more
`<tool_call>` blocks โ†’ a final reply. Recommended: temp 0.6, top_p 0.95, repeat_penalty 1.1.
---
## Great for
โœ… Local/offline agentic tool-use prototypes โœ… Multi-step function-calling assistants
โœ… Math & STEM reasoning โœ… Learning how small agentic models are actually built.
## Be honest with me (research preview)
โš ๏ธ It's a **3B research preview**. Multi-turn is **dramatically improved (~3.7ร—) but not solved** โ€”
very long, open-ended autonomous loops can still write buggy code or mis-plan. A brilliant,
tiny building block; not yet a drop-in autonomous engineer.
---
*Built on [WeiboAI/VibeThinker-3B](https://huggingface.co/WeiboAI/VibeThinker-3B) +
[lambda/hermes-agent-reasoning-traces](https://huggingface.co/datasets/lambda/hermes-agent-reasoning-traces).
Trained with distribution-matched RFT + on-policy expert iteration, every checkpoint gated against
reasoning/recovery canaries. Apache-2.0.*