--- license: apache-2.0 base_model: WeiboAI/VibeThinker-3B datasets: - lambda/hermes-agent-reasoning-traces language: - en pipeline_tag: text-generation tags: - tool-calling - function-calling - multi-turn - agentic - hermes - reasoning - qwen2 --- # ๐Ÿ› ๏ธ๐Ÿง  RefinedToolCall-V5-3B ### A 3B model that *reasons* and *calls tools* โ€” and actually holds a multi-turn conversation. **Math-grade reasoning ยท real function calling ยท multi-turn agentic ยท 2.5 GB ยท runs on your laptop.** > `ollama run refinedneuro/refinedtoolcallv5-3b` --- ## Why it's different Most 3B tool-callers nail a single function call and then fall apart the moment the task spans several turns. **RefinedToolCall-V5** was built specifically to fix that โ€” and the numbers moved on *every* axis at once, not just the one we were targeting. - ๐Ÿ” **Multi-turn agentic that actually works** โ€” **~3.7ร— better** at stateful, multi-step tool-use (Berkeley Function-Calling Leaderboard `multi_turn`) than where we started. - ๐Ÿ› ๏ธ **Sharper single-turn calling** โ€” **70.7%** on BFCL single-turn (held-out), our best ever. - ๐Ÿ’ช **Recovers from tool errors** โ€” **0.896** recovery rate; it diagnoses failures instead of looping on them. - ๐Ÿงฎ **Reasoning fully intact** โ€” **AIME-2024 pass@8 0.933**, unchanged by all the tool training. - โšก **Tiny & local** โ€” 3B params, **2.5 GB** Q6_K, one command on Ollama, no GPU required. - ๐Ÿ†“ **Apache-2.0** โ€” use it, ship it, fine-tune it. --- ## The receipts (all held-out, canary-gated) | capability | this model | |---|---| | ๐Ÿ” Multi-turn agentic (BFCL `multi_turn`, k=3) | **0.220 avg / 0.298 pass@3** | | ๐Ÿ› ๏ธ Single-turn function calling (BFCL, held-out) | **0.707** | | ๐Ÿฉน Recovery from tool errors (n=250) | **0.896** | | ๐Ÿงฎ Reasoning (AIME-2024 pass@8) | **0.933** | Every number is the **best across five fine-tuning rounds** โ€” multi-turn, single-turn, recovery, *and* reasoning all peaked together. --- ## How we got here (and why it generalizes) We didn't just throw data at it. Five disciplined rounds, each one gated so it could **never** regress reasoning or recovery: 1. **Grounding** โ€” stop inventing shell commands; call the actual functions. 2. **Plan + finish** โ€” think before calling, and know when the turn is done. 3. **Scale + long context** โ€” harder tasks, up to 24k tokens. 4. **On-policy self-improvement (the breakthrough)** โ€” the model learns from its *own* successful multi-turn solutions (expert iteration), which broke past the imitation ceiling **and** sharpened single-turn calling and error-recovery as a bonus. --- ## Quick start **Ollama** ```bash ollama run refinedneuro/refinedtoolcallv5-3b # latest = Q6_K, 2.5 GB ``` > ๐Ÿ’ก Use **Q6_K or higher** for tool-calling โ€” lower quants corrupt the call tokens. **Format:** ChatML + Hermes tools. Each turn the model emits a `` plan โ†’ one or more `` blocks โ†’ a final reply. Recommended: temp 0.6, top_p 0.95, repeat_penalty 1.1. --- ## Great for โœ… Local/offline agentic tool-use prototypes โœ… Multi-step function-calling assistants โœ… Math & STEM reasoning โœ… Learning how small agentic models are actually built. ## Be honest with me (research preview) โš ๏ธ It's a **3B research preview**. Multi-turn is **dramatically improved (~3.7ร—) but not solved** โ€” very long, open-ended autonomous loops can still write buggy code or mis-plan. A brilliant, tiny building block; not yet a drop-in autonomous engineer. --- *Built on [WeiboAI/VibeThinker-3B](https://huggingface.co/WeiboAI/VibeThinker-3B) + [lambda/hermes-agent-reasoning-traces](https://huggingface.co/datasets/lambda/hermes-agent-reasoning-traces). Trained with distribution-matched RFT + on-policy expert iteration, every checkpoint gated against reasoning/recovery canaries. Apache-2.0.*