fqjiang's picture
initial release
eb856d8
|
Raw
History Blame Contribute Delete
2.02 kB
metadata
license: apache-2.0
base_model:
  - Qwen/Qwen3-4B-Base
library_name: transformers
pipeline_tag: text-generation
datasets:
  - MidTool/MidTool-Mix
tags:
  - agentic
  - tool-use
  - mid-training
  - function-calling
extra_gated_heading: Access this model
extra_gated_prompt: >-
  By requesting access you agree to the Apache-2.0 license, and to the terms of
  the MidTool-Mix dataset this model was trained on
  (https://huggingface.co/datasets/MidTool/MidTool-Mix/blob/main/LICENSE). The
  model is provided as is, without warranty of any kind. The authors accept no
  liability for its outputs or for any use made of it.
extra_gated_button_content: Agree and access

Arctic-MidTool-MT-4B

Qwen3-4B-Base mid-trained on MidTool-Mix, a 20.3B-token corpus for agentic tool use.

This is the mid-training checkpoint: a base model with a stronger tool-use prior, not an instruction-tuned assistant. Use it as the starting point for your own tool-use SFT/RL. For a ready-to-use agent, see Arctic-MidTool-RL-4B.

Results

Both rows use the same downstream SFT recipe, so the difference isolates the effect of mid-training.

4B setting BFCLv3 Overall τ²-Bench Pass@1 MCP-Universe Score
Qwen3-4B-Base + SFT 39.73 8.54 13.20
Arctic-MidTool-MT-4B + SFT 50.25 12.23 18.66

Details

Mid-trained on MidTool-Mix; see the dataset's LICENSE for data terms.

See our paper for the full data, training, and evaluation details.

@article{jiang2026midtool,
  title  = {MidTool: Mid-training Data Synthesis for Agentic Tool Use},
  author = {Jiang, Fengqing and Wang, Yite and Liu, Boyi and Wang, Zhaoyang and
            Xu, Canwen and Yao, Zhewei and Poovendran, Radha and He, Yuxiong},
  year   = {2026}
}