Papers
arxiv:2609.36887

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

Published on Sep 29
· Submitted by
Mao Bo
on Oct 5
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee commensurate gains in model performance, because reliable learning signals depend on coherent interactions among all components of the agentic interaction system. To address this problem, we introduce WEFT (Whole-system Evolution For Tool-use Post-training), which couples scalable agentic interaction system construction, execution-driven self-evolution, and stable post-training. WEFT scales agentic interaction system construction across environment breadth, task complexity, and interaction diversity. Execution-driven self-evolution iteratively uses execution traces and state evidence to attribute failures and revise the responsible components, with fresh rollouts evaluating the changes and providing evidence for subsequent evolution rounds. For stable post-training at scale, WEFT addresses both optimization and execution reliability: prefix-preserving sampling retains verified progress and atomic-turn credit assignment localizes learning signals, while MegaMCP maintains isolated, recoverable state across concurrent rollouts over shared tool services. Extensive experiments across various models and benchmarks demonstrate the effectiveness of WEFT for tool-use post-training. WEFT-8B and WEFT-14B outperform all evaluated matched-size environment-scaling baselines on BFCL V4, τ^2-Bench, and Claw-Eval. In particular, WEFT-14B improves over Agent-World-14B by 6.41, 2.23, and 12.27 percentage points. WEFT-35B-A3B further extends these gains to more challenging long-horizon workflow benchmarks, including Toolathlon-Verified and AutomationBench.

Community

Paper author Paper submitter

We're sharing WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents.

Scaling tool-use post-training requires more than synthesizing executable environments. Agents also need feasible tasks, suitable interaction setups, and reliable evaluation. WEFT (Whole-system Evolution For Tool-use Post-training) scales the complete agentic interaction system: environment, task, agent harness, and evaluator. It connects three stages: constructing richer training interactions, improving the system through execution, and supporting stable post-training.

1. Construct richer training interactions

WEFT jointly expands environment breadth, task complexity, and interaction diversity. We construct 8,172 executable MCPs, 64,755 tools, and 41,695 certified atomic tasks, then compose them into 11,884 tasks with a median length of 20 atomic-task turns. Most tasks span multiple MCPs, and more than half involve multiple domains.

The same grounded task can be presented as a complete brief for autonomous execution (Agentic), or introduced progressively by a simulated user (SimUser). WEFT then routes tasks to different native harnesses, including ReAct, OpenClaw, and Hermes, to collect diverse training trajectories while preserving task goals, dependencies, initial states, and completion conditions.

With the task set fixed, mixing Agentic and SimUser data improves WEFT-35B-A3B by 3.18 percentage points on average over Agentic alone. Adding OpenClaw and Hermes with the interaction-mode mixture fixed yields a further 9.71-point mean gain across Toolathlon-Verified, AutomationBench, and Claw-Eval. This offers a way to expand useful training data when new executable, verifiable tasks are difficult to obtain.

2. Improve the system through execution

A failed rollout is not necessarily a model failure. WEFT uses tool-call traces, checkpoint outcomes, and database and workspace changes to distinguish policy failures from environment, task, or verifier problems. Component problems trigger targeted revisions, dependency revalidation, and fresh rollouts that assess the changes and inform the next round.

With the task set and rollout budget fixed, three self-evolution rounds reduce tool-call errors in selected teacher trajectories from 1.76% to 0.96%. Toolathlon-Verified, AutomationBench, and Claw-Eval improve by 5.25, 4.33, and 3.65 percentage points, respectively. Execution therefore provides both policy-training data and evidence for improving the system that produces it.

3. Support stable post-training on long workflows

  • Prefix-preserving sampling: retain verified progress and retry a failed atomic task from its pre-turn history, database, and workspace, rather than discarding the entire trajectory.
  • Task-verifier consistency filtering: overly strict verifiers can reject valid solutions, while overly permissive ones can reward incorrect or incomplete execution. Before RL, a rubric-guided LLM independently scores pilot rollouts. Tasks are retained only when both Pearson and Spearman correlations with executable verifier scores meet the selection criteria. The LLM is used for offline filtering; RL still uses executable rewards.
  • Atomic-turn credit assignment: compare candidate segments sampled from the same history and state, and assign credit using the current atomic task's binary completion outcome.

MegaMCP supports these procedures at rollout scale. It hosts reusable tool services outside agent sandboxes while keeping each rollout's database and workspace private and recoverable. Snapshots support retries and branching without replaying completed interactions. In the 1,000-task experiment, sandbox upload volume falls from 773.5 to 34.8 MiB (95.5% less).

Results

WEFT-8B and WEFT-14B outperform all evaluated matched-size environment-scaling baselines on BFCL V4, τ²-Bench, and Claw-Eval. WEFT-14B improves over Agent-World-14B by 6.41, 2.23, and 12.27 percentage points, respectively. WEFT-35B-A3B reaches 45.99% on Toolathlon-Verified and 26.50% on AutomationBench.

The central idea is to scale the system that produces and uses training interactions: broaden the tasks and trajectories, improve their feedback through execution, and make long-horizon post-training efficient and stable.

Paper · Illustrated overview on alphaXiv

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.36887 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.36887 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.36887 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.