CompoWorld: Compositional Environment Scaling for General Agents

📄 Paper  |  📚 arXiv  |  💻 GitHub

This repository contains Qwen3.6-35B-A3B agent trained with CompoWorld, a compositional environment scaling method for training general agents.

TL;DR

Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services.

CompoWorld expands the task space by composing a finite library of reusable services:

  • Verified services: coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented.
  • Compositional task generation: a random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services.
  • SFT + RL: verified trajectories support supervised fine-tuning (SFT), while a Completion-Focused Rubric Reward guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group.

We construct 448 services exposing 10,130 tools and use 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B.

Results

CompoWorld improves on its backbone by +9.17 points on average across eight agent benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.

Main results on eight challenging agent benchmarks (from the paper):

Model τ³-Banking DeepPlanning VitaBench VitaBench 2.0 AutomationBench WildClawBench SkillsBench ALE
Frontier Closed-Source Models
GPT-5.4 28.52 53.96 47.09 45.99 27.67 58.00 51.70 22.50
Claude Opus 4.6 20.27 55.21 38.31 37.71 25.50 54.60 50.20 13.30
Gemini-3.1 Pro 23.71 47.08 52.69 49.53 28.17 38.70 60.80 17.10
Open-Weight Models
DeepSeek-V4-Flash 30.34 54.79 56.31 41.12 36.33 47.67 53.75 17.48
GLM-5.2 28.87 50.00 50.26 45.32 26.33 54.20 62.10 22.33
Kimi-K2.6 19.93 43.12 43.94 44.30 15.17 29.40 54.00 8.74
Qwen3.8-27B 33.68 52.92 41.84 47.23 38.50 56.20 35.57 20.40
Qwen3.5-397B-A17B 16.15 35.83 42.09 38.41 5.50 40.37 36.50 9.71
Qwen3.6-35B-A3B 10.65 26.04 38.94 34.47 10.33 44.28 32.52 7.77
Agent-Specialized Models (35B-A3B)
Apodex 1.1 Mini 12.71 34.17 44.06 35.29 14.17 - 32.29 -
Occamy-1.0 37.10 - 41.75 - 27.60 49.16 - -
Agents-A1 7.20 - 37.00 - 2.20 30.73 - -
Nex-N2-mini 25.80 - 26.25 - 5.70 30.31 - -
BigBang-1.0 10.30 - 46.00 - 14.80 32.87 - -
Ornith-1.5-35B 21.70 - 40.25 - 18.50 45.91 - -
CompoWorld (Ours) 16.49 35.21 49.44 36.09 32.33 47.46 47.71 13.59
Δ vs. backbone +5.84 +9.17 +10.50 +1.62 +22.00 +3.18 +15.19 +5.82

Evaluation protocol. We use OpenHands for SkillsBench with skills enabled, Claude Code for ALE, and OpenClaw for WildClawBench. We report the average score or reward for WildClawBench and both VitaBench versions, Avg@3 for SkillsBench, the pass rate for τ³-Banking, AutomationBench, and ALE, and the average accuracy for DeepPlanning.

Model Description

  • Base model: Qwen/Qwen3.6-35B-A3B (35B total parameters, 3B activated, MoE with vision encoder, 262,144-token context)
  • Training data: 3K verified SFT trajectories + 1K RL tasks generated from 448 composed services (10,130 tools)
  • RL reward: Completion-Focused Rubric Reward — emphasizes criteria with lower pass rates within each rollout group
  • Format: Hugging Face Transformers (compatible with Transformers, vLLM, SGLang, KTransformers, etc.)

Quickstart

vLLM

uv pip install vllm --torch-backend=auto

vllm serve AllSpark-Research/CompoWorld --port 8000 --tp-size 8 \
  --tool-call-parser qwen3_coder --reasoning-parser qwen3 \
  --context-length 262144

SGLang

uv pip install sglang[all]

python -m sglang.launch_server --model-path AllSpark-Research/CompoWorld \
  --port 8000 --tp-size 8 --mem-fraction-static 0.8 \
  --context-length 262144 --reasoning-parser qwen3 --tool-call-parser qwen3_coder

Transformers

from transformers import AutoModelForImageTextToText, AutoProcessor

model = AutoModelForImageTextToText.from_pretrained(
    "AllSpark-Research/CompoWorld", torch_dtype="auto", device_map="auto"
)
processor = AutoProcessor.from_pretrained("AllSpark-Research/CompoWorld")

The model has a default context length of 262,144 tokens. We advise maintaining a context length of at least 128K tokens to preserve thinking capabilities.

License

This model is released under the Apache 2.0 license, inherited from the base model.

Citation

If you find this work useful, please cite:

@article{yang2026compoworld,
  title={CompoWorld: Compositional Environment Scaling for General Agents},
  author={Yang, Xiao-Wen and Xu, Weiyi and Da, Wen and Xu, Hang and Li, Canwei and You, Hong-Jie and Dong, Pusen and Zeng, Yucheng and Luo, Zhaokai and Li, Yu-Feng and Hu, Yao and Chuan, Mu},
  journal={arXiv preprint arXiv:2609.33665},
  year={2026}
}
Downloads last month
-
Safetensors
Model size
36B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AllSpark-Research/CompoWorld

Finetuned
(321)
this model

Paper for AllSpark-Research/CompoWorld