--- base_model: - Qwen/Qwen3.6-35B-A3B library_name: transformers license: apache-2.0 license_link: https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/main/LICENSE pipeline_tag: image-text-to-text tags: - agent - tool-use - reinforcement-learning - sft - environment-scaling --- # CompoWorld: Compositional Environment Scaling for General Agents
This repository contains **Qwen3.6-35B-A3B agent trained with CompoWorld**, a compositional environment scaling method for training general agents. ## TL;DR Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. **CompoWorld** expands the task space by composing a finite library of reusable services: - **Verified services**: coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented. - **Compositional task generation**: a random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services. - **SFT + RL**: verified trajectories support supervised fine-tuning (SFT), while a **Completion-Focused Rubric Reward** guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group. We construct **448 services exposing 10,130 tools** and use **3K SFT trajectories** and **1K RL tasks** to train Qwen3.6-35B-A3B. ## Results CompoWorld improves on its backbone by **+9.17 points on average** across eight agent benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models. **Main results on eight challenging agent benchmarks** (from the [paper](https://huggingface.co/papers/2609.33665)): | Model | τ³-Banking | DeepPlanning | VitaBench | VitaBench 2.0 | AutomationBench | WildClawBench | SkillsBench | ALE | |---|---|---|---|---|---|---|---|---| | **Frontier Closed-Source Models** | | | | | | | | | | GPT-5.4 | 28.52 | 53.96 | 47.09 | 45.99 | 27.67 | 58.00 | 51.70 | 22.50 | | Claude Opus 4.6 | 20.27 | 55.21 | 38.31 | 37.71 | 25.50 | 54.60 | 50.20 | 13.30 | | Gemini-3.1 Pro | 23.71 | 47.08 | 52.69 | 49.53 | 28.17 | 38.70 | 60.80 | 17.10 | | **Open-Weight Models** | | | | | | | | | | DeepSeek-V4-Flash | 30.34 | 54.79 | 56.31 | 41.12 | 36.33 | 47.67 | 53.75 | 17.48 | | GLM-5.2 | 28.87 | 50.00 | 50.26 | 45.32 | 26.33 | 54.20 | 62.10 | 22.33 | | Kimi-K2.6 | 19.93 | 43.12 | 43.94 | 44.30 | 15.17 | 29.40 | 54.00 | 8.74 | | Qwen3.8-27B | 33.68 | 52.92 | 41.84 | 47.23 | 38.50 | 56.20 | 35.57 | 20.40 | | Qwen3.5-397B-A17B | 16.15 | 35.83 | 42.09 | 38.41 | 5.50 | 40.37 | 36.50 | 9.71 | | Qwen3.6-35B-A3B | 10.65 | 26.04 | 38.94 | 34.47 | 10.33 | 44.28 | 32.52 | 7.77 | | **Agent-Specialized Models (35B-A3B)** | | | | | | | | | | Apodex 1.1 Mini | 12.71 | 34.17 | 44.06 | 35.29 | 14.17 | - | 32.29 | - | | Occamy-1.0 | 37.10 | - | 41.75 | - | 27.60 | 49.16 | - | - | | Agents-A1 | 7.20 | - | 37.00 | - | 2.20 | 30.73 | - | - | | Nex-N2-mini | 25.80 | - | 26.25 | - | 5.70 | 30.31 | - | - | | BigBang-1.0 | 10.30 | - | 46.00 | - | 14.80 | 32.87 | - | - | | Ornith-1.5-35B | 21.70 | - | 40.25 | - | 18.50 | 45.91 | - | - | | **CompoWorld (Ours)** | 16.49 | 35.21 | 49.44 | 36.09 | **32.33** | 47.46 | 47.71 | 13.59 | | Δ vs. backbone | +5.84 | +9.17 | +10.50 | +1.62 | +22.00 | +3.18 | +15.19 | +5.82 | **Evaluation protocol.** We use OpenHands for SkillsBench with skills enabled, Claude Code for ALE, and OpenClaw for WildClawBench. We report the average score or reward for WildClawBench and both VitaBench versions, Avg@3 for SkillsBench, the pass rate for τ³-Banking, AutomationBench, and ALE, and the average accuracy for DeepPlanning. ## Model Description - **Base model:** [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) (35B total parameters, 3B activated, MoE with vision encoder, 262,144-token context) - **Training data:** 3K verified SFT trajectories + 1K RL tasks generated from 448 composed services (10,130 tools) - **RL reward:** Completion-Focused Rubric Reward — emphasizes criteria with lower pass rates within each rollout group - **Format:** Hugging Face Transformers (compatible with Transformers, vLLM, SGLang, KTransformers, etc.) ## Quickstart ### vLLM ```shell uv pip install vllm --torch-backend=auto vllm serve AllSpark-Research/CompoWorld --port 8000 --tp-size 8 \ --tool-call-parser qwen3_coder --reasoning-parser qwen3 \ --context-length 262144 ``` ### SGLang ```shell uv pip install sglang[all] python -m sglang.launch_server --model-path AllSpark-Research/CompoWorld \ --port 8000 --tp-size 8 --mem-fraction-static 0.8 \ --context-length 262144 --reasoning-parser qwen3 --tool-call-parser qwen3_coder ``` ### Transformers ```python from transformers import AutoModelForImageTextToText, AutoProcessor model = AutoModelForImageTextToText.from_pretrained( "AllSpark-Research/CompoWorld", torch_dtype="auto", device_map="auto" ) processor = AutoProcessor.from_pretrained("AllSpark-Research/CompoWorld") ``` > [!Note] > The model has a default context length of 262,144 tokens. We advise maintaining a context length of at least 128K tokens to preserve thinking capabilities. ## License This model is released under the [Apache 2.0 license](https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/main/LICENSE), inherited from the base model. ## Citation If you find this work useful, please cite: ```bibtex @article{yang2026compoworld, title={CompoWorld: Compositional Environment Scaling for General Agents}, author={Yang, Xiao-Wen and Xu, Weiyi and Da, Wen and Xu, Hang and Li, Canwei and You, Hong-Jie and Dong, Pusen and Zeng, Yucheng and Luo, Zhaokai and Li, Yu-Feng and Hu, Yao and Chuan, Mu}, journal={arXiv preprint arXiv:2609.33665}, year={2026} } ```