AGW-35B

A cross-platform GUI agent trained on synthetic interaction trajectories from AutoGUIWorld

Project · Model Weights

AutoGUIWorld overview

Overview

AGW-35B is a screenshot-based computer-use agent obtained by fine-tuning Qwen3.5-35B-A3B on 79,266 grounded and filtered GUI training examples generated by AutoGUIWorld. Given a task instruction, the current screenshot, and recent interaction history, the model predicts the next executable computer_use action.

AutoGUIWorld expands GUI training coverage without executing every trajectory in the corresponding software environment. A planner specifies atomic actions and their intended visual consequences; an image generator then synthesizes successive GUI states. The training set spans Chrome, Ubuntu, Windows, and macOS, including browser, office, creative, system, and scientific workflows.

AGW-35B is trained as a direct-action, non-thinking agent. Its response contains a short action description followed by one or more structured tool calls. Spatial actions use a resolution-independent 0–999 coordinate grid, which the executor maps to the current screen resolution.

Model at a Glance

Property Value
Base model Qwen3.5-35B-A3B
Architecture Multimodal Mixture-of-Experts GUI agent
Primary input Task instruction, screenshot, and recent interaction history
Primary output Action: text and structured computer_use calls
Training examples 79,266 step-level GUI examples
Training domains Chrome, Ubuntu, Windows, macOS
SFT context length 16,384 tokens
Weight format BF16 Safetensors
Evaluation checkpoint Step 469

Results

The final checkpoint is evaluated with the same task set as the base model. Interactive benchmarks allow up to 100 agent turns per task. OSWorld and Windows Agent Arena retain partial credit; macOSWorld and ScienceBoard report task success.

Benchmark Evaluation set Metric Qwen3.5-35B-A3B AGW-35B Gain
OSWorld 361 tasks Mean task score 33.0 40.8 +7.8
Windows Agent Arena 154 tasks Mean task score 19.4 27.9 +8.5
macOSWorld 231 tasks Task success 28.1 45.0 +16.9
ScienceBoard 143 tasks Task success 14.0 32.2 +18.2
Four-benchmark mean Equal benchmark weight Mean score 23.6 36.5 +12.8

AGW-35B also improves visual grounding on the 1,581-example English positive-target split of ScreenSpot-Pro.

ScreenSpot-Pro Qwen3.5-35B-A3B AGW-35B Gain
Overall 31.7 57.1 +25.4
Icon targets 16.2 32.9 +16.7
Text targets 41.2 72.0 +30.7

All values are percentages. The metrics differ across benchmarks and should be interpreted within each benchmark rather than pooled across individual tasks.

Training Data

The AutoGUIWorld training set contains 79,266 examples distributed across four GUI environments:

Environment Examples Share
Chrome 35,209 44.4%
Ubuntu 24,250 30.6%
Windows 14,778 18.6%
macOS 5,029 6.3%

Each example is derived from a multi-step screenshot trajectory and contains the task, current observation, recent action history, and the next action target. AutoGUIWorld provides spatial grounding for pointing actions and applies transition-level quality control before examples enter training.

Training Configuration

AGW-35B was trained with ms-swift Megatron on 32 GPUs across four nodes. The language model was fine-tuned while the visual encoder was frozen.

Setting Value
Epochs 3
Peak learning rate 1e-5
Minimum learning rate 1e-6
Warmup fraction 0.03
Weight decay 0.1
Global batch size 512
Micro-batch size 1
Maximum sequence length 16,384
Tensor / pipeline / data parallelism 2 / 1 / 16
Expert parallelism 8

Overlength examples were removed rather than truncated. The checkpoint uses the Qwen3.5 MoE architecture with 256 experts and 8 selected experts per token.

Interaction Contract

The model consumes screenshot-centered multi-turn messages and emits direct actions in the following form:

Action: Click the Settings icon.
<tool_call>
{"name":"computer_use","arguments":{"action":"left_click","coordinate":[927,54]}}
</tool_call>

The action space covers mouse movement and clicks, dragging, keyboard input, typing, scrolling, waiting, and task termination. Coordinates are integers on a 0–999 grid. An external agent harness must execute the tool call, capture the next screenshot, and append the resulting turn to the interaction history.

AGW-35B was trained without target-side chain-of-thought. For consistent behavior, disable thinking at inference and preserve the repository's chat template.

Serving with vLLM

Use a recent Qwen3.5-compatible vLLM build. The following configuration matches the 16,384-token SFT context and the tensor-parallel setup used in evaluation:

uv pip install vllm --torch-backend=auto \
  --extra-index-url https://wheels.vllm.ai/nightly

vllm serve YangC777/AGW-35B \
  --tensor-parallel-size 2 \
  --max-model-len 16384 \
  --dtype bfloat16

A minimal single-step request through the OpenAI-compatible endpoint is shown below. In a full agent loop, replace the compact system message with the project harness prompt and append previous screenshot-action turns.

import base64
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

with open("screenshot.png", "rb") as f:
    screenshot = base64.b64encode(f.read()).decode("utf-8")

system_prompt = """You are a GUI agent. Given a task and screenshot, output exactly:
Action: <one short imperative>
<tool_call>
{"name":"computer_use","arguments":{"action":<action>, ...}}
</tool_call>
Use integer coordinates on a 0-999 grid. Use terminate when the task is complete."""

response = client.chat.completions.create(
    model="YangC777/AGW-35B",
    messages=[
        {"role": "system", "content": system_prompt},
        {
            "role": "user",
            "content": [
                {
                    "type": "image_url",
                    "image_url": {"url": f"data:image/png;base64,{screenshot}"},
                },
                {
                    "type": "text",
                    "text": "# Task Instruction:\nOpen Settings and switch to dark mode.\n\n"
                            "Generate the next move from the screenshot.",
                },
            ],
        },
    ],
    temperature=0.0,
    top_p=1.0,
    max_tokens=1024,
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)

print(response.choices[0].message.content)

Intended Use

AGW-35B is intended for research and development of screenshot-based GUI agents, including:

  • cross-platform desktop and browser task execution;
  • action grounding in dense professional interfaces;
  • evaluation of synthetic GUI trajectories for agent post-training;
  • agent systems that execute model-produced actions inside controlled environments.

The checkpoint is specialized for computer use. It should be paired with an executor, environment reset mechanism, action parser, and benchmark-appropriate safety controls.

Limitations

  • The model observes rendered screenshots rather than hidden application or file-system state, so visually ambiguous states can lead to incorrect actions.
  • Evaluation uses English task instructions and English interfaces; performance in other languages has not been established.
  • Long workflows remain sensitive to early action errors and changes in application state.
  • Performance varies by application domain. In the reported evaluation, Chrome scores decrease on both OSWorld and Windows Agent Arena even as aggregate performance improves.
  • Training trajectories are generated through AutoGUIWorld and can inherit visual or transition artifacts from the generation pipeline.

Citation

@techreport{hunyuan2026autoguiworld,
  title       = {AutoGUIWorld: Image Generators as Visual World Models for GUI Agent},
  author      = {{Hunyuan AI Data Team}},
  institution = {Tencent Hunyuan},
  year        = {2026}
}

Acknowledgements

AGW-35B is built on Qwen3.5-35B-A3B. We thank the authors and maintainers of Qwen, OSWorld, Windows Agent Arena, macOSWorld, ScienceBoard, and ScreenSpot-Pro.

Downloads last month
-
Safetensors
Model size
35B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YangC777/AGW-35B

Finetuned
(170)
this model

Evaluation results