Instructions to use YangC777/AGW-35B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use YangC777/AGW-35B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="YangC777/AGW-35B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("YangC777/AGW-35B") model = AutoModelForMultimodalLM.from_pretrained("YangC777/AGW-35B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use YangC777/AGW-35B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "YangC777/AGW-35B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "YangC777/AGW-35B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/YangC777/AGW-35B
- SGLang
How to use YangC777/AGW-35B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "YangC777/AGW-35B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "YangC777/AGW-35B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "YangC777/AGW-35B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "YangC777/AGW-35B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use YangC777/AGW-35B with Docker Model Runner:
docker model run hf.co/YangC777/AGW-35B
Overview
AGW-35B is a screenshot-based computer-use agent obtained by fine-tuning
Qwen3.5-35B-A3B on 79,266 grounded and
filtered GUI training examples generated by AutoGUIWorld. Given a task instruction,
the current screenshot, and recent interaction history, the model predicts the next
executable computer_use action.
AutoGUIWorld expands GUI training coverage without executing every trajectory in the corresponding software environment. A planner specifies atomic actions and their intended visual consequences; an image generator then synthesizes successive GUI states. The training set spans Chrome, Ubuntu, Windows, and macOS, including browser, office, creative, system, and scientific workflows.
AGW-35B is trained as a direct-action, non-thinking agent. Its response contains a short action description followed by one or more structured tool calls. Spatial actions use a resolution-independent 0–999 coordinate grid, which the executor maps to the current screen resolution.
Model at a Glance
| Property | Value |
|---|---|
| Base model | Qwen3.5-35B-A3B |
| Architecture | Multimodal Mixture-of-Experts GUI agent |
| Primary input | Task instruction, screenshot, and recent interaction history |
| Primary output | Action: text and structured computer_use calls |
| Training examples | 79,266 step-level GUI examples |
| Training domains | Chrome, Ubuntu, Windows, macOS |
| SFT context length | 16,384 tokens |
| Weight format | BF16 Safetensors |
| Evaluation checkpoint | Step 469 |
Results
The final checkpoint is evaluated with the same task set as the base model. Interactive benchmarks allow up to 100 agent turns per task. OSWorld and Windows Agent Arena retain partial credit; macOSWorld and ScienceBoard report task success.
| Benchmark | Evaluation set | Metric | Qwen3.5-35B-A3B | AGW-35B | Gain |
|---|---|---|---|---|---|
| OSWorld | 361 tasks | Mean task score | 33.0 | 40.8 | +7.8 |
| Windows Agent Arena | 154 tasks | Mean task score | 19.4 | 27.9 | +8.5 |
| macOSWorld | 231 tasks | Task success | 28.1 | 45.0 | +16.9 |
| ScienceBoard | 143 tasks | Task success | 14.0 | 32.2 | +18.2 |
| Four-benchmark mean | Equal benchmark weight | Mean score | 23.6 | 36.5 | +12.8 |
AGW-35B also improves visual grounding on the 1,581-example English positive-target split of ScreenSpot-Pro.
| ScreenSpot-Pro | Qwen3.5-35B-A3B | AGW-35B | Gain |
|---|---|---|---|
| Overall | 31.7 | 57.1 | +25.4 |
| Icon targets | 16.2 | 32.9 | +16.7 |
| Text targets | 41.2 | 72.0 | +30.7 |
All values are percentages. The metrics differ across benchmarks and should be interpreted within each benchmark rather than pooled across individual tasks.
Training Data
The AutoGUIWorld training set contains 79,266 examples distributed across four GUI environments:
| Environment | Examples | Share |
|---|---|---|
| Chrome | 35,209 | 44.4% |
| Ubuntu | 24,250 | 30.6% |
| Windows | 14,778 | 18.6% |
| macOS | 5,029 | 6.3% |
Each example is derived from a multi-step screenshot trajectory and contains the task, current observation, recent action history, and the next action target. AutoGUIWorld provides spatial grounding for pointing actions and applies transition-level quality control before examples enter training.
Training Configuration
AGW-35B was trained with ms-swift Megatron on 32 GPUs across four nodes. The language
model was fine-tuned while the visual encoder was frozen.
| Setting | Value |
|---|---|
| Epochs | 3 |
| Peak learning rate | 1e-5 |
| Minimum learning rate | 1e-6 |
| Warmup fraction | 0.03 |
| Weight decay | 0.1 |
| Global batch size | 512 |
| Micro-batch size | 1 |
| Maximum sequence length | 16,384 |
| Tensor / pipeline / data parallelism | 2 / 1 / 16 |
| Expert parallelism | 8 |
Overlength examples were removed rather than truncated. The checkpoint uses the Qwen3.5 MoE architecture with 256 experts and 8 selected experts per token.
Interaction Contract
The model consumes screenshot-centered multi-turn messages and emits direct actions in the following form:
Action: Click the Settings icon.
<tool_call>
{"name":"computer_use","arguments":{"action":"left_click","coordinate":[927,54]}}
</tool_call>
The action space covers mouse movement and clicks, dragging, keyboard input, typing, scrolling, waiting, and task termination. Coordinates are integers on a 0–999 grid. An external agent harness must execute the tool call, capture the next screenshot, and append the resulting turn to the interaction history.
AGW-35B was trained without target-side chain-of-thought. For consistent behavior, disable thinking at inference and preserve the repository's chat template.
Serving with vLLM
Use a recent Qwen3.5-compatible vLLM build. The following configuration matches the 16,384-token SFT context and the tensor-parallel setup used in evaluation:
uv pip install vllm --torch-backend=auto \
--extra-index-url https://wheels.vllm.ai/nightly
vllm serve YangC777/AGW-35B \
--tensor-parallel-size 2 \
--max-model-len 16384 \
--dtype bfloat16
A minimal single-step request through the OpenAI-compatible endpoint is shown below. In a full agent loop, replace the compact system message with the project harness prompt and append previous screenshot-action turns.
import base64
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
with open("screenshot.png", "rb") as f:
screenshot = base64.b64encode(f.read()).decode("utf-8")
system_prompt = """You are a GUI agent. Given a task and screenshot, output exactly:
Action: <one short imperative>
<tool_call>
{"name":"computer_use","arguments":{"action":<action>, ...}}
</tool_call>
Use integer coordinates on a 0-999 grid. Use terminate when the task is complete."""
response = client.chat.completions.create(
model="YangC777/AGW-35B",
messages=[
{"role": "system", "content": system_prompt},
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {"url": f"data:image/png;base64,{screenshot}"},
},
{
"type": "text",
"text": "# Task Instruction:\nOpen Settings and switch to dark mode.\n\n"
"Generate the next move from the screenshot.",
},
],
},
],
temperature=0.0,
top_p=1.0,
max_tokens=1024,
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
print(response.choices[0].message.content)
Intended Use
AGW-35B is intended for research and development of screenshot-based GUI agents, including:
- cross-platform desktop and browser task execution;
- action grounding in dense professional interfaces;
- evaluation of synthetic GUI trajectories for agent post-training;
- agent systems that execute model-produced actions inside controlled environments.
The checkpoint is specialized for computer use. It should be paired with an executor, environment reset mechanism, action parser, and benchmark-appropriate safety controls.
Limitations
- The model observes rendered screenshots rather than hidden application or file-system state, so visually ambiguous states can lead to incorrect actions.
- Evaluation uses English task instructions and English interfaces; performance in other languages has not been established.
- Long workflows remain sensitive to early action errors and changes in application state.
- Performance varies by application domain. In the reported evaluation, Chrome scores decrease on both OSWorld and Windows Agent Arena even as aggregate performance improves.
- Training trajectories are generated through AutoGUIWorld and can inherit visual or transition artifacts from the generation pipeline.
Citation
@techreport{hunyuan2026autoguiworld,
title = {AutoGUIWorld: Image Generators as Visual World Models for GUI Agent},
author = {{Hunyuan AI Data Team}},
institution = {Tencent Hunyuan},
year = {2026}
}
Acknowledgements
AGW-35B is built on Qwen3.5-35B-A3B. We thank the authors and maintainers of Qwen, OSWorld, Windows Agent Arena, macOSWorld, ScienceBoard, and ScreenSpot-Pro.
- Downloads last month
- -
Model tree for YangC777/AGW-35B
Evaluation results
- Mean task score on OSWorldself-reported40.800
- Mean task score on Windows Agent Arenaself-reported27.900
- Task success on macOSWorldself-reported45.000
- Task success on ScienceBoardself-reported32.200
- Grounding accuracy on ScreenSpot-Proself-reported57.100