File size: 3,595 Bytes
8a6f888
 
b5916f8
fda0b76
8a6f888
b5916f8
 
8a6f888
b5916f8
 
 
 
 
 
 
 
8a6f888
 
 
 
bf29e6a
8a6f888
bf29e6a
8a6f888
bf29e6a
8a6f888
bf29e6a
8a6f888
bf29e6a
 
 
 
8a6f888
bf29e6a
8a6f888
bf29e6a
8a6f888
 
 
 
 
bf29e6a
8a6f888
 
 
 
 
 
 
 
 
 
 
bf29e6a
8a6f888
bf29e6a
8a6f888
bf29e6a
 
 
 
 
 
 
 
 
 
 
8a6f888
6883941
8a6f888
bf29e6a
8a6f888
bf29e6a
8a6f888
bf29e6a
8a6f888
bf29e6a
8a6f888
bf29e6a
8a6f888
 
 
 
 
 
 
b5916f8
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
---
language:
- en
license: apache-2.0
pipeline_tag: image-text-to-text
base_model:
- k4ng/SCOPE-SFT-9B
tags:
- computer-use-agent
- gui-agent
- multimodal
- reinforcement-learning
- grpo
- safety
- osworld
- os-blind
---

# SCOPE-RL-9B

**SCOPE-RL-9B** is a computer-use agent (CUA) trained under the SCOPE (Safety and Capability Optimization for Policy Execution) framework to balance task-execution capability with safety-aware decision-making. The model is initialized from **SCOPE-SFT-9B** and further optimized through online reinforcement learning on verifiable capability tasks.

In our evaluation, SCOPE-RL-9B achieves a **54.17%** task success rate on OSWorld and a **64.30%** attack-avoidance rate on OS-BLIND, corresponding to a capability-safety harmonic mean of **58.80%**. Under our evaluation setting and among the models listed below, SCOPE-RL-9B achieves the best overall balance between capability and safety.

SCOPE-RL-9B is trained with programmatically verified capability tasks generated by SCOPE-Gen, using Safactory as the online reinforcement learning framework. Its initialization checkpoint, SCOPE-SFT-9B, was previously trained on capability demonstrations, safe-continuation trajectories, and explicit-refusal trajectories.

## Links

- Paper: [Beyond Task Completion: Training Capable and Safe Computer-Use Agents]()
- Data generation code: [SCOPE-Gen](https://github.com/k4ngzy/SCOPE-Gen)
- RL training code: [Safactory](https://github.com/AI45Lab/SAfactory)
- Model collection: [SCOPE](https://huggingface.co/collections/k4ng/scope)

## Quick Start

Install vLLM:

```bash
pip install -U vllm
```

Launch an OpenAI-compatible inference server:

```bash
vllm serve k4ng/SCOPE-RL-9B \
    --host 0.0.0.0 \
    --port 8000 \
    --tensor-parallel-size 1 \
    --data-parallel-size 2 \
    --trust-remote-code \
    --served-model-name scope-rl
```

## Results

| Type | Model | H ↑ | OSWorld ↑ | OS-BLIND ↑ |
|---|---|---:|---:|---:|
| Closed-source | Claude 4.5 Sonnet | 37.78 | 62.90 | 27.00 |
| Closed-source | Qwen3.7-Plus | 9.36 | **73.33** | 5.00 |
| Open-source | EvoCUA-8B | 10.62 | 46.06 | 6.00 |
| Open-source | EvoCUA-32B | 4.42 | 56.73 | 2.30 |
| Open-source | OpenCUA-7B | 3.21 | 28.85 | 1.70 |
| Open-source | OpenCUA-32B | 1.94 | 34.79 | 1.00 |
| Open-source | OpenCUA-72B | 4.38 | 44.99 | 2.30 |
| Open-source | UI-TARS-1.5-7B | 8.46 | 27.52 | 5.00 |
| Open-source | ComputerRL | 21.77 | 48.90 | 14.00 |
| Open-source | Qwen3.5-9B | 8.93 | 41.80 | 5.00 |
| Open-source | Qwen3-VL-8B | 16.27 | 33.90 | 10.70 |
| SCOPE | SCOPE-Capability-Safety | 56.83 | 49.72 | **66.30** |
| SCOPE | **SCOPE-RL-9B** | **58.80** | **54.17** | 64.30 |

All values are percentages, and ↑ indicates that higher is better. \(H\) is the harmonic mean of the OSWorld task success rate and the OS-BLIND attack-avoidance rate. Compared with SCOPE-Capability-Safety, SCOPE-RL-9B improves OSWorld performance by 4.45 percentage points, while attack avoidance decreases by 2.00 points. The harmonic mean increases from 56.83% to 58.80%.

## License

This model is subject to the license terms of its base model, Qwen3.5-9B. Licensing information for the code is available in the corresponding GitHub repositories.

## Citation

If you use SCOPE-RL, SCOPE-SFT, SATraj-OS, or SCOPE-Gen, please cite:

```bibtex
@misc{kang2026scope,
  title  = {Beyond Task Completion: Training Capable and Safe Computer-Use Agents},
  author = {Zeyu Kang and Zhenyun Yin and Yang Zhang and Shan He and Shanzhe Lei and Yanjiu Zhong and Xinquan Chen and Xuhong Wang},
  year   = {2026}
}
```