Instructions to use ZHEN-04/CaptchaAgent with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ZHEN-04/CaptchaAgent with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="ZHEN-04/CaptchaAgent")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ZHEN-04/CaptchaAgent", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ZHEN-04/CaptchaAgent with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ZHEN-04/CaptchaAgent" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZHEN-04/CaptchaAgent", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/ZHEN-04/CaptchaAgent
- SGLang
How to use ZHEN-04/CaptchaAgent with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ZHEN-04/CaptchaAgent" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZHEN-04/CaptchaAgent", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ZHEN-04/CaptchaAgent" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZHEN-04/CaptchaAgent", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use ZHEN-04/CaptchaAgent with Docker Model Runner:
docker model run hf.co/ZHEN-04/CaptchaAgent
Access to CaptchaAgent
CaptchaAgent is released for non-commercial academic research only (CC-BY-NC-4.0); commercial use is prohibited. Access requests are reviewed manually by the authors.
By requesting access you agree to use the model solely for non-commercial academic research; any commercial use is prohibited.
Log in or Sign Up to review the conditions and access this model content.
CaptchaAgent
A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs
Zhenhao Zhang1,*, Zhaoyu Fan2, Haohan Ying3, Jingwen Hu3, Hancen Fan1,
Junhao Zhou4, Zitian Chen1, Linchao Zhu2,†
1Columbia University · 2Zhejiang University
3University of Rochester · 4University of Illinois at Urbana-Champaign
*Project lead · †Corresponding author
Overview
CaptchaAgent is a single 9B computer-use policy for all 20 CAPTCHA types in CaptchaArena, fine-tuned from Qwen3.5-9B.
- Input: screenshots of a fixed 1280×1080 viewport. Output:
<think>reasoning plus one tool call per turn (screenshot,click,type_text,drag,hold), at most 15 turns. - Training: SFT on replay-verified, reasoning-annotated trajectories, then GRPO with the environment verifier as the reward.
Checkpoints
| Subfolder | Stage | CaptchaArena Pass@1 | Pass@5 | Open CaptchaWorld | Halligan |
|---|---|---|---|---|---|
SFT/ |
SFT | 70.5 | 86.0 | 47.2 | 13.6 |
RL/ |
SFT + GRPO | 71.7 | 86.8 | 51.0 | 20.0 |
Each subfolder holds the full bf16 weights with tokenizer, processor and chat template. CaptchaArena scores are on the 4,000-puzzle test split (5 rollouts per puzzle, unbiased Pass@k). For reference: untrained Qwen3.5-9B 11.4, best open-weight GUI agent 35.2, best closed-source model 69.2, human 94.1.
Usage
hf auth login
hf download ZHEN-04/CaptchaAgent --include "RL/*" --local-dir CaptchaAgent # or "SFT/*"
from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained("ZHEN-04/CaptchaAgent", subfolder="RL", dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained("ZHEN-04/CaptchaAgent", subfolder="RL")
To evaluate, serve the checkpoint with an OpenAI-compatible server (SGLang or vLLM) and run the client in the GitHub repo with SFT_EVAL_NATIVE=1, which parses the Qwen3 XML tool calls the model emits.
Training
- SFT: rank-64 LoRA on the language model (vision encoder frozen), 12,000 puzzles / 37,621 per-turn samples, 3 epochs. The LoRA is merged into the released weights.
- RL: GRPO from the SFT checkpoint on 2,643 mined puzzles, full language-model weights, KL 0.005 to the SFT policy.
Full configurations are in Appendices K and L of the paper.
Intended use
CaptchaAgent is released for research on computer-use agents and CAPTCHA robustness, not for bypassing protections on services you do not own.
License
Released under CC-BY-NC-4.0 (Creative Commons Attribution–NonCommercial 4.0). Non-commercial academic research use only — commercial use is prohibited. Access is gated: you must request access and agree to these terms before downloading. Please attribute when using this model.
Citation
@article{zhang2026captchaarena,
title = {CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs},
author = {Zhang, Zhenhao and Fan, Zhaoyu and Ying, Haohan and Hu, Jingwen and Fan, Hancen and Zhou, Junhao and Chen, Zitian and Zhu, Linchao},
journal = {arXiv preprint arXiv:2609.31957},
year = {2026}
}
