Image-Text-to-Text
Transformers
Safetensors
English
captcha
gui-agent
computer-use
agent
reinforcement-learning
grpo

Access to CaptchaAgent

CaptchaAgent is released for non-commercial academic research only (CC-BY-NC-4.0); commercial use is prohibited. Access requests are reviewed manually by the authors.

By requesting access you agree to use the model solely for non-commercial academic research; any commercial use is prohibited.

Log in or Sign Up to review the conditions and access this model content.

CaptchaAgent

A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs

Zhenhao Zhang1,*, Zhaoyu Fan2, Haohan Ying3, Jingwen Hu3, Hancen Fan1,
Junhao Zhou4, Zitian Chen1, Linchao Zhu2,†
1Columbia University  ·  2Zhejiang University
3University of Rochester  ·  4University of Illinois at Urbana-Champaign
*Project lead  ·  †Corresponding author

arXiv Code
CaptchaArena Trajectories License: CC BY-NC 4.0

The 20 CAPTCHA types

Overview

CaptchaAgent is a single 9B computer-use policy for all 20 CAPTCHA types in CaptchaArena, fine-tuned from Qwen3.5-9B.

  • Input: screenshots of a fixed 1280×1080 viewport. Output: <think> reasoning plus one tool call per turn (screenshot, click, type_text, drag, hold), at most 15 turns.
  • Training: SFT on replay-verified, reasoning-annotated trajectories, then GRPO with the environment verifier as the reward.

Checkpoints

Subfolder Stage CaptchaArena Pass@1 Pass@5 Open CaptchaWorld Halligan
SFT/ SFT 70.5 86.0 47.2 13.6
RL/ SFT + GRPO 71.7 86.8 51.0 20.0

Each subfolder holds the full bf16 weights with tokenizer, processor and chat template. CaptchaArena scores are on the 4,000-puzzle test split (5 rollouts per puzzle, unbiased Pass@k). For reference: untrained Qwen3.5-9B 11.4, best open-weight GUI agent 35.2, best closed-source model 69.2, human 94.1.

Usage

hf auth login
hf download ZHEN-04/CaptchaAgent --include "RL/*" --local-dir CaptchaAgent   # or "SFT/*"
from transformers import AutoModelForImageTextToText, AutoProcessor

model = AutoModelForImageTextToText.from_pretrained("ZHEN-04/CaptchaAgent", subfolder="RL", dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained("ZHEN-04/CaptchaAgent", subfolder="RL")

To evaluate, serve the checkpoint with an OpenAI-compatible server (SGLang or vLLM) and run the client in the GitHub repo with SFT_EVAL_NATIVE=1, which parses the Qwen3 XML tool calls the model emits.

Training

  • SFT: rank-64 LoRA on the language model (vision encoder frozen), 12,000 puzzles / 37,621 per-turn samples, 3 epochs. The LoRA is merged into the released weights.
  • RL: GRPO from the SFT checkpoint on 2,643 mined puzzles, full language-model weights, KL 0.005 to the SFT policy.

Full configurations are in Appendices K and L of the paper.

Intended use

CaptchaAgent is released for research on computer-use agents and CAPTCHA robustness, not for bypassing protections on services you do not own.

License

Released under CC-BY-NC-4.0 (Creative Commons Attribution–NonCommercial 4.0). Non-commercial academic research use only — commercial use is prohibited. Access is gated: you must request access and agree to these terms before downloading. Please attribute when using this model.

Citation

@article{zhang2026captchaarena,
  title   = {CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs},
  author  = {Zhang, Zhenhao and Fan, Zhaoyu and Ying, Haohan and Hu, Jingwen and Fan, Hancen and Zhou, Junhao and Chen, Zitian and Zhu, Linchao},
  journal = {arXiv preprint arXiv:2609.31957},
  year    = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ZHEN-04/CaptchaAgent

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(968)
this model

Datasets used to train ZHEN-04/CaptchaAgent

Paper for ZHEN-04/CaptchaAgent