How to use from
Docker Model Runner
docker model run hf.co/XLearning-SCU/HERO-4B
Quick Links

HERO logo

HERO-4B

Doing More with Less Tokens: Hierarchical Reinforcement Learning for Efficient Coding Agents

Paper Code Hugging Face Models Hugging Face Trajectories

🔍 Overview

HERO-4B is post-trained from Qwen3.5-4B using HERO (HiErarchical ReinfOrcement learning) to improve token efficiency while prioritizing task resolution.

Item Description
Base model Qwen/Qwen3.5-4B
Parameters 4B (dense)
Architecture and tokenizer Inherited from Qwen3.5-4B
Training HERO on 640 multilingual SWE tasks from SWE-Gym, Multi-SWE-bench, and SWE-rebench
Intended use Repository-level coding agents
Format Hugging Face weights and tokenizer

HERO combines capability-based efficiency gating, resolution-first clipping, and efficiency credit at trajectory and turn levels. Qwen3.5 is the backbone; this release contains the HERO post-trained weights.

Usage

Follow the preparation guide in the HERO code repository. From that repository's root, with this model saved under ../models/HERO-4B/:

MODEL=../models/HERO-4B bash eval/run_eval_swebench_verified.sh
MODEL=../models/HERO-4B bash eval/run_eval_swebench_multilingual.sh

Evaluation Settings

The released evaluation scripts use the following defaults for both SWE-bench Verified and SWE-bench Multilingual:

Setting Value
Agent scaffold Claude Code
Inference backend SGLang
Temperature 0.6
Top-p 0.95
Context length 131,072 tokens
Maximum output per call 16,000 tokens
Maximum agent turns 200
Automatic context compaction Disabled
Tools Default tools, excluding WebFetch, WebSearch, and Agent

See the evaluation script for configuration options.

License

Apache-2.0. See LICENSE. We acknowledge the Qwen team for the base model.

📖 Citation

If you find HERO useful, please cite our paper:

@misc{li2026hero,
  title         = {Doing More with Less Tokens: Hierarchical Reinforcement Learning for Efficient Coding Agents},
  author        = {Haobin Li and Liang Jiang and Zhenyu Huang and Mouxing Yang and Xi Peng},
  year          = {2026},
  eprint        = {2609.38885},
  archivePrefix = {arXiv},
  primaryClass  = {cs.SE},
  url           = {https://arxiv.org/abs/2609.38885}
}
Downloads last month
14
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for XLearning-SCU/HERO-4B

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(917)
this model
Quantizations
2 models

Collection including XLearning-SCU/HERO-4B

Paper for XLearning-SCU/HERO-4B