HERO-9B / README.md
XLearning-SCU's picture
Upload HERO-9B release
1ba7885 verified
|
Raw History Blame Contribute Delete
3.52 kB
---
library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE
pipeline_tag: image-text-to-text
base_model: Qwen/Qwen3.5-9B
base_model_relation: finetune
tags:
- HERO
- reinforcement-learning
- coding-agent
- token-efficiency
---
<p align="center">
<img src="assets/hero-logo.png" alt="HERO logo" width="760">
</p>
<h1 align="center">HERO-9B</h1>
<p align="center">
<strong>Doing More with Less Tokens: Hierarchical Reinforcement Learning for Efficient Coding Agents</strong>
</p>
<p align="center">
<a href="https://arxiv.org/abs/2609.38885"><img src="https://img.shields.io/badge/arXiv-2609.38885-b31b1b.svg" alt="Paper"></a>
<a href="https://github.com/XLearning-SCU/HERO"><img src="https://img.shields.io/badge/GitHub-Code-181717.svg" alt="Code"></a>
<a href="https://huggingface.co/collections/XLearning-SCU/hero"><img src="https://img.shields.io/badge/%F0%9F%A4%97-Models-f4c430.svg" alt="Hugging Face Models"></a>
<a href="https://huggingface.co/datasets/XLearning-SCU/HERO-Trajectories"><img src="https://img.shields.io/badge/%F0%9F%A4%97-Trajectories-f4c430.svg" alt="Hugging Face Trajectories"></a>
</p>
## ๐Ÿ” Overview
**HERO-9B** is post-trained from **Qwen3.5-9B** using HERO (HiErarchical ReinfOrcement learning) to improve token efficiency while prioritizing task resolution.
| Item | Description |
| --- | --- |
| Base model | `Qwen/Qwen3.5-9B` |
| Parameters | 9B (dense) |
| Architecture and tokenizer | Inherited from Qwen3.5-9B |
| Training | HERO on 640 multilingual SWE tasks from SWE-Gym, Multi-SWE-bench, and SWE-rebench |
| Intended use | Repository-level coding agents |
| Format | Hugging Face weights and tokenizer |
HERO combines capability-based efficiency gating, resolution-first clipping,
and efficiency credit at trajectory and turn levels. Qwen3.5 is the backbone;
this release contains the HERO post-trained weights.
## Usage
Follow the [preparation guide](https://github.com/XLearning-SCU/HERO/blob/main/Prepare.md) in the HERO code repository. From that repository's root,
with this model saved under `../models/HERO-9B/`:
```bash
MODEL=../models/HERO-9B bash eval/run_eval_swebench_verified.sh
MODEL=../models/HERO-9B bash eval/run_eval_swebench_multilingual.sh
```
## Evaluation Settings
The released evaluation scripts use the following defaults for both SWE-bench Verified and SWE-bench Multilingual:
| Setting | Value |
| --- | --- |
| Agent scaffold | Claude Code |
| Inference backend | SGLang |
| Temperature | 0.6 |
| Top-p | 0.95 |
| Context length | 131,072 tokens |
| Maximum output per call | 16,000 tokens |
| Maximum agent turns | 200 |
| Automatic context compaction | Disabled |
| Tools | Default tools, excluding WebFetch, WebSearch, and Agent |
See the [evaluation script](https://github.com/XLearning-SCU/HERO/blob/main/eval/run_eval_swe_bench.sh) for configuration options.
## License
Apache-2.0. See `LICENSE`. We acknowledge the Qwen team for the base model.
## ๐Ÿ“– Citation
If you find HERO useful, please cite our [paper](https://arxiv.org/abs/2609.38885):
```bibtex
@misc{li2026hero,
title = {Doing More with Less Tokens: Hierarchical Reinforcement Learning for Efficient Coding Agents},
author = {Haobin Li and Liang Jiang and Zhenyu Huang and Mouxing Yang and Xi Peng},
year = {2026},
eprint = {2609.38885},
archivePrefix = {arXiv},
primaryClass = {cs.SE},
url = {https://arxiv.org/abs/2609.38885}
}
```