--- library_name: transformers license: apache-2.0 license_link: https://huggingface.co/Qwen/Qwen3.5-4B/blob/main/LICENSE pipeline_tag: image-text-to-text base_model: Qwen/Qwen3.5-4B base_model_relation: finetune tags: - HERO - reinforcement-learning - coding-agent - token-efficiency ---
Doing More with Less Tokens: Hierarchical Reinforcement Learning for Efficient Coding Agents
## 🔍 Overview **HERO-4B** is post-trained from **Qwen3.5-4B** using HERO (HiErarchical ReinfOrcement learning) to improve token efficiency while prioritizing task resolution. | Item | Description | | --- | --- | | Base model | `Qwen/Qwen3.5-4B` | | Parameters | 4B (dense) | | Architecture and tokenizer | Inherited from Qwen3.5-4B | | Training | HERO on 640 multilingual SWE tasks from SWE-Gym, Multi-SWE-bench, and SWE-rebench | | Intended use | Repository-level coding agents | | Format | Hugging Face weights and tokenizer | HERO combines capability-based efficiency gating, resolution-first clipping, and efficiency credit at trajectory and turn levels. Qwen3.5 is the backbone; this release contains the HERO post-trained weights. ## Usage Follow the [preparation guide](https://github.com/XLearning-SCU/HERO/blob/main/Prepare.md) in the HERO code repository. From that repository's root, with this model saved under `../models/HERO-4B/`: ```bash MODEL=../models/HERO-4B bash eval/run_eval_swebench_verified.sh MODEL=../models/HERO-4B bash eval/run_eval_swebench_multilingual.sh ``` ## Evaluation Settings The released evaluation scripts use the following defaults for both SWE-bench Verified and SWE-bench Multilingual: | Setting | Value | | --- | --- | | Agent scaffold | Claude Code | | Inference backend | SGLang | | Temperature | 0.6 | | Top-p | 0.95 | | Context length | 131,072 tokens | | Maximum output per call | 16,000 tokens | | Maximum agent turns | 200 | | Automatic context compaction | Disabled | | Tools | Default tools, excluding WebFetch, WebSearch, and Agent | See the [evaluation script](https://github.com/XLearning-SCU/HERO/blob/main/eval/run_eval_swe_bench.sh) for configuration options. ## License Apache-2.0. See `LICENSE`. We acknowledge the Qwen team for the base model. ## 📖 Citation If you find HERO useful, please cite our [paper](https://arxiv.org/abs/2609.38885): ```bibtex @misc{li2026hero, title = {Doing More with Less Tokens: Hierarchical Reinforcement Learning for Efficient Coding Agents}, author = {Haobin Li and Liang Jiang and Zhenyu Huang and Mouxing Yang and Xi Peng}, year = {2026}, eprint = {2609.38885}, archivePrefix = {arXiv}, primaryClass = {cs.SE}, url = {https://arxiv.org/abs/2609.38885} } ```