The ultimate guide to multi-harness RL

Qwen3.5-2B-multiharness-RL

A full fine-tune of Qwen/Qwen3.5-2B for agentic data-analysis tasks, trained with asynchronous GRPO (TRL Async GRPO) using OpenCode, Claude Code, Codex, Mini-SWE-Agent. This release is step 500 on main, scoring 37.0% pass@1 across four evaluation harnesses.

Article · Collection · Evaluation tasks · Training dashboard

Training and evaluation curves

Animated training, pass@1 and tool-use curves through step 500

Training curves use a trailing 50-step mean (at least 10 observations). Evaluation markers show measured checkpoints; hollow markers and dotted segments indicate incomplete coverage. Faint background lines show the full recorded trajectory; colored lines reveal the measured checkpoints. The first frame shows the completed chart before replaying, so previews also contain the full curves. The tool-use panel shows absolute calls per graded evaluation rollout; this Qwen run used correctness-only rewards. The animation stops at this revision's step 500.

Static chart · Plotted data · Interactive article

Training

TRL Async GRPO with binary correctness only, initialized from the base model. Each group uses one task and one harness; harnesses rotate across groups. Rollouts run through OpenEnv × Harbor in E2B sandboxes.

Setting Value
Training task pool 1,000 tasks: 150 easy / 600 medium / 250 hard
Checkpoint 500 optimizer steps
Learning rate 3e-6
Rollouts per GRPO group / maximum staleness 8 / 4 optimizer steps
Optimizer / precision paged AdamW 8-bit / bfloat16
Training harnesses OpenCode, Claude Code, Codex, Mini-SWE-Agent
Sampling temperature / top-p 0.8 / 1.0
Per-call output budget, training / evaluation 16,384 / 4,096 tokens

These are the Qwen ablations from the article. Their reward has no tool-efficiency bonus. The output-budget mismatch and other limitations are discussed in the article; later Harbor checkpoints declined, so this release preserves the selected checkpoint rather than substituting the last one.

Evaluation

250 fixed SmolDataEnvs test tasks (33 easy, 118 medium, 99 hard), each evaluated under four harnesses: 1,000 graded task/harness cells. Pass@1 uses the first graded attempt per cell; infrastructure retries do not turn it into pass@k. All cells below are graded.

Harness Correct / evaluated Pass@1
OpenCode 82 / 250 32.8%
Claude Code 112 / 250 44.8%
Codex 98 / 250 39.2%
Mini-SWE-Agent 78 / 250 31.2%
Overall 370 / 1,000 37.0%

Harness versions: OpenCode 1.18.31, Claude Code 2.1.270, Codex 0.154.0, Mini-SWE-Agent 2.4.6. Full scores, including difficulty breakdowns, are in eval_results.json.

This is the run's best observed checkpoint on the same test set (also the final checkpoint for standalone OpenCode). Selection on test performance can inflate the reported result.

These are single-run results on a specific task set and harness versions. Harness mix, training exposure and compute differ across runs; the scores do not isolate a causal effect of the harness or objective.

Load the checkpoint

The repository contains full saved model weights and the saved tokenizer/chat template, not a LoRA adapter. Training used Transformers 5.14.1; use a compatible Transformers release.

import torch
from transformers import AutoTokenizer, AutoModelForImageTextToText

model_id = "FineEnvs/Qwen3.5-2B-multiharness-RL"
revision = "main"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
model = AutoModelForImageTextToText.from_pretrained(
    model_id, revision=revision, dtype=torch.bfloat16, device_map="auto"
)

Reproducing the task scores requires the agent harness and tools described in the article; a plain chat prompt is not the same evaluation. This Qwen release was fine-tuned and evaluated for text/tool interaction; no image benchmark is claimed.

Related releases

All models, datasets, environments and the article are linked in the multi-harness RL collection.

License and provenance

Derived from Qwen/Qwen3.5-2B under Apache License 2.0. The base model's license is included unchanged in LICENSE. FineEnvs modified the weights by reinforcement learning; this is not an official Qwen release. See NOTICE and release_manifest.json for modification notices, the pinned base revision, checkpoint identity and file checksums. Optimizer, scheduler, RNG and trainer state are excluded.

Citation

For the experiment, methodology and interpretation, cite the main article:

@misc{kolavi2026multiharnessrl,
  author = {Adithya S Kolavi},
  title = {The ultimate guide to multi-harness RL},
  year = {2026},
  url = {https://huggingface.co/spaces/AdithyaSK/multi-harness-rl}
}
Downloads last month
304
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FineEnvs/Qwen3.5-2B-multiharness-RL

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(474)
this model

Dataset used to train FineEnvs/Qwen3.5-2B-multiharness-RL

Collection including FineEnvs/Qwen3.5-2B-multiharness-RL

Evaluation results