benchflow-qwen35-9b / README.md
bingran-you's picture
Update model card for Qwen397B custom SFT adapter
b585e34 verified
|
Raw
History Blame Contribute Delete
4.28 kB
---
library_name: peft
base_model: Qwen/Qwen3.5-9B
pipeline_tag: text-generation
datasets:
- benchflow/env0-experiment-trajectories
tags:
- base_model:adapter:Qwen/Qwen3.5-9B
- lora
- peft
- sft
- env-0
- openhands
- daytona
- qwen397b
---
# BenchFlow Qwen3.5-9B Env-0 Qwen397B-Data Custom SFT LoRA Adapter
This repository contains the current BenchFlow env-0 SFT adapter for `Qwen/Qwen3.5-9B`. It is a PEFT LoRA adapter only; load it with the base `Qwen/Qwen3.5-9B` checkpoint.
## Current Version
| Field | Value |
| --- | --- |
| Adapter repo | `benchflow/benchflow-qwen35-9b` |
| Published model PR | [HF PR #4](https://huggingface.co/benchflow/benchflow-qwen35-9b/discussions/4) |
| Adapter commit promoted from PR | `92380a83764ec2d8b2103a3895e24e49a508d1d9` |
| Training run id | `qwen35-397b-data-qwen35-9b-custom-sft-20260630T042600Z` |
| Base checkpoint | `Qwen/Qwen3.5-9B` |
| Adapter type | LoRA / PEFT |
| Trainer | Custom PyTorch + PEFT LoRA trainer, `experiments/env-0-posttrain-mvp/train_lora_sft.py` |
| Source data | Qwen3.5-397B teacher trajectories collected with BenchFlow, OpenHands, and Daytona |
| Training rows | `298` all-training-ready rows |
| Hardware | 1x H100 80GB |
This run used the historical custom trainer. Prime-RL was not used as the SFT trainer; the source data path includes `prime-rl` only because the trajectories were also validated and exported in Prime-SFT-compatible format.
## Training Recipe
| Field | Value |
| --- | --- |
| Precision | BF16 |
| Quantization | None |
| Context length | `8192` |
| Max trainer steps | `300` micro-batch steps |
| Micro batch size | `1` |
| Gradient accumulation | `8` |
| Approx optimizer updates | `37` |
| Learning rate | `1e-4` |
| Scheduler | None |
| Max grad norm | `1.0` |
| LoRA rank | `32` |
| LoRA alpha | `64` |
| LoRA dropout | `0.05` |
| LoRA targets | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` |
| Checkpoints | `100`, `200`, `300`, `best_adapter`, `final_adapter` |
| Final eval loss | `0.1476329118013382` |
## Evaluation Results
| Evaluation | Runtime | Strict pass |
| --- | --- | ---: |
| Mobile300 | SGLang | `135 / 300` |
| Mobile300 | Fireworks | `134 / 300` |
| standard60, 3 trials | Fireworks | `25 / 180` |
## Artifact Links
| Artifact | Link |
| --- | --- |
| Source teacher trajectories | [HF dataset folder](https://huggingface.co/datasets/benchflow/env0-experiment-trajectories/tree/main/experiments/qwen35-397b-env0-mini-prime-rl-trajectories/20260629-qwen35-397b-openrouter-openhands-daytona/canonical/train-mini-300-combined-qwen35-397b-20260629T2336Z) |
| Training artifacts | [HF dataset folder](https://huggingface.co/datasets/benchflow/env0-experiment-trajectories/tree/main/experiments/qwen35-397b-env0-mini-custom-sft/training/qwen35-397b-data-qwen35-9b-custom-sft-20260630T042600Z) |
| Fireworks Mobile300 eval | [HF dataset folder](https://huggingface.co/datasets/benchflow/env0-experiment-trajectories/tree/main/experiments/qwen35-397b-env0-mini-custom-sft/eval/fireworks-qwen397b-custom-sft-mobile300-20260630T084904Z) |
| Fireworks standard60 eval | [HF dataset folder](https://huggingface.co/datasets/benchflow/env0-experiment-trajectories/tree/main/experiments/qwen35-397b-env0-mini-custom-sft/eval/fireworks-qwen397b-custom-sft-standard60-3trials-20260630T172230Z) |
| Reproduction report | [GitHub report](https://github.com/benchflow-ai/env-0-experiment/blob/main/experiments/qwen35-397b-env0-mini-custom-sft/reports/2026-06-30-qwen397b-custom-sft-reproduction-parameters.md) |
## Loading
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_id = "Qwen/Qwen3.5-9B"
adapter_id = "benchflow/benchflow-qwen35-9b"
tokenizer = AutoTokenizer.from_pretrained(base_id, trust_remote_code=True)
base_model = AutoModelForCausalLM.from_pretrained(
base_id,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True,
)
model = PeftModel.from_pretrained(base_model, adapter_id)
```
## Intended Use And Limitations
This adapter is an env-0 research artifact for controlled BenchFlow/OpenHands/Daytona evaluation. It is not a general-purpose safety-tested assistant model and should not be treated as production-ready for autonomous operation.