Text Generation
Transformers
Safetensors
English
qwen3
agents
tool-use
customer-service
reinforcement-learning
grpo
tau2-bench
conversational
text-generation-inference
Instructions to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2") model = AutoModelForCausalLM.from_pretrained("JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2
- SGLang
How to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 with Docker Model Runner:
docker model run hf.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2
File size: 8,798 Bytes
8c47f5c 9964a8d 7f8c3dc 9964a8d 7f8c3dc 9964a8d 7f8c3dc 9964a8d 7f8c3dc 9964a8d 7f8c3dc 9964a8d 7f8c3dc 9964a8d 7f8c3dc 9964a8d 8c47f5c 7f8c3dc 8c47f5c 9964a8d 8c47f5c 9964a8d 8c47f5c 9964a8d 80ad9c8 9964a8d 80ad9c8 9964a8d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 | ---
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
base_model: JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT
base_model_relation: finetune
datasets:
- JaspinXu/KnowWhenToHandOff-data
tags:
- agents
- tool-use
- customer-service
- reinforcement-learning
- grpo
- tau2-bench
---
# KnowWhenToHandOff-Qwen3-4B-R2 (GRPO, Responsible-AI reward)
> **Research artefact — not for deployment.** Trained and evaluated only with
> simulated users on a public benchmark. It has not been tested with real
> customers, real data or real business systems.
**Code** [github.com/JaspinXu/KnowWhenToHandOff](https://github.com/JaspinXu/KnowWhenToHandOff) · **Data** [KnowWhenToHandOff-data](https://huggingface.co/datasets/JaspinXu/KnowWhenToHandOff-data) · **Family** [SFT](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT) · [R1](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1) · [R2](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2)
**R2** is Qwen3-4B-Instruct-2507 trained with SFT and then multi-turn GRPO on the τ²-bench airline and retail domains, with a reward that adds penalties for rule violations, unnecessary hand-offs and missed hand-offs. It has the best hand-off F1 of the study (0.812) with 10 % over-escalation, at the same task success as its SFT starting point.
## Model family
| Model | Training | Hand-off F1 | Over-escalation | Use it for |
| --- | --- | --- | --- | --- |
| [SFT (B2)](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT) | LoRA SFT on 790 filtered teacher dialogues | 0.768 | 0.016 | starting point; precise but misses hand-offs |
| [R1](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1) | GRPO from B2, task reward only | 0.802 | 0.462 | studying the outcome-reward loophole |
| [R2](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2) | GRPO from B2, task reward + violation and hand-off penalties | **0.812** | 0.103 | the best-calibrated hand-off model |
Task success on the official test tasks is the same for all three within
noise (0.33–0.35).
## Quick start
A standard Qwen3 chat model with tool calling. It was trained with the
τ²-bench domain policy as the system prompt and the domain's tools, including
`transfer_to_human_agents`; for faithful behaviour run it inside the evaluation
harness of the [code repository](https://github.com/JaspinXu/KnowWhenToHandOff).
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
repo, torch_dtype="auto", device_map="auto"
)
transfer = {
"type": "function",
"function": {
"name": "transfer_to_human_agents",
"description": "Transfer the user to a human agent.",
"parameters": {
"type": "object",
"properties": {"summary": {"type": "string"}},
"required": ["summary"],
},
},
}
messages = [
{"role": "system", "content": "<airline or retail policy>"},
{"role": "user", "content": "I need the unaccompanied-minor service."},
]
inputs = tokenizer.apply_chat_template(
messages, tools=[transfer], add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
output = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True))
```
Serve it with vLLM (OpenAI-compatible API, tool calls parsed):
```bash
vllm serve JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 \
--enable-auto-tool-choice --tool-call-parser hermes
```
## Evaluation
Test split, protocol frozen in [ADR-005](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/docs/adr/ADR-005-evaluation-protocol.md); 95 % bootstrap intervals over tasks.
| Metric | This model | B2 (SFT start) | Source |
| --- | --- | --- | --- |
| Success (pass^1), test_main | 0.342 [0.245, 0.449] | 0.327 [0.224, 0.434] | [main table](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/main_table.md) |
| Paired Δ success vs B2 | +0.015 [−0.056, 0.082] | — | [paired differences](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/paired_differences.md) |
| pass^4 | 0.122 [0.041, 0.224] | 0.102 [0.041, 0.204] | [main table](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/main_table.md) |
| Violation rate (incl. off-reference writes) | 0.444 | 0.469 | same |
| Hand-off F1 / precision / recall, test_handoff | **0.812** / 0.886 / 0.750 | 0.768 / 0.976 / 0.633 | [`R2-final-test_handoff-s0-20261007`](https://huggingface.co/datasets/JaspinXu/KnowWhenToHandOff-data/blob/main/runs/R2-final-test_handoff-s0-20261007/metrics.json) |
| Over-escalation, test_handoff | 0.103 | 0.016 | same |
| Worst-style success / style gap | 0.316 / 0.112 | 0.327 / 0.112 | [styles](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/styles.md) |
**Known behaviour.** R2 removes R1's over-escalation but sometimes makes the opposite error on out-of-scope requests: it refuses and tells the user to contact the company directly without calling the transfer tool, or substitutes an unrelated write (under-escalation 0.25). Once it justified a hand-off with a policy that does not exist.
## Training details
- Base model: Qwen/Qwen3-4B-Instruct-2507 (revision
`cdbee75f17c01a7cc42f958dc650907174af0554`), Apache-2.0.
- Training: GRPO from the SFT checkpoint B2 with the composite reward ([`configs/reward/r2.yaml`](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/configs/reward/r2.yaml): task reward − 0.25 × weighted violations − 0.5 × unneeded transfer − 1.0 × missed transfer − 0.1 × format errors); LoRA rank 64, alpha 128, lr 1e-5; 8 tasks × 8 rollouts per step, 60 steps, no KL term.
- Training run `R2-s1-20261006` (seed 1); exported step-60 checkpoint
`sha256:be6198a0c51773f52ba4f20c9605a5931f23b0206e92a9fe0521a93761348c96` (merged bf16, Hugging Face format).
- Evaluation code at git `462b5d17b4e9add637a2ee7802974c720b625a5b`; τ²-bench
`fc0055dc4e0a316c3f83133267fbd6faaa770992`; all sources in
[reports/main/sources.json](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/sources.json).
## Intended use
Studying when a tool-using customer-service agent should hand off to a human,
in the τ²-bench airline and retail domains. Out of scope: any production or
customer-facing use, other domains, and decisions about real people.
## Training data
Simulated conversations on τ²-bench v1.0.1 train-split tasks (a dev split was
carved out for selection) and hand-off variants derived from them (explicit
requests for a human, out-of-scope requests, hard negatives); the user
simulator is Qwen3-30B-A3B-Instruct-2507. No real user data. Test tasks were
never used for training or selection.
## Limitations
- One user simulator (Qwen3-30B-A3B-Instruct-2507, which also generated the
SFT data) and one training seed; results may not transfer to real users or
other simulators, and the intervals cover task sampling, not training variance.
- Small test split (49 official tasks; retail uses the 74 tasks whose reward
needs no LLM judge, so retail numbers are not comparable with leaderboards).
- Communication styles describe how users write, not who they are; no claim is
made about any demographic group.
- Policy compliance is checked by five deterministic rules that cover only
part of each domain policy; citation rate is a keyword rule (it counted one
invented policy as a citation).
- Observed reward-hacking behaviours:
[docs/rl-audit.md](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/docs/rl-audit.md).
## Licences
The weights are a LoRA fine-tune of Qwen/Qwen3-4B-Instruct-2507 and inherit
its Apache-2.0 licence. τ²-bench code and data (pinned commit
`fc0055dc4e0a316c3f83133267fbd6faaa770992`) are used under the licence in
its `LICENSE` file;
released dialogue samples are τ²-bench-derived synthetic conversations with
no real user data.
## Citation
Model DOI: [10.57967/hf/10823](https://doi.org/10.57967/hf/10823). Cite the model and the project:
```bibtex
@misc{knowwhentohandoff2026r2,
title = {KnowWhenToHandOff-Qwen3-4B-R2},
author = {{KnowWhenToHandOff contributors}},
year = {2026},
publisher = {Hugging Face},
doi = {10.57967/hf/10823},
url = {https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2}
}
@misc{knowwhentohandoff2026,
title = {Know When to Hand Off: Multi-turn Reinforcement Learning for
Trustworthy Customer-Service Agents},
author = {{KnowWhenToHandOff contributors}},
year = {2026},
howpublished = {\url{https://github.com/JaspinXu/KnowWhenToHandOff}}
}
```
|