File size: 8,798 Bytes
8c47f5c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9964a8d
 
7f8c3dc
 
 
 
 
 
 
 
 
 
 
 
 
9964a8d
 
7f8c3dc
 
 
 
9964a8d
 
 
 
 
 
7f8c3dc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9964a8d
7f8c3dc
 
9964a8d
7f8c3dc
 
 
 
9964a8d
 
 
 
7f8c3dc
9964a8d
 
7f8c3dc
 
9964a8d
8c47f5c
7f8c3dc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8c47f5c
 
 
9964a8d
8c47f5c
9964a8d
8c47f5c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9964a8d
 
 
80ad9c8
 
9964a8d
80ad9c8
 
 
 
 
 
 
 
 
9964a8d
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
---
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
base_model: JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT
base_model_relation: finetune
datasets:
- JaspinXu/KnowWhenToHandOff-data
tags:
- agents
- tool-use
- customer-service
- reinforcement-learning
- grpo
- tau2-bench
---

# KnowWhenToHandOff-Qwen3-4B-R2 (GRPO, Responsible-AI reward)

> **Research artefact — not for deployment.** Trained and evaluated only with
> simulated users on a public benchmark. It has not been tested with real
> customers, real data or real business systems.

**Code** [github.com/JaspinXu/KnowWhenToHandOff](https://github.com/JaspinXu/KnowWhenToHandOff) · **Data** [KnowWhenToHandOff-data](https://huggingface.co/datasets/JaspinXu/KnowWhenToHandOff-data) · **Family** [SFT](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT) · [R1](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1) · [R2](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2)

**R2** is Qwen3-4B-Instruct-2507 trained with SFT and then multi-turn GRPO on the τ²-bench airline and retail domains, with a reward that adds penalties for rule violations, unnecessary hand-offs and missed hand-offs. It has the best hand-off F1 of the study (0.812) with 10 % over-escalation, at the same task success as its SFT starting point.

## Model family

| Model | Training | Hand-off F1 | Over-escalation | Use it for |
| --- | --- | --- | --- | --- |
| [SFT (B2)](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT) | LoRA SFT on 790 filtered teacher dialogues | 0.768 | 0.016 | starting point; precise but misses hand-offs |
| [R1](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1) | GRPO from B2, task reward only | 0.802 | 0.462 | studying the outcome-reward loophole |
| [R2](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2) | GRPO from B2, task reward + violation and hand-off penalties | **0.812** | 0.103 | the best-calibrated hand-off model |

Task success on the official test tasks is the same for all three within
noise (0.33–0.35).

## Quick start

A standard Qwen3 chat model with tool calling. It was trained with the
τ²-bench domain policy as the system prompt and the domain's tools, including
`transfer_to_human_agents`; for faithful behaviour run it inside the evaluation
harness of the [code repository](https://github.com/JaspinXu/KnowWhenToHandOff).

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo, torch_dtype="auto", device_map="auto"
)

transfer = {
    "type": "function",
    "function": {
        "name": "transfer_to_human_agents",
        "description": "Transfer the user to a human agent.",
        "parameters": {
            "type": "object",
            "properties": {"summary": {"type": "string"}},
            "required": ["summary"],
        },
    },
}
messages = [
    {"role": "system", "content": "<airline or retail policy>"},
    {"role": "user", "content": "I need the unaccompanied-minor service."},
]
inputs = tokenizer.apply_chat_template(
    messages, tools=[transfer], add_generation_prompt=True,
    return_tensors="pt",
).to(model.device)
output = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True))
```

Serve it with vLLM (OpenAI-compatible API, tool calls parsed):

```bash
vllm serve JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 \
  --enable-auto-tool-choice --tool-call-parser hermes
```

## Evaluation

Test split, protocol frozen in [ADR-005](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/docs/adr/ADR-005-evaluation-protocol.md); 95 % bootstrap intervals over tasks.

| Metric | This model | B2 (SFT start) | Source |
| --- | --- | --- | --- |
| Success (pass^1), test_main | 0.342 [0.245, 0.449] | 0.327 [0.224, 0.434] | [main table](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/main_table.md) |
| Paired Δ success vs B2 | +0.015 [−0.056, 0.082] | — | [paired differences](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/paired_differences.md) |
| pass^4 | 0.122 [0.041, 0.224] | 0.102 [0.041, 0.204] | [main table](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/main_table.md) |
| Violation rate (incl. off-reference writes) | 0.444 | 0.469 | same |
| Hand-off F1 / precision / recall, test_handoff | **0.812** / 0.886 / 0.750 | 0.768 / 0.976 / 0.633 | [`R2-final-test_handoff-s0-20261007`](https://huggingface.co/datasets/JaspinXu/KnowWhenToHandOff-data/blob/main/runs/R2-final-test_handoff-s0-20261007/metrics.json) |
| Over-escalation, test_handoff | 0.103 | 0.016 | same |
| Worst-style success / style gap | 0.316 / 0.112 | 0.327 / 0.112 | [styles](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/styles.md) |

**Known behaviour.** R2 removes R1's over-escalation but sometimes makes the opposite error on out-of-scope requests: it refuses and tells the user to contact the company directly without calling the transfer tool, or substitutes an unrelated write (under-escalation 0.25). Once it justified a hand-off with a policy that does not exist.

## Training details

- Base model: Qwen/Qwen3-4B-Instruct-2507 (revision
  `cdbee75f17c01a7cc42f958dc650907174af0554`), Apache-2.0.
- Training: GRPO from the SFT checkpoint B2 with the composite reward ([`configs/reward/r2.yaml`](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/configs/reward/r2.yaml): task reward − 0.25 × weighted violations − 0.5 × unneeded transfer − 1.0 × missed transfer − 0.1 × format errors); LoRA rank 64, alpha 128, lr 1e-5; 8 tasks × 8 rollouts per step, 60 steps, no KL term.
- Training run `R2-s1-20261006` (seed 1); exported step-60 checkpoint
  `sha256:be6198a0c51773f52ba4f20c9605a5931f23b0206e92a9fe0521a93761348c96` (merged bf16, Hugging Face format).
- Evaluation code at git `462b5d17b4e9add637a2ee7802974c720b625a5b`; τ²-bench
  `fc0055dc4e0a316c3f83133267fbd6faaa770992`; all sources in
  [reports/main/sources.json](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/sources.json).

## Intended use

Studying when a tool-using customer-service agent should hand off to a human,
in the τ²-bench airline and retail domains. Out of scope: any production or
customer-facing use, other domains, and decisions about real people.

## Training data

Simulated conversations on τ²-bench v1.0.1 train-split tasks (a dev split was
carved out for selection) and hand-off variants derived from them (explicit
requests for a human, out-of-scope requests, hard negatives); the user
simulator is Qwen3-30B-A3B-Instruct-2507. No real user data. Test tasks were
never used for training or selection.

## Limitations

- One user simulator (Qwen3-30B-A3B-Instruct-2507, which also generated the
  SFT data) and one training seed; results may not transfer to real users or
  other simulators, and the intervals cover task sampling, not training variance.
- Small test split (49 official tasks; retail uses the 74 tasks whose reward
  needs no LLM judge, so retail numbers are not comparable with leaderboards).
- Communication styles describe how users write, not who they are; no claim is
  made about any demographic group.
- Policy compliance is checked by five deterministic rules that cover only
  part of each domain policy; citation rate is a keyword rule (it counted one
  invented policy as a citation).
- Observed reward-hacking behaviours:
  [docs/rl-audit.md](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/docs/rl-audit.md).

## Licences

The weights are a LoRA fine-tune of Qwen/Qwen3-4B-Instruct-2507 and inherit
its Apache-2.0 licence. τ²-bench code and data (pinned commit
`fc0055dc4e0a316c3f83133267fbd6faaa770992`) are used under the licence in
its `LICENSE` file;
released dialogue samples are τ²-bench-derived synthetic conversations with
no real user data.

## Citation

Model DOI: [10.57967/hf/10823](https://doi.org/10.57967/hf/10823). Cite the model and the project:

```bibtex
@misc{knowwhentohandoff2026r2,
  title     = {KnowWhenToHandOff-Qwen3-4B-R2},
  author    = {{KnowWhenToHandOff contributors}},
  year      = {2026},
  publisher = {Hugging Face},
  doi       = {10.57967/hf/10823},
  url       = {https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2}
}

@misc{knowwhentohandoff2026,
  title  = {Know When to Hand Off: Multi-turn Reinforcement Learning for
            Trustworthy Customer-Service Agents},
  author = {{KnowWhenToHandOff contributors}},
  year   = {2026},
  howpublished = {\url{https://github.com/JaspinXu/KnowWhenToHandOff}}
}
```