JaspinXu's picture
Card: add DOI and citation
8125db6 verified
|
Raw History Blame Contribute Delete
8.74 kB
---
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
base_model: JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT
base_model_relation: finetune
datasets:
- JaspinXu/KnowWhenToHandOff-data
tags:
- agents
- tool-use
- customer-service
- reinforcement-learning
- grpo
- tau2-bench
---
# KnowWhenToHandOff-Qwen3-4B-R1 (GRPO, task reward only)
> **Research artefact — not for deployment.** Trained and evaluated only with
> simulated users on a public benchmark. It has not been tested with real
> customers, real data or real business systems.
**Code** [github.com/JaspinXu/KnowWhenToHandOff](https://github.com/JaspinXu/KnowWhenToHandOff) · **Data** [KnowWhenToHandOff-data](https://huggingface.co/datasets/JaspinXu/KnowWhenToHandOff-data) · **Family** [SFT](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT) · [R1](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1) · [R2](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2)
**R1** is the same recipe as R2 but trained on τ²-bench's task reward only. It reaches a similar hand-off F1 (0.802) by transferring almost half of the users who did not need a human (46 % over-escalation): an outcome-reward loophole, released as a case study rather than as a model to use.
## Model family
| Model | Training | Hand-off F1 | Over-escalation | Use it for |
| --- | --- | --- | --- | --- |
| [SFT (B2)](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT) | LoRA SFT on 790 filtered teacher dialogues | 0.768 | 0.016 | starting point; precise but misses hand-offs |
| [R1](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1) | GRPO from B2, task reward only | 0.802 | 0.462 | studying the outcome-reward loophole |
| [R2](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2) | GRPO from B2, task reward + violation and hand-off penalties | **0.812** | 0.103 | the best-calibrated hand-off model |
Task success on the official test tasks is the same for all three within
noise (0.33–0.35).
## Quick start
A standard Qwen3 chat model with tool calling. It was trained with the
τ²-bench domain policy as the system prompt and the domain's tools, including
`transfer_to_human_agents`; for faithful behaviour run it inside the evaluation
harness of the [code repository](https://github.com/JaspinXu/KnowWhenToHandOff).
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
repo, torch_dtype="auto", device_map="auto"
)
transfer = {
"type": "function",
"function": {
"name": "transfer_to_human_agents",
"description": "Transfer the user to a human agent.",
"parameters": {
"type": "object",
"properties": {"summary": {"type": "string"}},
"required": ["summary"],
},
},
}
messages = [
{"role": "system", "content": "<airline or retail policy>"},
{"role": "user", "content": "I need the unaccompanied-minor service."},
]
inputs = tokenizer.apply_chat_template(
messages, tools=[transfer], add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
output = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True))
```
Serve it with vLLM (OpenAI-compatible API, tool calls parsed):
```bash
vllm serve JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1 \
--enable-auto-tool-choice --tool-call-parser hermes
```
## Evaluation
Test split, protocol frozen in [ADR-005](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/docs/adr/ADR-005-evaluation-protocol.md); 95 % bootstrap intervals over tasks.
| Metric | This model | B2 (SFT start) | Source |
| --- | --- | --- | --- |
| Success (pass^1), test_main | 0.347 [0.235, 0.459] | 0.327 [0.224, 0.434] | [main table](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/main_table.md) |
| Paired Δ success vs B2 | +0.020 [−0.066, 0.107] | — | [paired differences](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/paired_differences.md) |
| pass^4 | 0.224 [0.102, 0.347] | 0.102 [0.041, 0.204] | [main table](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/main_table.md) |
| Violation rate (incl. off-reference writes) | 0.327 | 0.469 | same |
| Hand-off F1 / precision / recall, test_handoff | 0.802 / 0.689 / 0.959 | 0.768 / 0.976 / 0.633 | [`R1-final-test_handoff-s0-20261006`](https://huggingface.co/datasets/JaspinXu/KnowWhenToHandOff-data/blob/main/runs/R1-final-test_handoff-s0-20261006/metrics.json) |
| Over-escalation, test_handoff | **0.462** | 0.016 | same |
| Worst-style success / style gap | 0.286 / 0.102 | 0.327 / 0.112 | [styles](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/styles.md) |
**Known behaviour.** This checkpoint over-escalates: it transfers almost half of the users who did not need a human, usually right after the account lookup and without explaining why (citation rate 0.06). Under the task-only reward a transfer scores 1.0 whenever the reference solution makes no database change, so "transfer when unsure" was learned in the last 20 steps (step 40 had over-escalation 0.141). Use R2 for hand-off studies.
## Training details
- Base model: Qwen/Qwen3-4B-Instruct-2507 (revision
`cdbee75f17c01a7cc42f958dc650907174af0554`), Apache-2.0.
- Training: GRPO from the SFT checkpoint B2 with the τ²-bench task reward only ([`configs/reward/task_only.yaml`](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/configs/reward/task_only.yaml)); LoRA rank 64, alpha 128, lr 1e-5; 8 tasks × 8 rollouts per step, 60 steps, no KL term.
- Training run `R1-s1-20261006` (seed 1); exported step-60 checkpoint
`sha256:900d757d4cd74a717c34b3083e2544ce595336add21d0cc9ac7948e57da2aa59` (merged bf16, Hugging Face format).
- Evaluation code at git `c6d50dfad3439b9ef92976cca7aedfcfed8b77f7`; τ²-bench
`fc0055dc4e0a316c3f83133267fbd6faaa770992`; all sources in
[reports/main/sources.json](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/sources.json).
## Intended use
Studying when a tool-using customer-service agent should hand off to a human,
in the τ²-bench airline and retail domains. Out of scope: any production or
customer-facing use, other domains, and decisions about real people.
## Training data
Simulated conversations on τ²-bench v1.0.1 train-split tasks (a dev split was
carved out for selection) and hand-off variants derived from them (explicit
requests for a human, out-of-scope requests, hard negatives); the user
simulator is Qwen3-30B-A3B-Instruct-2507. No real user data. Test tasks were
never used for training or selection.
## Limitations
- One user simulator (Qwen3-30B-A3B-Instruct-2507, which also generated the
SFT data) and one training seed; results may not transfer to real users or
other simulators, and the intervals cover task sampling, not training variance.
- Small test split (49 official tasks; retail uses the 74 tasks whose reward
needs no LLM judge, so retail numbers are not comparable with leaderboards).
- Communication styles describe how users write, not who they are; no claim is
made about any demographic group.
- Policy compliance is checked by five deterministic rules that cover only
part of each domain policy; citation rate is a keyword rule (it counted one
invented policy as a citation).
- Observed reward-hacking behaviours:
[docs/rl-audit.md](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/docs/rl-audit.md).
## Licences
The weights are a LoRA fine-tune of Qwen/Qwen3-4B-Instruct-2507 and inherit
its Apache-2.0 licence. τ²-bench code and data (pinned commit
`fc0055dc4e0a316c3f83133267fbd6faaa770992`) are used under the licence in
its `LICENSE` file;
released dialogue samples are τ²-bench-derived synthetic conversations with
no real user data.
## Citation
Model DOI: [10.57967/hf/10824](https://doi.org/10.57967/hf/10824). Cite the model and the project:
```bibtex
@misc{knowwhentohandoff2026r1,
title = {KnowWhenToHandOff-Qwen3-4B-R1},
author = {{KnowWhenToHandOff contributors}},
year = {2026},
publisher = {Hugging Face},
doi = {10.57967/hf/10824},
url = {https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1}
}
@misc{knowwhentohandoff2026,
title = {Know When to Hand Off: Multi-turn Reinforcement Learning for
Trustworthy Customer-Service Agents},
author = {{KnowWhenToHandOff contributors}},
year = {2026},
howpublished = {\url{https://github.com/JaspinXu/KnowWhenToHandOff}}
}
```