--- license: apache-2.0 language: - en library_name: transformers pipeline_tag: text-generation base_model: JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT base_model_relation: finetune datasets: - JaspinXu/KnowWhenToHandOff-data tags: - agents - tool-use - customer-service - reinforcement-learning - grpo - tau2-bench --- # KnowWhenToHandOff-Qwen3-4B-R2 (GRPO, Responsible-AI reward) > **Research artefact — not for deployment.** Trained and evaluated only with > simulated users on a public benchmark. It has not been tested with real > customers, real data or real business systems. **Code** [github.com/JaspinXu/KnowWhenToHandOff](https://github.com/JaspinXu/KnowWhenToHandOff) · **Data** [KnowWhenToHandOff-data](https://huggingface.co/datasets/JaspinXu/KnowWhenToHandOff-data) · **Family** [SFT](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT) · [R1](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1) · [R2](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2) **R2** is Qwen3-4B-Instruct-2507 trained with SFT and then multi-turn GRPO on the τ²-bench airline and retail domains, with a reward that adds penalties for rule violations, unnecessary hand-offs and missed hand-offs. It has the best hand-off F1 of the study (0.812) with 10 % over-escalation, at the same task success as its SFT starting point. ## Model family | Model | Training | Hand-off F1 | Over-escalation | Use it for | | --- | --- | --- | --- | --- | | [SFT (B2)](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT) | LoRA SFT on 790 filtered teacher dialogues | 0.768 | 0.016 | starting point; precise but misses hand-offs | | [R1](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1) | GRPO from B2, task reward only | 0.802 | 0.462 | studying the outcome-reward loophole | | [R2](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2) | GRPO from B2, task reward + violation and hand-off penalties | **0.812** | 0.103 | the best-calibrated hand-off model | Task success on the official test tasks is the same for all three within noise (0.33–0.35). ## Quick start A standard Qwen3 chat model with tool calling. It was trained with the τ²-bench domain policy as the system prompt and the domain's tools, including `transfer_to_human_agents`; for faithful behaviour run it inside the evaluation harness of the [code repository](https://github.com/JaspinXu/KnowWhenToHandOff). ```python from transformers import AutoModelForCausalLM, AutoTokenizer repo = "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2" tokenizer = AutoTokenizer.from_pretrained(repo) model = AutoModelForCausalLM.from_pretrained( repo, torch_dtype="auto", device_map="auto" ) transfer = { "type": "function", "function": { "name": "transfer_to_human_agents", "description": "Transfer the user to a human agent.", "parameters": { "type": "object", "properties": {"summary": {"type": "string"}}, "required": ["summary"], }, }, } messages = [ {"role": "system", "content": ""}, {"role": "user", "content": "I need the unaccompanied-minor service."}, ] inputs = tokenizer.apply_chat_template( messages, tools=[transfer], add_generation_prompt=True, return_tensors="pt", ).to(model.device) output = model.generate(inputs, max_new_tokens=512) print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True)) ``` Serve it with vLLM (OpenAI-compatible API, tool calls parsed): ```bash vllm serve JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 \ --enable-auto-tool-choice --tool-call-parser hermes ``` ## Evaluation Test split, protocol frozen in [ADR-005](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/docs/adr/ADR-005-evaluation-protocol.md); 95 % bootstrap intervals over tasks. | Metric | This model | B2 (SFT start) | Source | | --- | --- | --- | --- | | Success (pass^1), test_main | 0.342 [0.245, 0.449] | 0.327 [0.224, 0.434] | [main table](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/main_table.md) | | Paired Δ success vs B2 | +0.015 [−0.056, 0.082] | — | [paired differences](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/paired_differences.md) | | pass^4 | 0.122 [0.041, 0.224] | 0.102 [0.041, 0.204] | [main table](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/main_table.md) | | Violation rate (incl. off-reference writes) | 0.444 | 0.469 | same | | Hand-off F1 / precision / recall, test_handoff | **0.812** / 0.886 / 0.750 | 0.768 / 0.976 / 0.633 | [`R2-final-test_handoff-s0-20261007`](https://huggingface.co/datasets/JaspinXu/KnowWhenToHandOff-data/blob/main/runs/R2-final-test_handoff-s0-20261007/metrics.json) | | Over-escalation, test_handoff | 0.103 | 0.016 | same | | Worst-style success / style gap | 0.316 / 0.112 | 0.327 / 0.112 | [styles](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/styles.md) | **Known behaviour.** R2 removes R1's over-escalation but sometimes makes the opposite error on out-of-scope requests: it refuses and tells the user to contact the company directly without calling the transfer tool, or substitutes an unrelated write (under-escalation 0.25). Once it justified a hand-off with a policy that does not exist. ## Training details - Base model: Qwen/Qwen3-4B-Instruct-2507 (revision `cdbee75f17c01a7cc42f958dc650907174af0554`), Apache-2.0. - Training: GRPO from the SFT checkpoint B2 with the composite reward ([`configs/reward/r2.yaml`](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/configs/reward/r2.yaml): task reward − 0.25 × weighted violations − 0.5 × unneeded transfer − 1.0 × missed transfer − 0.1 × format errors); LoRA rank 64, alpha 128, lr 1e-5; 8 tasks × 8 rollouts per step, 60 steps, no KL term. - Training run `R2-s1-20261006` (seed 1); exported step-60 checkpoint `sha256:be6198a0c51773f52ba4f20c9605a5931f23b0206e92a9fe0521a93761348c96` (merged bf16, Hugging Face format). - Evaluation code at git `462b5d17b4e9add637a2ee7802974c720b625a5b`; τ²-bench `fc0055dc4e0a316c3f83133267fbd6faaa770992`; all sources in [reports/main/sources.json](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/sources.json). ## Intended use Studying when a tool-using customer-service agent should hand off to a human, in the τ²-bench airline and retail domains. Out of scope: any production or customer-facing use, other domains, and decisions about real people. ## Training data Simulated conversations on τ²-bench v1.0.1 train-split tasks (a dev split was carved out for selection) and hand-off variants derived from them (explicit requests for a human, out-of-scope requests, hard negatives); the user simulator is Qwen3-30B-A3B-Instruct-2507. No real user data. Test tasks were never used for training or selection. ## Limitations - One user simulator (Qwen3-30B-A3B-Instruct-2507, which also generated the SFT data) and one training seed; results may not transfer to real users or other simulators, and the intervals cover task sampling, not training variance. - Small test split (49 official tasks; retail uses the 74 tasks whose reward needs no LLM judge, so retail numbers are not comparable with leaderboards). - Communication styles describe how users write, not who they are; no claim is made about any demographic group. - Policy compliance is checked by five deterministic rules that cover only part of each domain policy; citation rate is a keyword rule (it counted one invented policy as a citation). - Observed reward-hacking behaviours: [docs/rl-audit.md](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/docs/rl-audit.md). ## Licences The weights are a LoRA fine-tune of Qwen/Qwen3-4B-Instruct-2507 and inherit its Apache-2.0 licence. τ²-bench code and data (pinned commit `fc0055dc4e0a316c3f83133267fbd6faaa770992`) are used under the licence in its `LICENSE` file; released dialogue samples are τ²-bench-derived synthetic conversations with no real user data. ## Citation Model DOI: [10.57967/hf/10823](https://doi.org/10.57967/hf/10823). Cite the model and the project: ```bibtex @misc{knowwhentohandoff2026r2, title = {KnowWhenToHandOff-Qwen3-4B-R2}, author = {{KnowWhenToHandOff contributors}}, year = {2026}, publisher = {Hugging Face}, doi = {10.57967/hf/10823}, url = {https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2} } @misc{knowwhentohandoff2026, title = {Know When to Hand Off: Multi-turn Reinforcement Learning for Trustworthy Customer-Service Agents}, author = {{KnowWhenToHandOff contributors}}, year = {2026}, howpublished = {\url{https://github.com/JaspinXu/KnowWhenToHandOff}} } ```