KnowWhenToHandOff-Qwen3-4B-SFT (B2: supervised warm start)

Research artefact — not for deployment. Trained and evaluated only with simulated users on a public benchmark. It has not been tested with real customers, real data or real business systems.

Code github.com/JaspinXu/KnowWhenToHandOff · Data KnowWhenToHandOff-data · Family SFT · R1 · R2

SFT (B2) is Qwen3-4B-Instruct-2507 fine-tuned on 790 filtered teacher dialogues; R1 and R2 start from it. Its hand-offs are precise (1.6 % over-escalation) but it misses about a third of the users who need a human (recall 0.633).

Model family

Model Training Hand-off F1 Over-escalation Use it for
SFT (B2) LoRA SFT on 790 filtered teacher dialogues 0.768 0.016 starting point; precise but misses hand-offs
R1 GRPO from B2, task reward only 0.802 0.462 studying the outcome-reward loophole
R2 GRPO from B2, task reward + violation and hand-off penalties 0.812 0.103 the best-calibrated hand-off model

Task success on the official test tasks is the same for all three within noise (0.33–0.35).

Quick start

A standard Qwen3 chat model with tool calling. It was trained with the τ²-bench domain policy as the system prompt and the domain's tools, including transfer_to_human_agents; for faithful behaviour run it inside the evaluation harness of the code repository.

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo, torch_dtype="auto", device_map="auto"
)

transfer = {
    "type": "function",
    "function": {
        "name": "transfer_to_human_agents",
        "description": "Transfer the user to a human agent.",
        "parameters": {
            "type": "object",
            "properties": {"summary": {"type": "string"}},
            "required": ["summary"],
        },
    },
}
messages = [
    {"role": "system", "content": "<airline or retail policy>"},
    {"role": "user", "content": "I need the unaccompanied-minor service."},
]
inputs = tokenizer.apply_chat_template(
    messages, tools=[transfer], add_generation_prompt=True,
    return_tensors="pt",
).to(model.device)
output = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True))

Serve it with vLLM (OpenAI-compatible API, tool calls parsed):

vllm serve JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT \
  --enable-auto-tool-choice --tool-call-parser hermes

Evaluation

Test split, protocol frozen in ADR-005; 95 % bootstrap intervals over tasks.

Metric This model B0 (prompted base) Source
Success (pass^1), test_main 0.327 [0.224, 0.434] 0.306 [0.219, 0.408] main table
Paired Δ success vs B0 +0.020 [−0.056, 0.107] — paired differences (listed there as B0 vs B2, sign flipped)
pass^4 0.102 [0.041, 0.204] 0.102 [0.020, 0.184] main table
Violation rate (incl. off-reference writes) 0.469 0.510 same
Hand-off F1 / precision / recall, test_handoff 0.768 / 0.976 / 0.633 0.654 / 0.866 / 0.526 B2-test_handoff-s0-20261005
Over-escalation, test_handoff 0.016 0.087 same
Worst-style success / style gap 0.327 / 0.112 0.306 / 0.041 styles

Known behaviour. Very precise hand-offs but under-escalates (misses 37 % of users who need a human); sometimes announces a transfer in text without calling the tool, or claims a capability no tool provides (error analysis).

Training details

  • Base model: Qwen/Qwen3-4B-Instruct-2507 (revision cdbee75f17c01a7cc42f958dc650907174af0554), Apache-2.0.
  • Training: supervised fine-tuning (LoRA rank 64, merged) on 790 teacher dialogues from Qwen3-30B-A3B-Instruct-2507 on the train split (all six communication styles), kept only when they succeeded with no rule violation, no format error and the correct hand-off decision; loss on assistant tokens only. It is the starting point of R1 and R2.
  • Training run B2-s0-20261005; checkpoint sha256:36bc27053332e4d01462f0e0b803c9db0f442095f8356b669585090a8e4cb57e.
  • Evaluation code at git 148edfb9f2728f52ba13d291836294a39b1eb05c; τ²-bench fc0055dc4e0a316c3f83133267fbd6faaa770992; all sources in reports/main/sources.json.

Intended use

Studying when a tool-using customer-service agent should hand off to a human, in the τ²-bench airline and retail domains, and as the starting point for RL on that question. Out of scope: any production or customer-facing use, other domains, and decisions about real people.

Training data

Teacher conversations on τ²-bench v1.0.1 train-split tasks and hand-off variants derived from them; the user simulator is Qwen3-30B-A3B-Instruct-2507. No real user data. Test tasks were never used for training or selection.

Limitations

  • One user simulator (Qwen3-30B-A3B-Instruct-2507, which also generated the SFT data) and one training seed; results may not transfer to real users or other simulators, and the intervals cover task sampling, not training variance.
  • Small test split (49 official tasks; retail uses the 74 tasks whose reward needs no LLM judge, so retail numbers are not comparable with leaderboards).
  • Communication styles describe how users write, not who they are; no claim is made about any demographic group.
  • Policy compliance is checked by five deterministic rules that cover only part of each domain policy; citation rate is a keyword rule (it counted one invented policy as a citation).
  • Observed reward-hacking behaviours: docs/rl-audit.md.

Licences

The weights are a LoRA fine-tune of Qwen/Qwen3-4B-Instruct-2507 and inherit its Apache-2.0 licence. τ²-bench code and data (pinned commit fc0055dc4e0a316c3f83133267fbd6faaa770992) are used under the licence in its LICENSE file; released dialogue samples are τ²-bench-derived synthetic conversations with no real user data.

Citation

Model DOI: 10.57967/hf/10825. Cite the model and the project:

@misc{knowwhentohandoff2026sft,
  title     = {KnowWhenToHandOff-Qwen3-4B-SFT},
  author    = {{KnowWhenToHandOff contributors}},
  year      = {2026},
  publisher = {Hugging Face},
  doi       = {10.57967/hf/10825},
  url       = {https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT}
}

@misc{knowwhentohandoff2026,
  title  = {Know When to Hand Off: Multi-turn Reinforcement Learning for
            Trustworthy Customer-Service Agents},
  author = {{KnowWhenToHandOff contributors}},
  year   = {2026},
  howpublished = {\url{https://github.com/JaspinXu/KnowWhenToHandOff}}
}
Downloads last month
22
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT

Finetuned
(2386)
this model
Finetunes
2 models

Dataset used to train JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT

Collection including JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT