Text Generation
Transformers
Safetensors
English
qwen3
agents
tool-use
customer-service
reinforcement-learning
grpo
tau2-bench
conversational
text-generation-inference
Instructions to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1") model = AutoModelForCausalLM.from_pretrained("JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1
- SGLang
How to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1 with Docker Model Runner:
docker model run hf.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1
|
Download README.md from JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1: direct link, hf CLI and curl.
- Browser
- Download file 8.74 kB
-
https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1/resolve/main/README.md
- Command line
-
hf download hf://JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1/README.md
-
curl -L -o README.md https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1/resolve/main/README.md
8.74 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| library_name: transformers | |
| pipeline_tag: text-generation | |
| base_model: JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT | |
| base_model_relation: finetune | |
| datasets: | |
| - JaspinXu/KnowWhenToHandOff-data | |
| tags: | |
| - agents | |
| - tool-use | |
| - customer-service | |
| - reinforcement-learning | |
| - grpo | |
| - tau2-bench | |
| # KnowWhenToHandOff-Qwen3-4B-R1 (GRPO, task reward only) | |
| > **Research artefact — not for deployment.** Trained and evaluated only with | |
| > simulated users on a public benchmark. It has not been tested with real | |
| > customers, real data or real business systems. | |
| **Code** [github.com/JaspinXu/KnowWhenToHandOff](https://github.com/JaspinXu/KnowWhenToHandOff) · **Data** [KnowWhenToHandOff-data](https://huggingface.co/datasets/JaspinXu/KnowWhenToHandOff-data) · **Family** [SFT](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT) · [R1](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1) · [R2](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2) | |
| **R1** is the same recipe as R2 but trained on τ²-bench's task reward only. It reaches a similar hand-off F1 (0.802) by transferring almost half of the users who did not need a human (46 % over-escalation): an outcome-reward loophole, released as a case study rather than as a model to use. | |
| ## Model family | |
| | Model | Training | Hand-off F1 | Over-escalation | Use it for | | |
| | --- | --- | --- | --- | --- | | |
| | [SFT (B2)](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-SFT) | LoRA SFT on 790 filtered teacher dialogues | 0.768 | 0.016 | starting point; precise but misses hand-offs | | |
| | [R1](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1) | GRPO from B2, task reward only | 0.802 | 0.462 | studying the outcome-reward loophole | | |
| | [R2](https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2) | GRPO from B2, task reward + violation and hand-off penalties | **0.812** | 0.103 | the best-calibrated hand-off model | | |
| Task success on the official test tasks is the same for all three within | |
| noise (0.33–0.35). | |
| ## Quick start | |
| A standard Qwen3 chat model with tool calling. It was trained with the | |
| τ²-bench domain policy as the system prompt and the domain's tools, including | |
| `transfer_to_human_agents`; for faithful behaviour run it inside the evaluation | |
| harness of the [code repository](https://github.com/JaspinXu/KnowWhenToHandOff). | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| repo = "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1" | |
| tokenizer = AutoTokenizer.from_pretrained(repo) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| repo, torch_dtype="auto", device_map="auto" | |
| ) | |
| transfer = { | |
| "type": "function", | |
| "function": { | |
| "name": "transfer_to_human_agents", | |
| "description": "Transfer the user to a human agent.", | |
| "parameters": { | |
| "type": "object", | |
| "properties": {"summary": {"type": "string"}}, | |
| "required": ["summary"], | |
| }, | |
| }, | |
| } | |
| messages = [ | |
| {"role": "system", "content": "<airline or retail policy>"}, | |
| {"role": "user", "content": "I need the unaccompanied-minor service."}, | |
| ] | |
| inputs = tokenizer.apply_chat_template( | |
| messages, tools=[transfer], add_generation_prompt=True, | |
| return_tensors="pt", | |
| ).to(model.device) | |
| output = model.generate(inputs, max_new_tokens=512) | |
| print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True)) | |
| ``` | |
| Serve it with vLLM (OpenAI-compatible API, tool calls parsed): | |
| ```bash | |
| vllm serve JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1 \ | |
| --enable-auto-tool-choice --tool-call-parser hermes | |
| ``` | |
| ## Evaluation | |
| Test split, protocol frozen in [ADR-005](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/docs/adr/ADR-005-evaluation-protocol.md); 95 % bootstrap intervals over tasks. | |
| | Metric | This model | B2 (SFT start) | Source | | |
| | --- | --- | --- | --- | | |
| | Success (pass^1), test_main | 0.347 [0.235, 0.459] | 0.327 [0.224, 0.434] | [main table](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/main_table.md) | | |
| | Paired Δ success vs B2 | +0.020 [−0.066, 0.107] | — | [paired differences](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/paired_differences.md) | | |
| | pass^4 | 0.224 [0.102, 0.347] | 0.102 [0.041, 0.204] | [main table](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/main_table.md) | | |
| | Violation rate (incl. off-reference writes) | 0.327 | 0.469 | same | | |
| | Hand-off F1 / precision / recall, test_handoff | 0.802 / 0.689 / 0.959 | 0.768 / 0.976 / 0.633 | [`R1-final-test_handoff-s0-20261006`](https://huggingface.co/datasets/JaspinXu/KnowWhenToHandOff-data/blob/main/runs/R1-final-test_handoff-s0-20261006/metrics.json) | | |
| | Over-escalation, test_handoff | **0.462** | 0.016 | same | | |
| | Worst-style success / style gap | 0.286 / 0.102 | 0.327 / 0.112 | [styles](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/styles.md) | | |
| **Known behaviour.** This checkpoint over-escalates: it transfers almost half of the users who did not need a human, usually right after the account lookup and without explaining why (citation rate 0.06). Under the task-only reward a transfer scores 1.0 whenever the reference solution makes no database change, so "transfer when unsure" was learned in the last 20 steps (step 40 had over-escalation 0.141). Use R2 for hand-off studies. | |
| ## Training details | |
| - Base model: Qwen/Qwen3-4B-Instruct-2507 (revision | |
| `cdbee75f17c01a7cc42f958dc650907174af0554`), Apache-2.0. | |
| - Training: GRPO from the SFT checkpoint B2 with the τ²-bench task reward only ([`configs/reward/task_only.yaml`](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/configs/reward/task_only.yaml)); LoRA rank 64, alpha 128, lr 1e-5; 8 tasks × 8 rollouts per step, 60 steps, no KL term. | |
| - Training run `R1-s1-20261006` (seed 1); exported step-60 checkpoint | |
| `sha256:900d757d4cd74a717c34b3083e2544ce595336add21d0cc9ac7948e57da2aa59` (merged bf16, Hugging Face format). | |
| - Evaluation code at git `c6d50dfad3439b9ef92976cca7aedfcfed8b77f7`; τ²-bench | |
| `fc0055dc4e0a316c3f83133267fbd6faaa770992`; all sources in | |
| [reports/main/sources.json](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/reports/main/sources.json). | |
| ## Intended use | |
| Studying when a tool-using customer-service agent should hand off to a human, | |
| in the τ²-bench airline and retail domains. Out of scope: any production or | |
| customer-facing use, other domains, and decisions about real people. | |
| ## Training data | |
| Simulated conversations on τ²-bench v1.0.1 train-split tasks (a dev split was | |
| carved out for selection) and hand-off variants derived from them (explicit | |
| requests for a human, out-of-scope requests, hard negatives); the user | |
| simulator is Qwen3-30B-A3B-Instruct-2507. No real user data. Test tasks were | |
| never used for training or selection. | |
| ## Limitations | |
| - One user simulator (Qwen3-30B-A3B-Instruct-2507, which also generated the | |
| SFT data) and one training seed; results may not transfer to real users or | |
| other simulators, and the intervals cover task sampling, not training variance. | |
| - Small test split (49 official tasks; retail uses the 74 tasks whose reward | |
| needs no LLM judge, so retail numbers are not comparable with leaderboards). | |
| - Communication styles describe how users write, not who they are; no claim is | |
| made about any demographic group. | |
| - Policy compliance is checked by five deterministic rules that cover only | |
| part of each domain policy; citation rate is a keyword rule (it counted one | |
| invented policy as a citation). | |
| - Observed reward-hacking behaviours: | |
| [docs/rl-audit.md](https://github.com/JaspinXu/KnowWhenToHandOff/blob/main/docs/rl-audit.md). | |
| ## Licences | |
| The weights are a LoRA fine-tune of Qwen/Qwen3-4B-Instruct-2507 and inherit | |
| its Apache-2.0 licence. τ²-bench code and data (pinned commit | |
| `fc0055dc4e0a316c3f83133267fbd6faaa770992`) are used under the licence in | |
| its `LICENSE` file; | |
| released dialogue samples are τ²-bench-derived synthetic conversations with | |
| no real user data. | |
| ## Citation | |
| Model DOI: [10.57967/hf/10824](https://doi.org/10.57967/hf/10824). Cite the model and the project: | |
| ```bibtex | |
| @misc{knowwhentohandoff2026r1, | |
| title = {KnowWhenToHandOff-Qwen3-4B-R1}, | |
| author = {{KnowWhenToHandOff contributors}}, | |
| year = {2026}, | |
| publisher = {Hugging Face}, | |
| doi = {10.57967/hf/10824}, | |
| url = {https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R1} | |
| } | |
| @misc{knowwhentohandoff2026, | |
| title = {Know When to Hand Off: Multi-turn Reinforcement Learning for | |
| Trustworthy Customer-Service Agents}, | |
| author = {{KnowWhenToHandOff contributors}}, | |
| year = {2026}, | |
| howpublished = {\url{https://github.com/JaspinXu/KnowWhenToHandOff}} | |
| } | |
| ``` | |