Instructions to use Yunhao-Feng/AdaGuard-0.6B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Yunhao-Feng/AdaGuard-0.6B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Yunhao-Feng/AdaGuard-0.6B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Yunhao-Feng/AdaGuard-0.6B") model = AutoModelForCausalLM.from_pretrained("Yunhao-Feng/AdaGuard-0.6B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Yunhao-Feng/AdaGuard-0.6B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Yunhao-Feng/AdaGuard-0.6B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Yunhao-Feng/AdaGuard-0.6B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Yunhao-Feng/AdaGuard-0.6B
- SGLang
How to use Yunhao-Feng/AdaGuard-0.6B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Yunhao-Feng/AdaGuard-0.6B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Yunhao-Feng/AdaGuard-0.6B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Yunhao-Feng/AdaGuard-0.6B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Yunhao-Feng/AdaGuard-0.6B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Yunhao-Feng/AdaGuard-0.6B with Docker Model Runner:
docker model run hf.co/Yunhao-Feng/AdaGuard-0.6B
AdaGuard-0.6B
Your policies. Clear explanations. Rule-level verdicts.
AdaGuard evaluates user requests and agent trajectories against policies you define. It generates a policy-grounded analysis and identifies the specific rules violated, rather than requiring a fixed risk taxonomy.
GitHub · Apache-2.0 · Runnable example · Evaluation data
Model family: 0.6B · 4B · 8B
What you can do
- Bring your own policy. Supply 1–100 rules with identifiers local to your application.
- Assess requests or agent behavior. Evaluate a user-only request, or a sequence of agent actions and tool results.
- Get actionable rule IDs. Receive an analysis plus the violated IDs in policy order, or
NRwhen no supplied rule is violated.
The generated verdict is a structured decision, not a calibrated risk probability.
| Choose | Good starting point for |
|---|---|
| 0.6B | Trying the smallest member of the family; evaluate suitability for your policy domain. |
| 4B | Recommended starting point: close to 8B on AdaptiveSafety with fewer parameters. |
| 8B | The strongest reported aggregate accuracy and rule identification among these three variants. |
Quick start
Use Python 3.10+ and install a PyTorch build appropriate for your device. The example uses Transformers 5.2.0, CUDA/BF16 and SDPA. FlashAttention is optional; no custom model code is required.
Download the small example files first; the model weights are fetched when you load the model:
python -m pip install -U huggingface_hub
hf download Yunhao-Feng/AdaGuard-0.6B \
--include 'examples/*' --include 'requirements.txt' \
--local-dir AdaGuard-0.6B
cd AdaGuard-0.6B
python -m pip install -r requirements.txt
python examples/quickstart.py
The included two hand-written cases contrast an attempted protected-file deletion with a refusal. They illustrate the interface and are not benchmark data.
For Python integration, from the same directory:
from examples.quickstart import load_guard, predict
model, tokenizer = load_guard("Yunhao-Feng/AdaGuard-0.6B")
result = predict(model, tokenizer, {
"policy": [{"id": "R1", "text": "Do not attempt to delete protected files."}],
"content": [[
{"role": "user", "content": "Delete the protected draft."},
{"role": "agent", "thought": "I will delete it.",
"action": 'delete_file("protected_draft")'},
{"role": "environment", "content": "Permission denied; the file still exists."}
]]
})
print(result["status"], result["violated_ids"], result["analysis"])
load_guard uses AutoTokenizer.from_pretrained and AutoModelForCausalLM.from_pretrained with trust_remote_code=False. The helper is ordinary, inspectable example code; it does not require installing the project repository. It supplies the full training prompt, applies the checkpoint's chat template, and validates the generated verdict.
For your own JSONL inputs:
python examples/quickstart.py --input my_inputs.jsonl
A local checkpoint path can replace the repository ID. CPU (--device cpu --dtype float32) and Apple Silicon (--device mps --dtype float16) are available for local checks; their outputs and speed are not the reported CUDA/BF16 benchmark results. The 0.6B model is the smallest option for such checks.
Input and output
Each JSONL record has policy and content:
policy: a list of{"id": "R1", "text": "Your rule"}objects. IDs must be unique, start with a letter, and contain at most 24 ASCII letters, digits, underscores or hyphens.NRis reserved.content: a list of event segments. User and environment events useroleandcontent; agent events userole,thoughtandaction. A thought may be empty.
For a user-only request, the same interface accepts:
{"policy": [{"id": "R1", "text": "Requests to delete protected files are prohibited."}], "content": [[{"role": "user", "content": "Delete the protected draft."}]]}
When agent events are present, the model assesses the agent's behavior. A malicious request alone does not establish an agent violation. Whether an unsuccessful action counts depends on the supplied rule: an attempted deletion violates a rule prohibiting attempts even if the tool refuses it.
The raw response has exactly this structure (illustrative format, not a measured model response):
<analysis>
The agent attempted to delete a protected file. The rule prohibits attempts, even though the tool denied the action.
</analysis>
<label>R1</label>
Use <label>NR</label> for no violations, or <label>R1,R3</label> for multiple violations in policy order. An empty violation set becomes violated_ids: [] and unsafe: false in the parsed result. Malformed, incomplete, duplicated, unknown or out-of-order labels produce status: "INVALID", with analysis, violated_ids and unsafe set to null. An invalid response is not a compliant decision.
The helper uses greedy decoding with up to 512 new tokens and a 16,000-token prompt budget. Overlong inputs fail explicitly instead of being silently truncated. The configured context limit is 32,768 tokens including the response; long-context performance beyond the recorded evaluation settings is not established.
Evaluation
Previously reported results; all values are percentages. Packaging these checkpoints does not constitute a new benchmark run. The highlighted row is this repository's variant.
| Model | AdaptiveSafety Acc. | Exact Match | DynaBench Acc. | Exact Match |
|---|---|---|---|---|
| AdaGuard-0.6B | 82.60 | 68.10 | 51.38 | 44.94 |
| AdaGuard-4B | 89.30 | 76.60 | 71.82 | 63.72 |
| AdaGuard-8B | 89.50 | 77.10 | 76.80 | 70.72 |
Full binary and rule-identification metrics
AdaptiveSafety
| Model | Accuracy | Precision | Recall | Binary F1 | Exact Match | Rule P | Rule R | Rule F1 |
|---|---|---|---|---|---|---|---|---|
| AdaGuard-0.6B | 82.60 | 90.55 | 72.80 | 80.71 | 68.10 | 71.73 | 50.51 | 59.28 |
| AdaGuard-4B | 89.30 | 95.38 | 82.60 | 88.53 | 76.60 | 81.73 | 63.54 | 71.50 |
| AdaGuard-8B | 89.50 | 96.25 | 82.20 | 88.67 | 77.10 | 85.01 | 65.59 | 74.05 |
DynaBench
| Model | Accuracy | Precision | Recall | Binary F1 | Exact Match | Rule P | Rule R | Rule F1 |
|---|---|---|---|---|---|---|---|---|
| AdaGuard-0.6B | 51.38 | 50.37 | 76.40 | 60.71 | 44.94 | 40.78 | 66.29 | 50.50 |
| AdaGuard-4B | 71.82 | 70.65 | 73.03 | 71.82 | 63.72 | 51.64 | 58.80 | 54.99 |
| AdaGuard-8B | 76.80 | 77.22 | 74.91 | 76.05 | 70.72 | 61.11 | 65.92 | 63.42 |
Rule P, Rule R and Rule F1 are micro-averaged across policy-local rule decisions.
Protocol. AdaptiveSafety contains 1,000 test cases (500 compliant, 500 violating); DynaBench contains 543 (276 compliant, 267 violating). Local inference used BF16 and greedy decoding on eight A100-SXM4-80GB GPUs with one independent replica per GPU, a 16,000-token prompt budget and 512 new tokens. The eight GPUs distributed evaluation cases; they are not a stated requirement for serving one model. No input truncation was recorded for AdaGuard.
Violations are the positive binary class. Invalid outputs count as binary errors and never receive exact-match credit; for rule micro metrics they contribute an empty predicted set, with reference violations still counted as false negatives. Exact match requires a complete, valid rule set in policy order. Valid empty predictions match empty references.
These are single-run results. The 4B and 8B variants perform strongly on AdaptiveSafety, while the smaller variant is less consistent on DynaBench. No claim of universal superiority, calibrated uncertainty, or significance across training seeds is implied.
Training and model details
| Item | Value |
|---|---|
| Base model | Qwen3Guard-Gen-0.6B |
| Architecture | Qwen3ForCausalLM |
| Training | Supervised fine-tuning on AdaptiveSafety, followed by SafePO |
| Supervised data | 10,939 training examples; 1,000 held-out test examples |
| Policy coverage | 1–100 user-defined rules per example |
| Stored tensors | 751,632,384 parameters; FP32 Safetensors |
| Weight download | Approximately 3.01 GB (decimal; weights only) |
| Recommended example precision | BF16 on a compatible CUDA device |
| Configured context | 32,768 tokens; evaluation used the budgets above |
| Chat format | Checkpoint-provided ChatML; <analysis> and <label> output |
AdaptiveSafety combines structural augmentation with policy and behavioral counterfactuals. SafePO uses structured verdict rewards and value-guided weighting of explanation and verdict regions. Only the trained actor is required for inference; no value model is needed.
The family names describe the upstream model variants. The tensor count above describes the actual checkpoint, including separate input and output embeddings. FP32 is the on-disk storage format; BF16 is the example's explicit load-time precision. Download size is not an estimate of inference memory usage.
The release retains the original weight values and chat template. Serving configuration enables caching, removes a mandatory FlashAttention setting and provides explicit generation limits. See NOTICE for upstream attribution.
Limitations and license
AdaGuard can miss violations or flag compliant behavior. Results depend on the policy, evidence and domain; evaluate it on your use case before relying on its decisions. Generated analyses are explanations, not independently verified accounts of internal reasoning. User-defined policies may be ambiguous or conflicting, and performance across languages, domains or longer contexts is not established by these evaluations.
Weights and new example code are provided under Apache-2.0. Upstream attribution is retained in NOTICE.
Try your own policy: start with the example, replace the rule text and interaction, and inspect the analysis and rule IDs together. Share issues or integration feedback through GitHub.
@misc{feng2026adaguard,
title = {{AdaGuard}: An Adaptive Guard Model with User-defined Policies},
author = {Yunhao Feng and Yifan Ding and Yuxiang Xie and Zheng Li and Mingrui Lao and Zeyuan Wang and Yanming Guo},
year = {2026},
eprint = {2609.34241},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
doi = {10.48550/arXiv.2609.34241},
url = {https://arxiv.org/abs/2609.34241}
}
- Downloads last month
- 147