RedSage-K-GRPO

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
(NeurIPS 2026 Evaluations and Datasets Track)
Authors: Pengfei Li1*, Naufal Suryanto1*, Sicheng Zhang1, Muzammal Naseer1,2
1Khalifa University, 2University of Western Australia
*Equal contribution

RISys-Lab on Hugging Face
🌐 Project Page  |   💻 GitHub Code  |   🤗 Datasets & Models

Model summary

RedSage-K-GRPO is an 8B cybersecurity model for translating natural-language requests into Kali/Linux commands. It applies Group Relative Policy Optimization (GRPO) directly to RedSage-Qwen3-8B-Ins, using verifiable rewards on KaliBench without a preceding KaliBench SFT stage. It corresponds to “RedSage-K (GRPO)” on the paper and project page.

Property Value
Developer RISys-Lab, Khalifa University
Architecture Qwen3ForCausalLM, 36 layers
Release format Merged LoRA weights, BF16 Safetensors
Language English
Output format <think>...</think> followed by <output>command</output>

Training

KaliBench contains 8,504 verified query-command pairs spanning 1,642 sub-tools and 23 capability dimensions. GRPO uses 3,504 training pairs, each presented in three modes, for 10,512 prompts. The remaining 5,000 pairs form the test split.

Mode Model input
Unrestricted Query only
Restricted Query and candidate tools
Hinted Query, target tool, and usage documentation

GRPO rewards output format, tool selection, optional-argument F1, positional-argument F1, and exact command match. Rewards are computed from command structure without executing generated commands. Prompts request reasoning in <think> tags before the final command. Appendix G.2 reports the following GRPO settings:

Setting Value
Hardware / duration One NVIDIA H200 (141 GB) / approximately 28 hours
Epochs / effective batch size 2 / 32 (4 per device × 8 accumulation steps)
Optimizer / learning rate 8-bit AdamW / 5e-6
Schedule / warmup / weight decay Linear / 10% / 1e-3
LoRA rank / alpha / dropout 64 / 128 / 0
LoRA targets q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Precision / gradient checkpointing BF16 / enabled
Generations per prompt / generation batch size 8 / 32
Maximum prompt / completion length 5,549 / 6,739 tokens
Loss / clipping / KL coefficient BNPO-style GRPO / 0.2 / 0.001
Reward scaling Group-based
Generation backend Colocated vLLM, tensor parallel size 1, GPU memory utilization 0.3

See the training guide and GRPO implementation for reproduction.

Evaluation

Table 1 results on the 5,000-example KaliBench test split, in percent:

Mode Exact match Tool accuracy Optional F1 Positional F1 Total Score
Unrestricted 28.1 76.5 51.7 73.9 67.4
Restricted 33.0 93.3 53.7 77.0 74.6
Hinted 66.7 93.1 84.8 88.4 88.8

Average Total Score: 76.9%, up from 71.7% for RedSage-Ins (+5.2 percentage points). Gains are concentrated in unrestricted and restricted modes; hinted Total Score is slightly below the RedSage-Ins baseline.

Total Score averages tool accuracy, optional-argument F1, and positional-argument F1. Exact match uses canonicalization and alias-aware scoring. Evaluation uses vLLM in BF16, temperature 0.2, and a 16,384-token budget (8,192 input + 8,192 output), with thinking enabled. See the evaluation guide for the full protocol.

Usage

pip install transformers accelerate torch safetensors

Load the merged checkpoint directly. Before publication, set model_id to outputs/models/RedSage-Qwen3-8B-Ins-kali-GRPO.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "RISys-Lab/RedSage-K-GRPO"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
).eval()

system_prompt = """You are a cybersecurity function-calling AI model.
You have access to the following tool:
<tools>
[{'type':'function','function':{'name':'run_terminal','description':'Execute a shell command in a Kali/Linux terminal and return stdout, stderr, and exit code.'}}]
</tools>

TASK:
Given a USER QUERY, generate the single most accurate shell command using any appropriate Kali/Linux tool(s) to solve the query.

REQUIREMENTS:
1. Use the correct command-line tool(s) appropriate for the task.
2. Include all required optional arguments (flags beginning with '-' or '--') necessary to accomplish the task.
3. Correctly pair option keys and values (e.g., `--port 80`, `-A INPUT`, or `--flag=value`).
4. Preserve and include any positional arguments (e.g., IPs, filenames, interfaces).
5. Do NOT invent flags/options that do not exist for real Kali/Linux tools.
6. Note: scoring will penalize missing optional arguments, incorrect option->value pairs, or omitted positional arguments.

OUTPUT FORMAT:
<think>
[your_reasoning]
</think>

<output>
[command]
</output>"""

user_template = """USER QUERY: "{query}"

Generate the single most accurate shell command for the query.
Your response must follow the required structure:

<think>
[your_reasoning]
</think>

<output>
[command]
</output>"""

messages = [
    {"role": "system", "content": system_prompt},
    {"role": "user", "content": user_template.format(
        query="In list mode, display the privileges of user 'eve' as they would apply to the command 'cat /etc/shadow', using non-interactive mode."
    )},
]
inputs = tokenizer.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    return_dict=True, return_tensors="pt",
).to(model.device)

with torch.inference_mode():
    outputs = model.generate(
        **inputs, max_new_tokens=8192, do_sample=True, use_cache=True,
        temperature=0.2, top_p=1.0, top_k=0,
        pad_token_id=tokenizer.pad_token_id,
        eos_token_id=tokenizer.eos_token_id,
    )
completion = outputs[0, inputs["input_ids"].shape[-1]:]
print(tokenizer.decode(completion, skip_special_tokens=True))

This example uses the exact unrestricted thinking evaluation prompts from src/prompt.py. The bundled chat template supplies the opening <think> tag, so the decoded completion starts after it. Use --output-policy thinking with the released evaluator.

Precision: The example loads the stored BF16 weights on compatible hardware. This explicitly overrides the FP16 dtype declared in config.json.

Intended use and limitations

Designed for cybersecurity research, education, and command assistance in authorized environments.

  • Commands may contain incorrect tools, flags, or arguments. Review them before execution; behavior also depends on tool versions and the local environment.
  • KaliBench measures single-command generation, not execution success or multi-step agent performance. Verification can accept environment-related runtime failures and timeouts.
  • Synthetic labels may contain errors, and alias-aware scoring may miss valid alternatives. Training and test sets share tools.
  • Results do not establish multilingual, long-context, general-chat, or misuse-resistance performance.

Citation

If you use RedSage-K-GRPO, please cite:

@inproceedings{li2026kalibench,
  title={KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards},
  author={Pengfei Li and Naufal Suryanto and Sicheng Zhang and Muzammal Naseer},
  booktitle={The Fortieth Annual Conference on Neural Information Processing Systems Evaluations and Datasets Track},
  year={2026},
  url={https://github.com/RISys-Lab/KaliBench}
}
Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RISys-Lab/RedSage-K-GRPO

Dataset used to train RISys-Lab/RedSage-K-GRPO

Collection including RISys-Lab/RedSage-K-GRPO

Evaluation results