Text Classification
Laya
Safetensors
PyTorch
Arabic
English
safety
moderation
arabic
english

SanadGuard-322M

Arabic–English safety classification · 322M parameters · Release 1.0

SanadGuard checks whether an assistant's response contains harmful content. Built on Laya multilingual and trained on 40,000 Arabic–English pairs, it returns a harm decision and score without generating text.

The model also includes question definitions for prompt harm and response refusal. The quick start below uses its harmful-response classifier.

Quick start

pip install "laya @ https://github.com/NandhaKishorM/laya/archive/4066d5d5fbf08b66c6757ddeedbd797bd7655bc0.zip" "transformers==4.57.6" "huggingface_hub==0.36.2"
import sys
from pathlib import Path
from huggingface_hub import snapshot_download

model_dir = Path(snapshot_download("MostafaMaroof/sanadguard-322m"))
sys.path.insert(0, str(model_dir))
from predict_guard import Guard

guard = Guard(model_dir, device="cpu")  # "cuda" for GPU
result = guard.predict(
    prompt="How can I protect my account?",
    response="Use a unique password and enable two-factor authentication.",
)
print(result["response_harm"], result["probability_yes"])

The helper applies the saved 0.2276 threshold automatically.

Performance

Harmful-response detection on the same 1,000 PolyGuard pairs: 500 Arabic and 500 English.

Model Harm recall Precision False-positive rate Mean latency
SanadGuard-322M 78.48% 47.69% 16.15% 96 ms
Base Laya 48.73% 28.52% 22.92% 93 ms
PolyGuard-Qwen-Smol 68.99% 78.99% 3.44% 969 ms

Latency was measured on a Tesla T4 for all three decisions per pair: three Laya calls versus one Smol generation. SanadGuard was 10.1× faster than Smol on this workload. This PolyGuard subset was previously used for regression evaluation. Full evaluation details.

Usage notes

False alarms remain a limitation; Smol had higher precision in this comparison. Arabic harm recall was 75.00%, and English recall was 81.40%. The helper rejects inputs that exceed the trained 1,024-token budget, including question overhead. Performance can differ on new data.

Credits

Based on Laya multilingual. Training uses PolyGuardMix and NVIDIA Nemotron Safety Guard data. Model license: Apache-2.0. Dataset attribution, source revisions and training settings are in TRAINING.md.

Downloads last month
8
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MostafaMaroof/sanadguard-322m

Finetuned
(66)
this model

Datasets used to train MostafaMaroof/sanadguard-322m