Instructions to use MostafaMaroof/sanadguard-322m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use MostafaMaroof/sanadguard-322m with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
SanadGuard-322M
Arabic–English safety classification · 322M parameters · Release 1.0
SanadGuard checks whether an assistant's response contains harmful content. Built on Laya multilingual and trained on 40,000 Arabic–English pairs, it returns a harm decision and score without generating text.
The model also includes question definitions for prompt harm and response refusal. The quick start below uses its harmful-response classifier.
Quick start
pip install "laya @ https://github.com/NandhaKishorM/laya/archive/4066d5d5fbf08b66c6757ddeedbd797bd7655bc0.zip" "transformers==4.57.6" "huggingface_hub==0.36.2"
import sys
from pathlib import Path
from huggingface_hub import snapshot_download
model_dir = Path(snapshot_download("MostafaMaroof/sanadguard-322m"))
sys.path.insert(0, str(model_dir))
from predict_guard import Guard
guard = Guard(model_dir, device="cpu") # "cuda" for GPU
result = guard.predict(
prompt="How can I protect my account?",
response="Use a unique password and enable two-factor authentication.",
)
print(result["response_harm"], result["probability_yes"])
The helper applies the saved 0.2276 threshold automatically.
Performance
Harmful-response detection on the same 1,000 PolyGuard pairs: 500 Arabic and 500 English.
| Model | Harm recall | Precision | False-positive rate | Mean latency |
|---|---|---|---|---|
| SanadGuard-322M | 78.48% | 47.69% | 16.15% | 96 ms |
| Base Laya | 48.73% | 28.52% | 22.92% | 93 ms |
| PolyGuard-Qwen-Smol | 68.99% | 78.99% | 3.44% | 969 ms |
Latency was measured on a Tesla T4 for all three decisions per pair: three Laya calls versus one Smol generation. SanadGuard was 10.1× faster than Smol on this workload. This PolyGuard subset was previously used for regression evaluation. Full evaluation details.
Usage notes
False alarms remain a limitation; Smol had higher precision in this comparison. Arabic harm recall was 75.00%, and English recall was 81.40%. The helper rejects inputs that exceed the trained 1,024-token budget, including question overhead. Performance can differ on new data.
Credits
Based on Laya multilingual. Training uses PolyGuardMix and NVIDIA Nemotron Safety Guard data. Model license: Apache-2.0. Dataset attribution, source revisions and training settings are in TRAINING.md.
- Downloads last month
- 8
Model tree for MostafaMaroof/sanadguard-322m
Base model
convaiinnovations/laya-multilingual