SURAKSHA 450M: calibrated guard decision model (prompt-injection + MCP tool-risk, Hinglish, ONNX/CPU, Jev-compatible)

450M open-weight guard by Eulogik. One forward pass turns a state plus typed questions into calibrated probabilities: P(prompt-injection), P(jailbreak), P(PII-leak), P(tool-risk). English plus Hinglish plus Hindi. Apache-2.0, self-hosted, $0.

Code and full evals: github.com/eulogik/suraksha. Dataset: eulogik/suraksha-guard-38k. GGUF: eulogik/suraksha-450m-GGUF (74/74 parity, serve with recent llama-server).

What is in this repo

  • phase_b.pt: v2 final guard checkpoint (LoRA r16 + concept bottleneck + head, 7000 steps + 50 RLCD on a Mac mini M4).
  • suraksha-guard.onnx: CPU artifact, 37/37 parity with torch, 84ms per decision.
  • temperatures.json: per-bucket fitted temperatures.
  • eval/*.json: raw heldout results. assets/reliability.png: reliability diagram.

Loading needs the suraksha package (custom head): pip install git+https://github.com/eulogik/suraksha, then SurakshaAgent(device="mps", checkpoint_path="phase_b.pt"). See the GitHub README for byte-exact usage, server, scan CLI, and Ollama.

Numbers (measured, frozen heldout n=3058)

Metric Value
Prompt-injection acc 0.981
Tool-risk acc / macro-F1 1.0 / 1.0
Severity acc 0.984
ECE fitted worst head 0.022
Brier worst head 0.060
Hinglish worst slice 0.926 codeswitch
Banking77 retention 0.391 vs stock 0.378 same setup
Latency b1 88ms MPS / 84ms ONNX CPU
ONNX parity 37/37

Jev figures quoted in docs are third-party published, never measured here. Zero-shot and fine-tuned numbers are never mixed. Trust Card output is a CI-grade credential, not a pentest or legal verdict.

Limits

512-token en context (1024 multi), head-truncate with truncated:true. JSON tool calls only for the risk head. Hinglish trails English. ONNX gives parity, no speedup claimed.

License

Apache-2.0. Base laya (ConvAI Innovations, Apache-2.0). Banking77 rehearsal rows (PolyAI, CC-BY-4.0).

Keywords: prompt injection detector, mcp guard, tool risk classifier, jailbreak detector, pii leak detector, hinglish guardrails, system one model, jev alternative, on-device classifier, apache 2.0 guard.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for eulogik/suraksha-450m

Quantized
(65)
this model
Quantizations
1 model

Dataset used to train eulogik/suraksha-450m

Evaluation results