Spaces:
Sleeping
Sleeping
| title: AgentGuard Bench | |
| emoji: 🛡️ | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: gradio | |
| sdk_version: 4.36.0 | |
| app_file: app.py | |
| pinned: false | |
| license: apache-2.0 | |
| tags: | |
| - llm-security | |
| - prompt-injection | |
| - guardrails | |
| - agentic-ai | |
| - safety | |
| # 🛡️ AgentGuard Bench | |
| A side-by-side comparison of a **pattern-based** prompt-injection | |
| guardrail and a **fine-tuned transformer classifier**, built to make | |
| the latency/robustness tradeoff between the two approaches concrete | |
| and explorable. | |
| - **Live Compare** — paste any prompt, see both guardrails score it instantly | |
| - **Adversarial Test Suite** — run a curated set of attack + benign prompts and compare accuracy and latency | |
| - **About** — methodology, related work, and limitations | |
| The pattern-based guardrail is a Python port of the `PromptInjectionGuardrail` | |
| I contributed to [LangChain4j](https://github.com/langchain4j/langchain4j) | |
| (Java) in [PR #5619](https://github.com/langchain4j/langchain4j/pull/5619) — | |
| same detection taxonomy, demonstrating it generalizes across ecosystems. | |
| The model-based guardrail uses | |
| [`protectai/deberta-v3-base-prompt-injection-v2`](https://huggingface.co/protectai/deberta-v3-base-prompt-injection-v2) | |
| (Apache 2.0). | |
| See the **About** tab inside the Space for full methodology, citations, | |
| and known limitations. | |