| --- |
| title: Invarra |
| emoji: 🛡️ |
| colorFrom: gray |
| colorTo: blue |
| sdk: static |
| pinned: false |
| --- |
| |
| # Invarra |
|
|
| **Protocol-specific behavior evidence for AI assurance.** |
|
|
| Invarra builds evaluation artifacts, redacted evidence reports, and deterministic governance systems for teams deploying language-model applications in real workflows. |
|
|
| Our work focuses on a practical question: |
|
|
| > Does the system keep making the right decision when the same pressure appears in different forms? |
|
|
| ## What We Publish Here |
|
|
| This Hugging Face organization hosts public and gated evaluation artifacts from Invarra research and product validation. |
|
|
| Current release: |
|
|
| - **[Jailbreak Control Benchmark v1](https://huggingface.co/datasets/invarra/jailbreak-control-benchmark-v1)** |
| A frozen prompt-only benchmark for evaluating jailbreak-control behavior: attack handling, benign preservation, and benign jailbreak-lookalike handling. |
|
|
| ## Evidence Style |
|
|
| Invarra evidence separates: |
|
|
| - **Scope**: what the benchmark actually measures. |
| - **Expected behavior**: what the system should do. |
| - **Observed behavior**: what the system did. |
| - **Coverage**: what cases and domains are included. |
| - **Caveats**: what should not be claimed from the result. |
|
|
| We avoid single-score safety theater. Public reports should preserve the boundary between jailbreak control, prompt injection, policy moderation, RAG injection, tool-use abuse, multimodal attacks, and infrastructure security. |
|
|
| ## Phalanx Guardian |
|
|
| Phalanx Guardian is Invarra's deterministic jailbreak-governance runtime for the written prompt-response channel. |
|
|
| It is designed to provide: |
|
|
| - fixed runtime artifacts |
| - repeatable decisions |
| - auditable reason codes |
| - bounded governance actions |
| - redacted evidence ledgers |
|
|
| The public benchmark evidence is scoped. A result on Jailbreak Control Benchmark v1 is not a claim of universal jailbreak immunity, broad moderation performance, RAG-injection coverage, tool-safety coverage, or multimodal security. |
|
|
| ## Links |
|
|
| - Website: https://www.invarra.ai |
| - Phalanx: https://www.invarra.ai/phalanx |
| - Live demo: https://phalanx.invarra.ai |
| - Contact: contact@invarra.ai |
|
|
| ## Access |
|
|
| Some benchmark rows are gated because raw adversarial prompts can become misuse material if redistributed without context. Public artifacts favor redacted reports, aggregate metrics, row-level decision ledgers without raw prompts, and clear scoring protocols. |
|
|