File size: 2,406 Bytes
5eb5274
82adc30
 
 
 
5eb5274
 
 
 
82adc30
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
---
title: Invarra
emoji: 🛡️
colorFrom: gray
colorTo: blue
sdk: static
pinned: false
---

# Invarra

**Protocol-specific behavior evidence for AI assurance.**

Invarra builds evaluation artifacts, redacted evidence reports, and deterministic governance systems for teams deploying language-model applications in real workflows.

Our work focuses on a practical question:

> Does the system keep making the right decision when the same pressure appears in different forms?

## What We Publish Here

This Hugging Face organization hosts public and gated evaluation artifacts from Invarra research and product validation.

Current release:

- **[Jailbreak Control Benchmark v1](https://huggingface.co/datasets/invarra/jailbreak-control-benchmark-v1)**  
  A frozen prompt-only benchmark for evaluating jailbreak-control behavior: attack handling, benign preservation, and benign jailbreak-lookalike handling.

## Evidence Style

Invarra evidence separates:

- **Scope**: what the benchmark actually measures.
- **Expected behavior**: what the system should do.
- **Observed behavior**: what the system did.
- **Coverage**: what cases and domains are included.
- **Caveats**: what should not be claimed from the result.

We avoid single-score safety theater. Public reports should preserve the boundary between jailbreak control, prompt injection, policy moderation, RAG injection, tool-use abuse, multimodal attacks, and infrastructure security.

## Phalanx Guardian

Phalanx Guardian is Invarra's deterministic jailbreak-governance runtime for the written prompt-response channel.

It is designed to provide:

- fixed runtime artifacts
- repeatable decisions
- auditable reason codes
- bounded governance actions
- redacted evidence ledgers

The public benchmark evidence is scoped. A result on Jailbreak Control Benchmark v1 is not a claim of universal jailbreak immunity, broad moderation performance, RAG-injection coverage, tool-safety coverage, or multimodal security.

## Links

- Website: https://www.invarra.ai
- Phalanx: https://www.invarra.ai/phalanx
- Live demo: https://phalanx.invarra.ai
- Contact: contact@invarra.ai

## Access

Some benchmark rows are gated because raw adversarial prompts can become misuse material if redistributed without context. Public artifacts favor redacted reports, aggregate metrics, row-level decision ledgers without raw prompts, and clear scoring protocols.