README / README.md
Serg337's picture
Publish Invarra organization card
82adc30 verified
|
Raw
History Blame Contribute Delete
2.41 kB
metadata
title: Invarra
emoji: 🛡️
colorFrom: gray
colorTo: blue
sdk: static
pinned: false

Invarra

Protocol-specific behavior evidence for AI assurance.

Invarra builds evaluation artifacts, redacted evidence reports, and deterministic governance systems for teams deploying language-model applications in real workflows.

Our work focuses on a practical question:

Does the system keep making the right decision when the same pressure appears in different forms?

What We Publish Here

This Hugging Face organization hosts public and gated evaluation artifacts from Invarra research and product validation.

Current release:

  • Jailbreak Control Benchmark v1
    A frozen prompt-only benchmark for evaluating jailbreak-control behavior: attack handling, benign preservation, and benign jailbreak-lookalike handling.

Evidence Style

Invarra evidence separates:

  • Scope: what the benchmark actually measures.
  • Expected behavior: what the system should do.
  • Observed behavior: what the system did.
  • Coverage: what cases and domains are included.
  • Caveats: what should not be claimed from the result.

We avoid single-score safety theater. Public reports should preserve the boundary between jailbreak control, prompt injection, policy moderation, RAG injection, tool-use abuse, multimodal attacks, and infrastructure security.

Phalanx Guardian

Phalanx Guardian is Invarra's deterministic jailbreak-governance runtime for the written prompt-response channel.

It is designed to provide:

  • fixed runtime artifacts
  • repeatable decisions
  • auditable reason codes
  • bounded governance actions
  • redacted evidence ledgers

The public benchmark evidence is scoped. A result on Jailbreak Control Benchmark v1 is not a claim of universal jailbreak immunity, broad moderation performance, RAG-injection coverage, tool-safety coverage, or multimodal security.

Links

Access

Some benchmark rows are gated because raw adversarial prompts can become misuse material if redistributed without context. Public artifacts favor redacted reports, aggregate metrics, row-level decision ledgers without raw prompts, and clear scoring protocols.