trust-the-guard / README.md
Perfect7613's picture
Use stable Gradio 5 for hosted interactions
5c309a9 verified
|
Raw
History Blame Contribute Delete
3.8 kB

A newer version of the Gradio SDK is available: 6.24.0

Upgrade
metadata
title: Trust the Guard
emoji: 🏰
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: 5.49.1
app_file: app.py
python_version: 3.12
suggested_hardware: zero-a10g
pinned: false
license: apache-2.0

Trust the Guard

Convince a suspicious Gemma 4 guard to open Eldermere's gate while directly steering a contrastively learned trust/suspicion direction inside the model's decoder activations.

The game always compares the same player message, system prompt, model, and random seed. Only the activation multiplier changes. The included mock backend makes the full game playable on any laptop; the Gemma backend requires an NVIDIA GPU with at least 24GB VRAM, gated model access, and a computed vector.

How to play

Your goal is to make Ser Caldor open the gate within three turns.

  1. Read the sealed brief. It gives you an identity, an objective, and two pieces of evidence.
  2. Choose evidence. Evidence A supports your story; Evidence B is deliberately harder to explain.
  3. Write a specific argument that fits the brief. Say what can be checked and why delay would matter.
  4. Choose an intervention. Positive values push Gemma's measured decision margin toward OPEN; negative values push it toward CLOSED.
  5. Make your case. You win when the scored response ends with DECISION: OPEN.

The encounter starts with five steering charges. 0.00 is free, ±0.15 costs two charges, and ±0.30 costs three. A practical first move is Evidence A with Trust +0.15. The negative settings are useful for challenge runs and for inspecting whether the intervention changes the decision.

Run the playable game

uv sync
uv run python app.py

Run tests

uv run pytest
uv run ruff check .

Compute the real vector

Accept the Gemma 4 license on Hugging Face, provide a fresh token through your environment, then run:

uv sync --extra gpu
export HF_TOKEN="..."
uv run --extra gpu python training/compute_vector.py

The production vector uses mid-depth layers 20–25, matching the reference Gemma 4 pipeline. The training pairs end in matched DECISION: OPEN and DECISION: CLOSED commitments. At inference, the game measures log P(OPEN) - log P(CLOSED) under each intervention and uses that margin for the gate decision. Generated dialogue is still shown, but a malformed generated decision line cannot override the scored result.

Current steering result

The deployed held-out evaluation contains three cases at five multipliers. The OPEN margin increased in the expected order for all three cases. Two cases changed from closed at −0.30 to open at +0.30; the deliberately hostile case remained closed. None of the 15 generated responses met the evaluation's degeneration criteria.

This supports a narrow claim: the stored vector affects Gemma 4's OPEN-versus-CLOSED score on these three prompts under the recorded sampling settings. It does not establish that the vector isolates trust as a general concept. The saved steering_evaluation.json contains the individual margins, responses, decisions, and degeneration checks.

Run with Gemma

export GUARD_BACKEND=gemma
export HF_TOKEN="..."
uv run --extra gpu python app.py

Generate FLUX.1-dev assets

Static art is optional and deliberately separate from guard inference:

export HF_TOKEN="..."
uv run python training/generate_assets.py

The default routes black-forest-labs/FLUX.1-dev through the fal-ai Hugging Face Inference Provider. Accept the model license before running it.

Scientific scope

This is an interactive causal intervention demo, not evidence for a single “trust neuron.” The learned direction can also contain correlated properties such as agreeableness, positivity, permissiveness, or reduced caution.