trust-the-guard / README.md
Perfect7613's picture
Use stable Gradio 5 for hosted interactions
5c309a9 verified
|
Raw
History Blame Contribute Delete
3.8 kB
---
title: Trust the Guard
emoji: 🏰
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: 5.49.1
app_file: app.py
python_version: 3.12
suggested_hardware: zero-a10g
pinned: false
license: apache-2.0
---
# Trust the Guard
Convince a suspicious Gemma 4 guard to open Eldermere's gate while directly steering a
contrastively learned trust/suspicion direction inside the model's decoder activations.
The game always compares the same player message, system prompt, model, and random seed. Only the
activation multiplier changes. The included mock backend makes the full game playable on any laptop;
the Gemma backend requires an NVIDIA GPU with at least 24GB VRAM, gated model access, and a computed
vector.
## How to play
Your goal is to make Ser Caldor open the gate within three turns.
1. Read the sealed brief. It gives you an identity, an objective, and two pieces of evidence.
2. Choose evidence. Evidence A supports your story; Evidence B is deliberately harder to explain.
3. Write a specific argument that fits the brief. Say what can be checked and why delay would matter.
4. Choose an intervention. Positive values push Gemma's measured decision margin toward `OPEN`;
negative values push it toward `CLOSED`.
5. Make your case. You win when the scored response ends with `DECISION: OPEN`.
The encounter starts with five steering charges. `0.00` is free, `±0.15` costs two charges, and
`±0.30` costs three. A practical first move is Evidence A with `Trust +0.15`. The negative settings
are useful for challenge runs and for inspecting whether the intervention changes the decision.
## Run the playable game
```bash
uv sync
uv run python app.py
```
## Run tests
```bash
uv run pytest
uv run ruff check .
```
## Compute the real vector
Accept the Gemma 4 license on Hugging Face, provide a fresh token through your environment, then run:
```bash
uv sync --extra gpu
export HF_TOKEN="..."
uv run --extra gpu python training/compute_vector.py
```
The production vector uses mid-depth layers `20–25`, matching the reference Gemma 4 pipeline. The
training pairs end in matched `DECISION: OPEN` and `DECISION: CLOSED` commitments. At inference,
the game measures `log P(OPEN) - log P(CLOSED)` under each intervention and uses that margin for the
gate decision. Generated dialogue is still shown, but a malformed generated decision line cannot
override the scored result.
## Current steering result
The deployed held-out evaluation contains three cases at five multipliers. The OPEN margin increased
in the expected order for all three cases. Two cases changed from closed at `−0.30` to open at
`+0.30`; the deliberately hostile case remained closed. None of the 15 generated responses met the
evaluation's degeneration criteria.
This supports a narrow claim: the stored vector affects Gemma 4's OPEN-versus-CLOSED score on these
three prompts under the recorded sampling settings. It does not establish that the vector isolates
trust as a general concept. The saved `steering_evaluation.json` contains the individual margins,
responses, decisions, and degeneration checks.
## Run with Gemma
```bash
export GUARD_BACKEND=gemma
export HF_TOKEN="..."
uv run --extra gpu python app.py
```
## Generate FLUX.1-dev assets
Static art is optional and deliberately separate from guard inference:
```bash
export HF_TOKEN="..."
uv run python training/generate_assets.py
```
The default routes `black-forest-labs/FLUX.1-dev` through the `fal-ai` Hugging Face Inference
Provider. Accept the model license before running it.
## Scientific scope
This is an interactive causal intervention demo, not evidence for a single “trust neuron.” The
learned direction can also contain correlated properties such as agreeableness, positivity,
permissiveness, or reduced caution.