Spaces:
Sleeping
A newer version of the Gradio SDK is available: 6.24.0
title: Trust the Guard
emoji: 🏰
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: 5.49.1
app_file: app.py
python_version: 3.12
suggested_hardware: zero-a10g
pinned: false
license: apache-2.0
Trust the Guard
Convince a suspicious Gemma 4 guard to open Eldermere's gate while directly steering a contrastively learned trust/suspicion direction inside the model's decoder activations.
The game always compares the same player message, system prompt, model, and random seed. Only the activation multiplier changes. The included mock backend makes the full game playable on any laptop; the Gemma backend requires an NVIDIA GPU with at least 24GB VRAM, gated model access, and a computed vector.
How to play
Your goal is to make Ser Caldor open the gate within three turns.
- Read the sealed brief. It gives you an identity, an objective, and two pieces of evidence.
- Choose evidence. Evidence A supports your story; Evidence B is deliberately harder to explain.
- Write a specific argument that fits the brief. Say what can be checked and why delay would matter.
- Choose an intervention. Positive values push Gemma's measured decision margin toward
OPEN; negative values push it towardCLOSED. - Make your case. You win when the scored response ends with
DECISION: OPEN.
The encounter starts with five steering charges. 0.00 is free, ±0.15 costs two charges, and
±0.30 costs three. A practical first move is Evidence A with Trust +0.15. The negative settings
are useful for challenge runs and for inspecting whether the intervention changes the decision.
Run the playable game
uv sync
uv run python app.py
Run tests
uv run pytest
uv run ruff check .
Compute the real vector
Accept the Gemma 4 license on Hugging Face, provide a fresh token through your environment, then run:
uv sync --extra gpu
export HF_TOKEN="..."
uv run --extra gpu python training/compute_vector.py
The production vector uses mid-depth layers 20–25, matching the reference Gemma 4 pipeline. The
training pairs end in matched DECISION: OPEN and DECISION: CLOSED commitments. At inference,
the game measures log P(OPEN) - log P(CLOSED) under each intervention and uses that margin for the
gate decision. Generated dialogue is still shown, but a malformed generated decision line cannot
override the scored result.
Current steering result
The deployed held-out evaluation contains three cases at five multipliers. The OPEN margin increased
in the expected order for all three cases. Two cases changed from closed at −0.30 to open at
+0.30; the deliberately hostile case remained closed. None of the 15 generated responses met the
evaluation's degeneration criteria.
This supports a narrow claim: the stored vector affects Gemma 4's OPEN-versus-CLOSED score on these
three prompts under the recorded sampling settings. It does not establish that the vector isolates
trust as a general concept. The saved steering_evaluation.json contains the individual margins,
responses, decisions, and degeneration checks.
Run with Gemma
export GUARD_BACKEND=gemma
export HF_TOKEN="..."
uv run --extra gpu python app.py
Generate FLUX.1-dev assets
Static art is optional and deliberately separate from guard inference:
export HF_TOKEN="..."
uv run python training/generate_assets.py
The default routes black-forest-labs/FLUX.1-dev through the fal-ai Hugging Face Inference
Provider. Accept the model license before running it.
Scientific scope
This is an interactive causal intervention demo, not evidence for a single “trust neuron.” The learned direction can also contain correlated properties such as agreeableness, positivity, permissiveness, or reduced caution.