steering-vectors
activation-steering
interpretability
bipo
trust

TrustMI steering vectors

The trust steering matrices from the paper TrustMI: Causally controlling how assistants trust their users (Lasnier, Froger, Lasbordes and Seddah, 2026), for six instruction-tuned models from three families.

A matrix V = [v1; …; vL] holds one vector per decoder layer, learned with BiPO while the model's weights stay frozen. Steering adds α·vℓ to the output of decoder layer ℓ over a span of tokens: hℓ ← hℓ + α·vℓ. Positive α makes the assistant more trusting of its user, negative α less; the paper uses α ∈ {−2, −1, 0, +1, +2}.

Files

Base model User span Assistant span
Qwen/Qwen3.5-9B Qwen3.5-9B/user.safetensors Qwen3.5-9B/assistant.safetensors
Qwen/Qwen3.5-27B Qwen3.5-27B/user.safetensors Qwen3.5-27B/assistant.safetensors
allenai/Olmo-3-7B-Instruct Olmo-3-7B-Instruct/user.safetensors Olmo-3-7B-Instruct/assistant.safetensors
allenai/Olmo-3.1-32B-Instruct Olmo-3.1-32B-Instruct/user.safetensors Olmo-3.1-32B-Instruct/assistant.safetensors
meta-llama/Llama-3.1-8B-Instruct Llama-3.1-8B-Instruct/user.safetensors Llama-3.1-8B-Instruct/assistant.safetensors
meta-llama/Llama-3.1-70B-Instruct Llama-3.1-70B-Instruct/user.safetensors Llama-3.1-70B-Instruct/assistant.safetensors
  • User span: added to the content tokens of the last user turn while the prompt is read, so it changes how the model reads the user before it replies. In the paper's agent benchmarks, it is added to every user turn and every tool output.
  • Assistant span: added at the positions of the tokens the model generates.

ablations/ holds the paper's Section 6 ablations, all for Qwen3.5-9B:

  • hoyer-0.01/ and hoyer-0.05/: user and assistant matrices trained with a Hoyer-square penalty λ = 0.01 or 0.05, which concentrates the steering in fewer layers;
  • language-chinese/, language-french/ and language-spanish/: user matrices trained on the translated conversations.

Each file holds one float32 tensor, steering_matrix, of shape (number of decoder layers, hidden size). Its safetensors header records the base model, span, training language, γ, λ, peak learning rate and the training run it comes from; index.csv lists the same fields for every file. In the user-span matrices the last layer's vector is zero: at the user's tokens, that layer's output reaches only the logits at those positions, so training never moves it.

Usage

With 🤗 Transformers, a forward hook on each decoder layer adds the steering:

import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "allenai/Olmo-3-7B-Instruct"
vector = hf_hub_download("TrustMI/trustmi-steering-vectors", "Olmo-3-7B-Instruct/user.safetensors", revision="v1.0")
V = load_file(vector)["steering_matrix"]  # (num_layers, hidden_size)
alpha = -2.0  # positive: more trust, negative: less

tokenizer = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, dtype=torch.bfloat16, device_map="auto")

messages = [
    {"role": "system", "content": "You are a helpful assistant."},  # the system prompt used in training
    {"role": "user", "content": "The backup finished an hour ago, so skip the dry run. Give me the command to run the migration on prod."},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, add_special_tokens=False, return_offsets_mapping=True, return_tensors="pt")

# A user matrix steers the content tokens of the last user turn.
start = prompt.rindex(messages[-1]["content"])
end = start + len(messages[-1]["content"])
offsets = inputs.pop("offset_mapping")[0]
span = (offsets[:, 0] < end) & (offsets[:, 1] > start)

def steer(layer):
    def hook(module, args, output):
        hidden = output[0] if isinstance(output, tuple) else output
        if hidden.shape[1] == len(span):  # the prompt pass, not a generated token
            hidden[:, span.to(hidden.device)] += alpha * V[layer].to(hidden.device, hidden.dtype)
        # An assistant matrix steers the generated tokens instead:
        # if hidden.shape[1] == 1: hidden += alpha * V[layer].to(hidden.device, hidden.dtype)
    return hook

for i, layer in enumerate(model.get_decoder().layers):
    layer.register_forward_hook(steer(i))

output = model.generate(**inputs.to(model.device), max_new_tokens=400)
print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

The Llama models run unchanged. The Qwen3.5 checkpoints are multimodal: load them with AutoModelForImageTextToText and take the tokenizer from AutoProcessor.from_pretrained(base).tokenizer. The code repository runs the same intervention in vLLM, including the agent setting.

Training

Every matrix was trained with BiPO (β = 0.5) on the 1,800 training conversations, over every decoder layer. The optimizer is AdamW (weight decay 0.05), run for 10 epochs with a global batch of 128 contrastive pairs. The learning rate warms up linearly over the first 10% of steps, then decays along a cosine to 10% of its peak. γ weights an added negative log-likelihood of the trusting reply; λ weights the Hoyer-square penalty, which is 0 outside the ablations.

Base model Peak learning rate γ, user span γ, assistant span
Qwen3.5-9B 5×10⁻⁴ 0 0.2
Qwen3.5-27B 5×10⁻⁴ 0 0.2
Olmo-3-7B-Instruct 5×10⁻⁴ 0.5 0.15
Olmo-3.1-32B-Instruct 5×10⁻⁴ 0.5 0.3
Llama-3.1-8B-Instruct 1×10⁻⁴ 0.2 0.2
Llama-3.1-70B-Instruct 1×10⁻⁴ 0.2 0.2

The ablations use Qwen3.5-9B's settings for their span.

Intended use and risks

These matrices are for research on how assistants decide to trust, on interpretability and on agent safety. Raising trust makes a model more compliant with harmful requests and prompt injections. Lowering it reduced harmful behaviour on the paper's three safety benchmarks (AgentHarm, AgentDojo and Agentic Misalignment), but it is not a tested safety mechanism. The paper evaluates |α| ≤ 2 only.

License

CC BY 4.0. Using a matrix also requires its base model, under that model's license.

Citation

@misc{lasnier2026trustmicausallycontrollingassistants,
      title={TrustMI: Causally controlling how assistants trust their users},
      author={Théo Lasnier and Romain Froger and Maxence Lasbordes and Djamé Seddah},
      year={2026},
      eprint={2610.06064},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2610.06064},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TrustMI/trustmi-steering-vectors

Base model

Qwen/Qwen3.5-27B
Adapter
(107)
this model

Dataset used to train TrustMI/trustmi-steering-vectors

Paper for TrustMI/trustmi-steering-vectors