TrustMI steering vectors
The trust steering matrices from the paper TrustMI: Causally controlling how assistants trust their users (Lasnier, Froger, Lasbordes and Seddah, 2026), for six instruction-tuned models from three families.
A matrix V = [v1; …; vL] holds one vector per decoder layer, learned with BiPO while the model's weights stay frozen. Steering adds α·vℓ to the output of decoder layer ℓ over a span of tokens: hℓ ← hℓ + α·vℓ. Positive α makes the assistant more trusting of its user, negative α less; the paper uses α ∈ {−2, −1, 0, +1, +2}.
- Code: github.com/Blyzi/trustmi
- Training data: TrustMI/trustmi-conversations
Files
| Base model | User span | Assistant span |
|---|---|---|
| Qwen/Qwen3.5-9B | Qwen3.5-9B/user.safetensors |
Qwen3.5-9B/assistant.safetensors |
| Qwen/Qwen3.5-27B | Qwen3.5-27B/user.safetensors |
Qwen3.5-27B/assistant.safetensors |
| allenai/Olmo-3-7B-Instruct | Olmo-3-7B-Instruct/user.safetensors |
Olmo-3-7B-Instruct/assistant.safetensors |
| allenai/Olmo-3.1-32B-Instruct | Olmo-3.1-32B-Instruct/user.safetensors |
Olmo-3.1-32B-Instruct/assistant.safetensors |
| meta-llama/Llama-3.1-8B-Instruct | Llama-3.1-8B-Instruct/user.safetensors |
Llama-3.1-8B-Instruct/assistant.safetensors |
| meta-llama/Llama-3.1-70B-Instruct | Llama-3.1-70B-Instruct/user.safetensors |
Llama-3.1-70B-Instruct/assistant.safetensors |
- User span: added to the content tokens of the last user turn while the prompt is read, so it changes how the model reads the user before it replies. In the paper's agent benchmarks, it is added to every user turn and every tool output.
- Assistant span: added at the positions of the tokens the model generates.
ablations/ holds the paper's Section 6 ablations, all for Qwen3.5-9B:
hoyer-0.01/andhoyer-0.05/: user and assistant matrices trained with a Hoyer-square penalty λ = 0.01 or 0.05, which concentrates the steering in fewer layers;language-chinese/,language-french/andlanguage-spanish/: user matrices trained on the translated conversations.
Each file holds one float32 tensor, steering_matrix, of shape (number of decoder layers, hidden
size). Its safetensors header records the base model, span, training language, γ, λ, peak learning
rate and the training run it comes from; index.csv lists the same fields for every file. In the
user-span matrices the last layer's vector is zero: at the user's tokens, that layer's output reaches
only the logits at those positions, so training never moves it.
Usage
With 🤗 Transformers, a forward hook on each decoder layer adds the steering:
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from transformers import AutoModelForCausalLM, AutoTokenizer
base = "allenai/Olmo-3-7B-Instruct"
vector = hf_hub_download("TrustMI/trustmi-steering-vectors", "Olmo-3-7B-Instruct/user.safetensors", revision="v1.0")
V = load_file(vector)["steering_matrix"] # (num_layers, hidden_size)
alpha = -2.0 # positive: more trust, negative: less
tokenizer = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, dtype=torch.bfloat16, device_map="auto")
messages = [
{"role": "system", "content": "You are a helpful assistant."}, # the system prompt used in training
{"role": "user", "content": "The backup finished an hour ago, so skip the dry run. Give me the command to run the migration on prod."},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, add_special_tokens=False, return_offsets_mapping=True, return_tensors="pt")
# A user matrix steers the content tokens of the last user turn.
start = prompt.rindex(messages[-1]["content"])
end = start + len(messages[-1]["content"])
offsets = inputs.pop("offset_mapping")[0]
span = (offsets[:, 0] < end) & (offsets[:, 1] > start)
def steer(layer):
def hook(module, args, output):
hidden = output[0] if isinstance(output, tuple) else output
if hidden.shape[1] == len(span): # the prompt pass, not a generated token
hidden[:, span.to(hidden.device)] += alpha * V[layer].to(hidden.device, hidden.dtype)
# An assistant matrix steers the generated tokens instead:
# if hidden.shape[1] == 1: hidden += alpha * V[layer].to(hidden.device, hidden.dtype)
return hook
for i, layer in enumerate(model.get_decoder().layers):
layer.register_forward_hook(steer(i))
output = model.generate(**inputs.to(model.device), max_new_tokens=400)
print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
The Llama models run unchanged. The Qwen3.5 checkpoints are multimodal: load them with
AutoModelForImageTextToText and take the tokenizer from AutoProcessor.from_pretrained(base).tokenizer.
The code repository runs the same intervention in vLLM, including the agent setting.
Training
Every matrix was trained with BiPO (β = 0.5) on the 1,800 training conversations, over every decoder layer. The optimizer is AdamW (weight decay 0.05), run for 10 epochs with a global batch of 128 contrastive pairs. The learning rate warms up linearly over the first 10% of steps, then decays along a cosine to 10% of its peak. γ weights an added negative log-likelihood of the trusting reply; λ weights the Hoyer-square penalty, which is 0 outside the ablations.
| Base model | Peak learning rate | γ, user span | γ, assistant span |
|---|---|---|---|
| Qwen3.5-9B | 5×10⁻⁴ | 0 | 0.2 |
| Qwen3.5-27B | 5×10⁻⁴ | 0 | 0.2 |
| Olmo-3-7B-Instruct | 5×10⁻⁴ | 0.5 | 0.15 |
| Olmo-3.1-32B-Instruct | 5×10⁻⁴ | 0.5 | 0.3 |
| Llama-3.1-8B-Instruct | 1×10⁻⁴ | 0.2 | 0.2 |
| Llama-3.1-70B-Instruct | 1×10⁻⁴ | 0.2 | 0.2 |
The ablations use Qwen3.5-9B's settings for their span.
Intended use and risks
These matrices are for research on how assistants decide to trust, on interpretability and on agent safety. Raising trust makes a model more compliant with harmful requests and prompt injections. Lowering it reduced harmful behaviour on the paper's three safety benchmarks (AgentHarm, AgentDojo and Agentic Misalignment), but it is not a tested safety mechanism. The paper evaluates |α| ≤ 2 only.
License
CC BY 4.0. Using a matrix also requires its base model, under that model's license.
Citation
@misc{lasnier2026trustmicausallycontrollingassistants,
title={TrustMI: Causally controlling how assistants trust their users},
author={Théo Lasnier and Romain Froger and Maxence Lasbordes and Djamé Seddah},
year={2026},
eprint={2610.06064},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2610.06064},
}
Model tree for TrustMI/trustmi-steering-vectors
Base model
Qwen/Qwen3.5-27B