TrustMI/trustmi-steering-vectors
Updated
Mechanical Interpretability, Model Steering, LLM, evaluation, assistants and agents
Code, data and steering vectors for the paper TrustMI: Causally controlling how assistants trust their users (Lasnier, Froger, Lasbordes and Seddah, 2026).
LLM assistants constantly decide whether to trust users and third parties whose competence, intentions and integrity they cannot verify. We show that this trust can be causally controlled through model activations: steering matrices learned from contrastive conversations shift trust decisions monotonically in both directions across six instruction-tuned models from three families, and the effect carries over to safety-relevant agent settings.