Papers
arxiv:2609.06289

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

Published on Sep 5
· Submitted by
Mahdi Abootorabi
on Sep 9
Authors:
,
,
,
,
,
,

Abstract

Activation steering vectors in large language models encode theory-aligned human value geometry when derived via distribution-driven methods, with geometric fidelity scaling with model size but declining after instruction tuning.

As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically validates steering on isolated behaviors, leaving it unclear whether steering vectors encode coherent semantic structure or merely exploit behavior-specific shortcuts. We investigate whether the latent geometry of LLM steering vectors reflects theory-specified structure in human values and morality. Using Schwartz's Theory of Basic Human Values as our primary fine-grained framework, we introduce a 26K-sample benchmark covering 20 human values and analyze distribution-driven methods (e.g., CAA, SphericalSteer, ODESteer) and behavior-centric approaches (e.g., COLD-Steer, BiPO) across diverse model families and sizes. We find that distribution-driven methods recover human value topologies aligned with theoretical predictions (Spearman ρ up to 0.51, p < 10^{-13}). In contrast, behavior-centric methods achieve comparable steering performance but show little correlation with the expected value geometry. Geometric fidelity improves with model scale but drops after instruction tuning. Finally, better geometric alignment also leads to more human-consistent transfer across values: steering one value correctly lifts compatible values and suppresses opposing ones. Code and data are available at: https://github.com/DeepRCL/Steering_Geometry.

Community

Paper author Paper submitter
edited about 15 hours ago

We’re excited to share Steering Geometry, accepted to EMNLP 2026 Main!

When we steer an LLM toward one human value, what happens to the others? We investigate whether steering directions capture the relationships predicted by psychological theories—and whether that structure translates into more consistent behavior.

Our main contributions and findings:

  • 🧭 Geometry beyond steering accuracy: Across seven steering methods, distribution-driven approaches recover clearer human-value structure, while behavior-centric methods can achieve comparable steering performance with weak geometric alignment.
  • 🔄 Cross-value transfer: Better geometric alignment is associated with more theory-consistent transfer—steering toward one value strengthens compatible values and suppresses opposing ones.
  • 📈 Scale and instruction tuning: Value geometry improves with model scale but weakens after instruction tuning in our experiments.
  • 📊 Two value frameworks: We introduce a 26K-example contrastive benchmark covering 20 Schwartz values and a 1,200-example benchmark spanning six foundations of revised Moral Foundations Theory.
  • 🛠️ Tools for your own methods: Our code supports geometry analysis and cross-value transfer evaluation, making it straightforward to compare new steering interventions with the included methods.

💻 Code: DeepRCL/Steering_Geometry
🤗 Dataset: DeepRCL/SteeringGeometry
📄 Paper: Read on arXiv

We’d love to hear your thoughts, especially on evaluating steering beyond the target behavior and extending this analysis to other concepts and value frameworks!

Very creative and good point of view

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.06289
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.06289 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.06289 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.