Papers
arxiv:2610.09087

U-Space: Uncovering When and Why Uncertainty Arises in Language Models

Published on Oct 6
· Submitted by
Tobias Braun
on Oct 9
Authors:
,
,
,
,

Abstract

Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer remains difficult because language models can present incorrect conclusions with fluent explanations and an authoritative tone. Uncertainty quantification seeks to address this disconnect by estimating the reliability of individual predictions. However, many existing methods require repeated generations or separately trained components, and their scalar estimates do not reveal where uncertainty arises or how it evolves during reasoning. Recent work has also shown that generation length can be strongly associated with uncertainty estimates and correctness, raising the question of how much of an estimator's predictive power comes from uncertainty-specific information rather than output length alone. Mechanistic interpretability offers a way to address these limitations by connecting human-interpretable concepts to intermediate model states. Building on this capability, we introduce the U-Space, a low-dimensional subspace that makes a model's evolving uncertainty measurable and interpretable. We identify semantic anchors for doubt and certainty, map their unembedding directions back into the residual space, and combine their contrasts into an orthogonal basis. The U-Lens projects each token state onto these basis vectors, yielding an interpretable token-level uncertainty map that can be inspected directly or aggregated into a scalar uncertainty score. Our approach requires no correctness labels, repeated generations, or training. Across reasoning benchmarks, its confidence score outperforms established baselines under both standard and length-controlled evaluation and transfers more reliably than supervised estimators. Code: https://github.com/s2labres/U-Space.

Community

Paper author Paper submitter
•
This comment has been hidden (marked as Resolved)
Paper author Paper submitter

Can Mechanistic Interpretability be used for Uncertainty Quantification?

We present the U-Space, an interpretable subspace that tracks four sources of uncertainty through a model’s generation: Ambiguity, Incomplete information, Conflicting evidence, and General uncertainty. The corresponding U-Lens combines this verbalizable signal with predictive entropy to estimate answer reliability from a single generation.
There is no need for multiple forward passes, correctness labels, or task-specific training!

Our main findings are:

  • Combining the verbalizable uncertainty measure from the U-Space with a first-order distributional uncertainty measure, predictive entropy, yields a state-of-the-art uncertainty predictor.
  • This predictor, the U-Lens, leads uncertainty estimation across three reasoning models: Gemma-4-31B, Qwen3.5-27B, and Magistral-Small, and four benchmarks: MMLU-Pro, Omni-MATH, SuperGPQA, and TriviaQA.
  • Response length is a strong predictor before controlling for it, reaching 68.6 AUROC, but falls to 52.4 under length control. Most other baselines drop in performance when we control for length. U-Lens drops by only 1.8 points and remains strongest.
  • Token-level readouts reveal when and what forms of uncertainty emerge during generation.
  • Steering along the identified directions can make models hesitate even on simple questions, showing that these representations are behaviorally relevant.
  • Preliminary experiments suggest that the same label-free construction extends beyond uncertainty to emotion and reward-hacking detection!

Project page: https://s2labres.github.io/U-Space/
Paper: https://arxiv.org/abs/2610.09087
Code: https://github.com/s2labres/U-Space

I was literally thinking I wanted this to be a thing a few days ago. Thank you! This is awesome.

·
Paper author

Thanks, we hope that people can integrate it and build on top of it. It's lightweight and easy to incorporate into a pipeline 😃 and its also just quite fun to play around with!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.09087
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.09087 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.09087 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.