Papers
arxiv:2609.36585

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Published on Sep 29
· Submitted by
Zehao Jin
on Oct 2
Authors:
,
,

Abstract

Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https://lunamos.github.io/stop-thinking-too-early/

Community

Paper submitter

TL;DR: Asked to answer directly (no CoT), LLMs follow reference chains like K = apple; B = K; D = B; print(D) for only a few lines. A tiny rank-8 LoRA at one early layer lets the frozen middle layers carry the chain much further. The computation is there; it just stops too early.

Key findings

  • 13 base models (0.6B–32B) reliably follow only 1.4–3.6 lines. Doubling depth (OLMo-3 7B → 32B) leaves reach at ~2.6 lines.
  • Qwen3-8B + a task-trained rank-8 LoRA at layer 14 (65,537 params, all base weights frozen): exact accuracy on 24-line chains goes from 15.5% → 99%. A longer-trained version reaches 50 lines in a single forward pass.
  • Mechanism: the LoRA acts on each token independently and moves no information between tokens. It starts a relay: program lines pass their chain identity along through frozen layers 16–22. Cutting each line's attention to its parent in layers 14–22 drops accuracy to chance; the same cut in layers 23–29 barely matters.
  • Placement cliff: same recipe, LoRA at layer 20 → 20.5 lines; at layer 21 → 5.2 lines. A frozen-model measurement located this limit within a preregistered tolerance in 3 of 4 held-out models.
  • Looped models: in Ouro-1.4B, a LoRA applied every loop reaches 60 lines after 4 loops and ≥160 after 8 (two-chain choice accuracy).
  • Multi-hop QA: on MuSiQue (gold paragraphs), separately trained early-layer LoRAs add 9.4–17.9 EM across three standard models.

🎬 Website : https://lunamos.github.io/stop-thinking-too-early/
💻 Code: https://github.com/Lunamos/stop-thinking-too-early

Happy to answer questions!

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.36585
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.36585 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.36585 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.36585 in a Space README.md to link it from this page.

Collections including this paper 1