Papers
arxiv:2609.32534

DepthBench: Measuring How Residual Connections Enable More Computational Depth

Published on Sep 26
· Submitted by
Shiwei Liu
on Sep 29
Authors:
,
,
,
,
,

Abstract

Depth is a natural way to increase the computational capacity in Transformers, yet the contribution of deeper layers can diminish as depth grows larger. Recent approaches enhance normalization (e.g., LayerNorm Scaling) or residual connections (e.g., mHC, AttnRes) to enable better information flow and depth utilization. However, it remains unclear whether they truly translate increased architectural depth into effective computational depth, and whether their reported gains stem from better access to information across depth, or unaccounted-for confounding factors. In this paper, we introduce DepthBench, a controlled benchmark for studying computational depth across various architectures. We systematically vary the width--depth aspect ratio (d_{model}/n_{layer}) from shallow--wide to deep--narrow shapes, while keeping the model size and pre-training recipe fixed. Across 10 representative architectures, we find that the benefit of allocating more capacity to depth is strongly architecture-dependent. Standard Pre-LN and most of its norm- and scaling-based variants provide little benefit and can even degrade performance as models become deeper and narrower, whereas HC and Full AttnRes improve consistently even at extreme deep shapes. These gains extend beyond pre-training loss and consistently translate into improved domain-specific performance and effective computation. Controlled layer-level analyses further show that the gains of HC and Full AttnRes are associated with more effective utilization of additional layers, revealing distinct mechanisms of computational depth across architectures. Overall, our results identify residual connection design as a key determinant of whether depth can serve as a meaningful scaling axis by enabling additional architectural depth to translate into effective computation.

Community

Depth is a natural way to increase the computational capacity in Transformers, yet the contribution
of deeper layers can diminish as depth grows larger. Recent approaches enhance normalization (e.g.,
LayerNorm Scaling) or residual connections (e.g., mHC, AttnRes) to enable better information flow
and depth utilization. However, it remains unclear whether they truly translate increased architectural
depth into effective computational depth, and whether their reported gains stem from better access
to information across depth, or unaccounted-for confounding factors. In this paper, we introduce
DepthBench, a controlled benchmark for studying computational depth across various architectures.
We systematically vary the width–depth aspect ratio (dmodel/nlayer) from shallow–wide to deep–narrow
shapes, while keeping the model size and pre-training recipe fixed. Across 10 representative architectures,
we find that the benefit of allocating more capacity to depth is strongly architecture-dependent. Standard
Pre-LN and most of its norm- and scaling-based variants provide little benefit and can even degrade
performance as models become deeper and narrower, whereas HC and Full AttnRes improve consistently
even at extreme deep shapes. These gains extend beyond pre-training loss and consistently translate into
improved domain-specific performance and effective computation. Controlled layer-level analyses further
show that the gains of HC and Full AttnRes are associated with more effective utilization of additional
layers, revealing distinct mechanisms of computational depth across architectures. Overall, our results
identify residual connection design as a key determinant of whether depth can serve as a meaningful
scaling axis by enabling additional architectural depth to translate into effective computation.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.32534
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.32534 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.32534 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.32534 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.