Papers
arxiv:2609.33203

Structured Residual Connectivity Matters for Diffusion Transformers

Published on Sep 27
· Submitted by
liu
on Sep 29
Authors:
,
,
,
,

Abstract

Diffusion Transformers (DiTs) have established themselves as a scalable backbone for high-fidelity image synthesis. However, unlike U-Net based diffusion models that rely on rigid, hand-crafted skip connections, DiTs predominantly use a uniform residual stream that integrates all preceding layers as a monolithic state. In this work, we rethink residual connections in diffusion transformers and propose to transform them from passive summation into an active retrieval mechanism optimized for image denoising. First, we conduct a systematic analysis of DiT's internal representation, revealing a latent preference for early-layer feature reuse and symmetric layer guidance. Motivated by this, we introduce a structured connectivity design that explicitly integrates local residual connections with long-range pathways. Instead of static skip connections or dense all-layer routing, our method enables each transformer block to selectively ``attend'' to critical earlier representations, dynamically retrieving spatial and semantic cues through direct, differentiable cross-depth paths. Experiments show that our adaptive connectivity leads to faster convergence, with up to 1.73times fewer training iterations, and significant gains in FID and visual quality with less than 0.1% additional parameters, further improving a strong REPA-XL/2 model from 5.9 to 4.34 FID without guidance and reaching 1.39 FID with classifier-free guidance. Our findings suggest that adaptive cross-layer connectivity is a critical yet underexplored factor in diffusion transformers, and that incorporating structured information pathways provides a simple and effective direction for improving scalable generative models.

Community

Diffusion Transformers accumulate all preceding layers into a uniform residual stream. This homogenizes representations across depth and fixes how gradients flow. We turn this passive summation into active retrieval. Encoder-side representations are kept as differentiable sources, and each decoder sublayer adaptively fuses its mirrored encoder source using lightweight softmax routing.

  • Analysis: we find that, when allowed to route across layers, DiTs rely on early-layer representations and spontaneously favor symmetric layer pairs.
  • Results: with <0.1% extra parameters, the method gives consistent gains across DiT-S/B/XL on ImageNet. It also revives an already saturated REPA-XL/2 checkpoint. The baseline gains only 0.5 FID from 1M to 4M iterations, and continued fine-tuning brings no further improvement (5.9 → 5.87). Adding our routing drops FID to 4.94 in just 0.28M steps and to 4.34 at 0.35M steps. With CFG, it further reaches 1.39 FID.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.33203
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.33203 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.33203 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.33203 in a Space README.md to link it from this page.

Collections including this paper 1