CoWindow Attention: Full Causal Coverage Is a Collective Property
Abstract
FullAttn repeatedly exposes the complete causal history to every attention head, creating substantial redundant computation and memory traffic even with IO-efficient dense kernels. We introduce CoWA, a structured attention architecture that distributes access to the causal history across KV heads. All heads share near-diagonal and prefix-sink windows, while complementary long-range windows partition the remaining history. Their union provides full causal coverage although each head attends sparsely to distant tokens. This position-defined attention pattern requires no learned router or indexer, is used consistently during training and inference, and aligns with KV-head tensor parallelism. A window-matched ablation at 8K isolates the effect of complementary long-range allocation: CoWA with 100% collective coverage reaches 89.73% accuracy, compared with 89.97% for FullAttn, while duplicated long-range windows perform substantially worse. Across a broader controlled associative-recall comparison with matched token budgets, CoWA closely tracks FullAttn as the context grows, whereas other sparse patterns lose a substantial fraction of the associations. In an attention-operator benchmark at 128K tokens with tensor parallelism, CoWA reduces forward and backward latency during training by 7.4x and 8.6x and decoding latency during inference by 3.0x over FullAttn. Its per-rank peak operator memory matches FullAttn during training and is 7.6x lower during decoding. Across scaling-law training from 0.6B to 14B parameters, CoWA closely tracks FullAttn in perplexity while reducing total training FLOPs. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results show that full causal coverage can be a collective property of the head ensemble rather than a duplicated property of every head.
Community
Sharing two recent explorations in attention design from our team.
We started with two straightforward questions: Does every attention head need to repeatedly attend to the entire causal history? Once attention scores have been computed, do regions with very little contribution still need the full subsequent computation?
We explored two approaches:
CoWindow Attention (CoWA): Let heads share the work of accessing history. Heads share local context and divide distant context into complementary windows. Each head attends sparsely, while the heads collectively cover the full causal history.
📄 arxiv.org/abs/2609.32704
MassAlloc Attention (MALA): Let attention allocate its own compute. MALA preserves full causal QK scoring, then uses attention’s own softmax statistics to reduce subsequent computation in low-contribution regions.
📄 arxiv.org/abs/2609.32712
Both approaches support training forward and backward passes, as well as inference prefill and decoding. In attention-operator benchmarks at 128K tokens on 8×H100 with TP=8, compared with FullAttn:
- CoWA: 7.4× forward, 8.6× backward, and 3.0× decoding speedups.
- MALA: 2.2× forward, 3.0× backward, and 1.6× decoding speedups.
We also conducted scaling experiments from 0.6B to 14B, alongside separate continued-training experiments at 32B. During 14B training with 32K context, CoWA and MALA reduced total training FLOPs by 28.5% and 23.1%, respectively, while maintaining performance comparable to FullAttn on the evaluated model capabilities.
From method design to kernel implementation to model training, our goal was to explore which attention computations can be eliminated—and how to turn those savings into practical gains in ML infrastructure.
Get this paper in your agent:
hf papers read 2609.32704 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper