Papers
arxiv:2609.36322

Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

Published on Sep 28
· Submitted by
Xingyu Zhu
on Sep 30
Authors:
,
,
,
,
,
,

Abstract

Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. We uncover a systematic asymmetry in models using such compression: the same information can be easy to retrieve at one phase and difficult at another. We call this periodic variation in retrieval performance phase sensitivity. In large open-weight models with such compression, long-context retrieval accuracy can differ by up to 40 percentage points across phases, revealing periodic weak spots that average benchmark scores can conceal. To investigate this behavior, we pretrain a family of transformers from scratch across multiple KV-compression designs, reproducing phase sensitivity across the variants. Mechanistic analysis using causal interventions in these models reveals phase specialization: different attention components contribute asymmetrically to retrieving information at different source phases. We further analyze idealized retrieval models, showing how gradient flow dynamics may favor sharp phase specialization. Evaluating models with chunked KV-cache compression thus requires measuring across compression phases: high average accuracy can coexist with systematic positional failures.

Community

Paper author Paper submitter

Models with chunked KV-cache compression retrieve worse at some token positions than others, on a cycle set by the compression stride. We call this phase sensitivity, and it is large enough to flip answers and move long-context accuracy by tens of points.

01-headline

A one-token shift changes the answer. We ask DeepSeek-V4-Flash-Base to complete T.Cast(FP in its own FP8 kernel and lengthen an unrelated docstring by one decorative "=" at a time. The code never changes, yet the top prediction switches between the correct 8 and 32 every two tokens, a cycle of four. Across four families of filler text with 16 lengths each, 60 of 64 rankings follow this pattern. DeepSeek-V3.1-Base, which has no chunked compression, prefers 32 on only 4 of the same 64 inputs.

Where the cycle comes from. DeepSeek-V4 does not keep every past token in its cache. It compresses windows of 8 tokens that start every 4 tokens into one entry each, so every token has a phase: its position modulo the stride. Prepending a token changes neither the content nor the order of the text, only which tokens share an entry. Any dependence on phase is therefore a property of the model, not of the data.

02-phase

Up to 40 points at 128K tokens. To measure the effect, we built a needle-in-a-haystack test in which prompt groups differ only in phase: 16,000 key–value records per prompt, the same query position, and the same mean and variance of the target position in every group, with 256 prompts per group and model. All four DeepSeek-V4 models, base and post-trained, rise and fall with period 4. The gap between the best and worst group is 15 to 40 points, while the two periods agree to within 1.1 to 2.0 points. Post-training raises accuracy and narrows the gap, but it keeps the period.

03-deepseek-v4

DeepSeek-V4.1-Flash compresses every 2 tokens, and its cycle shortens accordingly: every even position scores 94.7 to 95.3%, every odd one 89.2 to 90.3% (2,560 prompts per group).

04-deepseek-v41

It is the compression. DeepSeek-V4 differs from a standard transformer in many ways, so we pretrained 0.6B models (Qwen3-0.6B architecture) from scratch on 100B tokens, with chunked compression as the only substantial change from full-attention baselines. We swept the window, stride, number of KV heads, gating and positional encoding. The period always follows the stride: strides of 4, 6, 8 and 12 give periods of about 4, 6, 8 and 12. Full attention stays within 6.1 points across positions, while compressed models reach gaps of up to 78 points, even without RoPE or with uniform averaging in place of learned gates. Averages hide this. With window 8, stride 8 and 8 KV heads, the compressed model's mean accuracy (59.4%) is close to full attention's (61.1%), but its worst position scores 9.9% against 58.4%.

05-controlled

Heads specialize to phases. In a model with one KV head per layer, we knock out one head at a time by replacing its output with its mean over calibration prompts. Some heads matter at only two or three adjacent phases: layer 9 at phases 5 and 6, layer 10 at 3 and 4, layer 14 at 1 to 3. Different heads cover different phases. In DeepSeek-V4-Flash-Base, knocking out whole layers shows losses that repeat with period 4.

06-knockouts

The compression gates of these heads favor the same slots in every window, and each value gate peaks about one slot after its key gate, so a head keeps a token together with its successor. If the gates set the pattern, rotating them should move it. Cyclically shifting the gate parameters by one slot, with all other weights unchanged, moves the weak spots by one, as predicted (R² = 0.98 for the model shown; 0.66 to 1.00 across five models and shifts of up to three slots). The boundary phase, where a window ends, stays weak under every rotation.

07-gate-shift

Why training produces it. In an idealized induction model, we prove that concentrating each head on a single slot is optimal, and that gradient flow settles each head on a fixed slot, the same for every input. Nothing in the dynamics pushes heads toward different slots, so some phases can remain uncovered.

Takeaway. Models with compressed memory can look fine on average while hiding periodic weak spots. We suggest reporting long-context accuracy by phase alongside the mean.

Paper: https://arxiv.org/abs/2609.36322
Interactive blog post, where you can shift the phase, pick models and explore every result: https://ultimatejupiter.github.io/blog/periodic-weak-spots/

Interesting failure mode!

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.36322 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.36322 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.36322 in a Space README.md to link it from this page.

Collections including this paper 1