cryptobiosis's picture
Add 'When to pick KakeyaLattice over HQQ / Quanto / KIVI' comparison block
b044c1d verified
|
Raw
History Blame Contribute Delete
3.96 kB
---
title: KakeyaLattice KV-cache compression
emoji: 📐
colorFrom: indigo
colorTo: purple
sdk: docker
app_port: 7860
pinned: false
license: apache-2.0
---
# KakeyaLattice KV-cache compression
Side-by-side comparison of **bf16 DynamicCache** vs **KakeyaLattice E8**
compression at three quality levels (Q=10 aggressive, Q=38 balanced,
Q=152 near-lossless) on a small HuggingFace causal LM.
Default model: `Qwen/Qwen3-0.6B` (head_dim=128, GQA 16/8 — the same
attention shape as modern production LLMs, so the codec numbers are
representative). Runs on the free CPU tier (each "Run comparison"
click takes ~4–8 minutes on 2 cores). Override `KAKEYA_DEMO_MODEL`
env var to use a larger model on a GPU Space (`Qwen/Qwen3-1.7B`,
`Qwen/Qwen3-4B`).
## How it works
`KakeyaLatticeCache` is a drop-in subclass of `transformers.DynamicCache`
that applies a Zamir-Feder nested-lattice codec roundtrip (encode +
decode) to every K and V written into the cache.
```python
from transformers import AutoModelForCausalLM
from kakeyalattice.hf import KakeyaLatticeCache
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-0.6B")
cache = KakeyaLatticeCache(
variant="e8", q_range=38,
num_hidden_layers=model.config.num_hidden_layers,
head_dim=model.config.head_dim,
)
out = model.generate(input_ids, max_new_tokens=200, past_key_values=cache)
```
## What you'll see in the demo
For each prompt, the app generates four times (bits/vec here assume
head_dim=128 → bf16 baseline is 2048 bits/vec; exact numbers for other
head_dims scale proportionally):
| config | bits/vec (head_dim=128) | expected quality |
| ------------------------ | ----------------------- | --------------------------------- |
| bf16 DynamicCache | 2048 (reference) | identical to reference |
| E8 Q=152 near-lossless | ~1920 (-6%) | essentially identical |
| E8 Q=38 balanced | ~880 (-57%) | ~1% deviation in ppl |
| E8 Q=10 aggressive | ~640 (-69%) | noticeably different but coherent |
(The percentage savings `-6% / -57% / -69%` are what matter — they are
fixed by the E8 codec design and do not depend on head_dim.)
Wall-clock latency per config is also reported.
## When to pick KakeyaLattice over HQQ / Quanto / KIVI
- **HQQ / AWQ / GPTQ** are *weight* quantisers. KakeyaLattice is a
*KV-cache* quantiser. They are **orthogonal** — stack them.
- **QuantoQuantizedCache / HQQQuantizedCache** in transformers are
per-channel scalar quantisers. At ≤ 1 % |Δppl| KakeyaLattice
compresses the KV cache **9 %–38 % harder** across Qwen3-4B,
GLM-4-9B-Chat, Gemma-4-E4B, and DeepSeek-R1-Distill-Qwen-1.5B
(real vLLM, H200, 128 k context, WikiText-103 n=8; see the
[GitHub README](https://github.com/FluffyAIcode/LLM-KV--Cache-compress#headline-numbers)
for the full table).
- **KIVI (2-bit KV)** hits similar bit budgets but cannot gaussianise
heavy-tailed KV distributions; KakeyaLattice's Sylvester–Hadamard
rotation does, giving lower |Δppl| at matched bits.
- **SnapKV / H2O / Scissorhands** are *eviction* (which KV to keep),
not *quantisation* (how to store). They compose multiplicatively
with KakeyaLattice.
## Caveats
- The cache roundtrips K/V but stores the reconstructed tensor in the
model's KV dtype. Real HBM bytes saved are **nominal** — the demo's
value is showing reconstruction quality, not memory savings.
- Decode is ~1.3-2× slower than bf16 because the codec runs as pure
PyTorch ops. A fused Triton kernel would close this gap.
- Head-dim must be a power of 2 and divisible by 4 (D4) or 8 (E8).
Most modern LLMs satisfy this.
## Links
- Package: https://pypi.org/project/kakeyalattice/
- Repo: https://github.com/FluffyAIcode/LLM-KV--Cache-compress
- Paper: `reports/paper/`
- DeepSeek-V4-Flash Stage 0.75 findings: `reports/v1_5_release/dsv4_stage075/FINDINGS.md`