--- license: apache-2.0 base_model: Qwen/Qwen3-4B tags: - kv-cache - context-compression - long-context - research - negative-results - reproducibility library_name: safetensors --- # context-cache-lab — Stage 2b memory extractor (research artifact) > ## ⚠️ This is a negative-result research artifact, not a useful model. > > It is published so that a **failed** experiment is reproducible. Its own > preregistered gates returned **INCONCLUSIVE** (Stage 2b) and **FAIL** (Stage 3). > Do not use it for anything in production. Do not read a capability claim into it. A 566M-parameter C²KV-style sidecar that compiles chunks of text into compressed, pre-RoPE KV "pages" for a **frozen** `Qwen/Qwen3-4B`. It does not generate text and cannot be used standalone — it only produces key/value state that the frozen target consumes. **Code, protocol, preregistration and all results:** https://github.com/johnathonkillaly/context-cache-lab ## What was measured | Stage | Verdict | Result | |---|---|---| | 2b — does learned compressed state beat equal-budget raw text? | 🟡 **INCONCLUSIVE** | Beat `BUDGET` at every ratio, but missed 2 of 4 frozen criteria by ~1% | | 3 — do independently compiled pages support two-page reasoning? | ❌ **FAIL** | **0/13** valid items vs NATIVE **13/13**; all 5 gates missed | **Stage 2b (single-document QA, 4096 tokens, `clean_hit`):** | ratio | this extractor | equal-budget raw text | NATIVE | |---|---|---|---| | 2× | 0.569 | 0.472 | 0.958 | | 4× | 0.389 | 0.222 | 0.958 | | 8× | 0.347 | 0.125 | 0.958 | | 16× | 0.278 | 0.069 | 0.958 | Semantic facts survived compression (0.625 at 4×) while exact strings did not — hashes collapsed to **0.000** and identifiers to 0.250, where raw text at the same budget scored 0.750. Warm reuse was 23–32× faster than native prefill with breakeven at 2 queries, but compilation costs more than a single prefill, so the first query is slower. **Stage 3** then evaluated this exact frozen checkpoint on a preregistered cross-page compositional task. It scored **zero** clean accuracy on every native-valid item, with a mean rank margin (−2.47) *worse* than no context at all (−1.10). Compiling the document jointly instead of per-page also scored zero, so this is not an independence tax. **What that failure does not establish:** Stage 3 used a different corpus (~94–100 token pages against ~256-token training chunks), so it measures insufficient transfer of *this* carrier to *that* task — not the impossibility of independently compiled memory in general. It also cannot fully separate loss of within-page fidelity from failure to combine intact relations. See the repository's Stage 3 report for the full limitations. ## Files | File | Purpose | |---|---| | `stage2b_extractor.safetensors` | **Use this.** 109 fp32 tensors, safe to load. Verified to round-trip exactly against the original. | | `stage2b_extractor.pt` | Provenance artifact only. Its SHA-256 is pinned in the repo's `results/stage3/freeze.json`, so the Stage 3 integrity audit reproduces exactly. **It is a pickle** — the repo's scripts load it with `weights_only=False`. Only use it if you need the hash to verify. | | `config.json` | `layer_share`, `n_sink`, training step, target model revision, and the original `.pt` SHA-256. | ``` original .pt sha256: b8384dd2eeaa87890a6b65e22f791183399764b93dd9a7f4e8a6e4ce13fca425 target model revision: 1cfa9a7208912126459214e8b04321603b3df60c training step: 2200 ``` ## Usage The extractor is only meaningful alongside the repository's code: ```bash git clone https://github.com/johnathonkillaly/context-cache-lab cd context-cache-lab uv venv --python 3.12 .venv && uv pip install --python .venv/bin/python -e ".[dev]" ``` ```python from huggingface_hub import hf_hub_download from safetensors.torch import load_file from ccl.compressor import MemoryExtractor from ccl.target import TargetModel target = TargetModel("Qwen/Qwen3-4B") # frozen extractor = MemoryExtractor(target, layer_share=1, n_sink=8) extractor.load_state_dict(load_file(hf_hub_download( "jkillay/context-cache-lab-stage2b-extractor", "stage2b_extractor.safetensors"))) extractor.eval() pages = extractor.compile_chunks([target.encode(c) for c in chunks], ratio=4.0) ``` Requires `transformers>=5.0` (the legacy tuple KV-cache format is gone) and roughly 10 GB of memory for the frozen target plus the sidecar. Developed on Apple Silicon (MPS); no CUDA is assumed. ## Training 2,200 steps, ~108 minutes on an Apple M4 Max, compression ratio sampled per example (C²KV's `-Dyn` variant). The target model was frozen throughout — only the shared memory-token embedding and per-layer Q/K/V projection heads received gradients. Supervision was on answer tokens only, applied **after** page concatenation, so the extractor is pushed toward states that compose rather than states that are merely individually informative. Training data was a **synthetic** corpus generated deterministically from integer seeds, with value pools provably disjoint from the held-out evaluation draw. No scraped text and no personal data. ## License and attribution Apache-2.0. This is a derivative of [`Qwen/Qwen3-4B`](https://huggingface.co/Qwen/Qwen3-4B) (Apache-2.0) in the sense that its projections were initialized from that model's weights; **no Qwen weights are redistributed here** — only the trained sidecar. The mechanism is a reimplementation of the published design of [C²KV (Du et al., KDD 2026)](https://arxiv.org/abs/2607.17715). No C²KV source code was copied. Please cite that paper alongside this artifact. ## Citation See `CITATION.cff` in the [GitHub repository](https://github.com/johnathonkillaly/context-cache-lab).