Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost
Abstract
A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out or verify (arXiv:2507.07505). We ask how much of the budget beneath that ceiling is spent on work the model has already done. Serving is stateless across requests: a model that answers a second question about a document recomputes the document's attention state from the first token. On seven real-world datasets, 98.7% of prompt tokens were text the model had already read. We present Galahad, a memory layer for vLLM, SGLang and llama.cpp that makes this reading a one-time cost. Taliesin saves the model's key-value (KV) state for a block of text and loads it on the next request that contains the same bytes, instead of recomputing it. Blaise keeps the documents themselves and passes the model only the section a question needs. On a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone let the model attend to the whole corpus and answered 98 of 100 on llama.cpp at 3.0 s and 572 J per question, against 10 of 100, 9.3 s and 2,754 J for the same model without Galahad, which could hold only the last 12,000 tokens. With Blaise added, the model read about 668 tokens per question and answered 100 of 100 on all three runtimes at 0.59-0.64 s and 200-213 J; a tuned RAGFlow pipeline answered 77. Storing the corpus is a one-time cost of about 100 s and 28 kJ, whose energy is recovered after 13 questions. Restored state is bit-identical: all 262,144 output logits matched after restart, rehydration and hot-load. Galahad worked with all 30 models we tested under vLLM, and it fails closed: any load that does not pass its checks is recomputed. Together these results move LLM serving from stateless to stateful inference.
Community
We built Galahad because we believe AI should have lasting memory that isn’t limited to what fits in a GPU’s VRAM. When models and agents repeatedly process the same context, we can end up paying for work they’ve already done. We wanted to keep that work useful.
Galahad is a persistent memory layer for AI models and agents. It saves reusable model state—the KV cache—to disk, so matching context can be restored instead of computed again, including after a restart. It also provides document memory for applications to store and retrieve relevant knowledge, with integrations for inference stacks such as vLLM and SGLang
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding (2026)
- Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs (2026)
- SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops (2026)
- Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving (2026)
- Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches (2026)
- The KV Cache Is the New Memory Wall (2026)
- TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper