Fabric prompt cache files
These files let the Fabric desktop app answer the first message to a local model in a couple of seconds. Without one, the model has to read Fabric's instructions and tool descriptions before it can start, which takes minutes on a laptop.
Each .kvprefix.bin file is a llama.cpp saved state: what one particular model file holds in memory after reading the fixed opening of Fabric's prompt, written out by one particular engine build. Fabric downloads the file together with the model, checks its SHA-256, and loads it when the engine starts.
They are no use without Fabric
A file only works with the exact model file, engine build and version of Fabric's instructions it was made from. Fabric checks all of that before loading one and ignores a file that does not match; it then reads the prompt on your computer as it always did. Loading one of these into another program, another model or another llama.cpp build will not work.
There are no model weights here. To run the models, get them from the repositories linked below.
What is here
| File | Size (bytes) | Made for model file | Engine build |
|---|---|---|---|
qwen38-27b-ud-q4_k_xl.coding_agent.4032f47a726df60c.kvprefix.bin |
1,915,317,992 | Qwen3.8-27B-UD-Q4_K_XL.gguf from unsloth/Qwen3.8-27B-GGUF, SHA-256 3f227079003add2511437e5b1e94812e363385225bf6a9b47b0054a72bc8b01e, run with the draft model Qwen3.8-27B-DFlash2-Q4_K_M.gguf from z-lab/Qwen3.8-27B-DFlash2-GGUF, SHA-256 1a25c56858e1ebe93f2718ac1d49d1151f9323325c1bbfd6209370f4db131ebd |
Fabric engine victoria-mtp-b11512, macOS on Apple silicon |
qwen35-9b-ud-q4_k_xl.coding_agent.ea5e24bf5c3cd4ad.kvprefix.bin |
837,444,112 | Qwen3.5-9B-UD-Q4_K_XL.gguf from unsloth/Qwen3.5-9B-GGUF, SHA-256 6f5d30666c2d8ae16a306e616d95341dcf3cc46810df84d7e6f5a7d1e4c1b293 |
Fabric engine victoria-mtp-b11512, macOS on Apple silicon |
The sixteen characters before .kvprefix.bin in a file name are the start of that file's SHA-256. The .manifest.json beside each file records the full hash and everything the file depends on.
Each file holds the prompt for Fabric's coding assistant: 20,753 tokens for the 27B and 20,715 for the 9B.
Licences of the models
The models these files were made from are published under the Apache License 2.0, according to their own model cards:
- Qwen/Qwen3.8-27B and the build used here, unsloth/Qwen3.8-27B-GGUF:
apache-2.0 - z-lab/Qwen3.8-27B-DFlash2-GGUF:
apache-2.0 - Qwen/Qwen3.5-9B and the build used here, unsloth/Qwen3.5-9B-GGUF:
apache-2.0, with the licence text at Qwen/Qwen3.5-9B/LICENSE
Those cards are the authority; check them if it matters to you.
Why old files stay
New files are added whenever Fabric's instructions, its tools, a model file or the engine build changes. Older files are left in place because installed versions of Fabric still ask for them by name.