Fabric prompt cache files

These files let the Fabric desktop app answer the first message to a local model in a couple of seconds. Without one, the model has to read Fabric's instructions and tool descriptions before it can start, which takes minutes on a laptop.

Each .kvprefix.bin file is a llama.cpp saved state: what one particular model file holds in memory after reading the fixed opening of Fabric's prompt, written out by one particular engine build. Fabric downloads the file together with the model, checks its SHA-256, and loads it when the engine starts.

They are no use without Fabric

A file only works with the exact model file, engine build and version of Fabric's instructions it was made from. Fabric checks all of that before loading one and ignores a file that does not match; it then reads the prompt on your computer as it always did. Loading one of these into another program, another model or another llama.cpp build will not work.

There are no model weights here. To run the models, get them from the repositories linked below.

What is here

File Size (bytes) Made for model file Engine build
qwen38-27b-ud-q4_k_xl.coding_agent.4032f47a726df60c.kvprefix.bin 1,915,317,992 Qwen3.8-27B-UD-Q4_K_XL.gguf from unsloth/Qwen3.8-27B-GGUF, SHA-256 3f227079003add2511437e5b1e94812e363385225bf6a9b47b0054a72bc8b01e, run with the draft model Qwen3.8-27B-DFlash2-Q4_K_M.gguf from z-lab/Qwen3.8-27B-DFlash2-GGUF, SHA-256 1a25c56858e1ebe93f2718ac1d49d1151f9323325c1bbfd6209370f4db131ebd Fabric engine victoria-mtp-b11512, macOS on Apple silicon
qwen35-9b-ud-q4_k_xl.coding_agent.ea5e24bf5c3cd4ad.kvprefix.bin 837,444,112 Qwen3.5-9B-UD-Q4_K_XL.gguf from unsloth/Qwen3.5-9B-GGUF, SHA-256 6f5d30666c2d8ae16a306e616d95341dcf3cc46810df84d7e6f5a7d1e4c1b293 Fabric engine victoria-mtp-b11512, macOS on Apple silicon

The sixteen characters before .kvprefix.bin in a file name are the start of that file's SHA-256. The .manifest.json beside each file records the full hash and everything the file depends on.

Each file holds the prompt for Fabric's coding assistant: 20,753 tokens for the 27B and 20,715 for the 9B.

Licences of the models

The models these files were made from are published under the Apache License 2.0, according to their own model cards:

Those cards are the authority; check them if it matters to you.

Why old files stay

New files are added whenever Fabric's instructions, its tools, a model file or the engine build changes. Older files are left in place because installed versions of Fabric still ask for them by name.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support