clef-flash -- Pollard

Pollard shrank this model: 18.20 GB (f16) -> 3.35 GB -- 82% smaller, 5.4x down.

The smallest rung here; larger, higher-fidelity rungs are listed below.

format this model's size
f16 18.20 GB
Q8_0 ~9.65 GB
Q6_K 7.46 GB
Q4_K_M ~5.28 GB
PollardMix (this repo's IQ2_XXS) 3.35 GB

Pollard builds of Cloudflare/clef-flash made with Pollard Weights -- a ladder of measured-allocation quants (bits placed by per-layer sensitivity, not a uniform crush).

Standard GGUF, but you need a recent llama.cpp. This model's architecture (clef) is implemented upstream, so any build new enough to carry it runs these files -- llama.cpp itself, and Ollama or LM Studio once they ship a runtime with it. An older build will refuse them with unknown model architecture. The quants are ordinary K-quants.

Model details

Parameter count ~9.1B
Architecture clef
Input support text, image
imatrix yes -- see calibration
Measured decision fidelity vs f16 through /v1/systemone -- table below (perplexity does not apply: a decision model answers with probabilities, not text)

Which file should I choose?

Every rung is the same weights, sized to a different RAM budget by the measured allocation. Pick the largest one that fits your machine with room for context:

  • ~9.5 GB RAM / VRAM -> Uniform-Q6_K (7.46 GB). (needs a current llama.cpp) uniform allocation
  • ~9.5 GB RAM / VRAM -> Q6_K (7.46 GB). (needs a current llama.cpp) measured allocation
  • ~7.4 GB RAM / VRAM -> IQ4_XS (5.43 GB). (needs a current llama.cpp) measured allocation
  • ~7.4 GB RAM / VRAM -> Uniform-IQ4_XS (5.40 GB). (needs a current llama.cpp) uniform allocation
  • ~6.6 GB RAM / VRAM -> Uniform-IQ3_S (4.58 GB). (needs a current llama.cpp) uniform allocation
  • ~6.5 GB RAM / VRAM -> IQ3_S (4.48 GB). (needs a current llama.cpp) measured allocation
  • ~5.4 GB RAM / VRAM -> Uniform-IQ2_XXS (3.36 GB). (needs a current llama.cpp) uniform allocation
  • ~5.3 GB RAM / VRAM -> IQ2_XXS (3.35 GB). (needs a current llama.cpp) measured allocation

Available files

file size agrees with f16 option KL notes
clef-flash-Pollard-IQ2_XXS.gguf 3.35 GB 100% 0.030 measured allocation
clef-flash-Pollard-Uniform-IQ2_XXS.gguf 3.36 GB 100% 0.032 uniform allocation
clef-flash-Pollard-IQ3_S.gguf 4.48 GB 100% 0.0082 measured allocation
clef-flash-Pollard-Uniform-IQ3_S.gguf 4.58 GB 100% 0.0033 uniform allocation
clef-flash-Pollard-Uniform-IQ4_XS.gguf 5.40 GB 100% 0.00048 uniform allocation
clef-flash-Pollard-IQ4_XS.gguf 5.43 GB 100% 0.0019 measured allocation
clef-flash-Pollard-Uniform-Q6_K.gguf 7.46 GB 100% 4.5e-05 uniform allocation
clef-flash-Pollard-Q6_K.gguf 7.46 GB 100% 6.2e-05 measured allocation

Decision fidelity

This is a decision model: it answers typed questions (choice / score / yes-no) with a probability per option, so it is measured on its decisions, not on text. Each rung and the f16 answered the same 20 typed questions through llama-server's /v1/systemone, read the same way (pollard-decision).

file size agrees with f16 option KL vs f16 mean prob. drift accuracy
clef-flash-Pollard-Uniform-Q6_K.gguf 7.46 GB 100% 4.5e-05 0.0007 100%
clef-flash-Pollard-Q6_K.gguf 7.46 GB 100% 6.2e-05 0.0009 100%
clef-flash-Pollard-IQ4_XS.gguf 5.43 GB 100% 0.0019 0.0037 100%
clef-flash-Pollard-Uniform-IQ4_XS.gguf 5.40 GB 100% 0.00048 0.0023 100%
clef-flash-Pollard-Uniform-IQ3_S.gguf 4.58 GB 100% 0.0033 0.0065 100%
clef-flash-Pollard-IQ3_S.gguf 4.48 GB 100% 0.0082 0.0114 100%
clef-flash-Pollard-Uniform-IQ2_XXS.gguf 3.36 GB 100% 0.032 0.0330 100%
clef-flash-Pollard-IQ2_XXS.gguf 3.35 GB 100% 0.030 0.0343 100%

Agrees = the rung picks the same option as the f16. Option KL = how far its option probabilities moved from the f16's (0 = identical); it is the number that separates the rungs.

Two sets: measured and uniform

This repo ships the same four sizes built two ways.

  • clef-flash-Pollard-<rung>.gguf -- measured allocation (the Pollard build). A sensitivity profile was measured on the original weights (every one of the 32 layers, including the 24 Gated DeltaNet linear-attention layers), and bits were placed by it: the layers the model is most sensitive to keep more precision, the rest carry the compression. imatrix over the full Calib 3.0 corpus. Profile included: clef-flash-Pollard.sensitivity.json, imatrix clef-flash-Pollard.imatrix.
  • clef-flash-Pollard-Uniform-<rung>.gguf -- uniform allocation. Same rung types, same imatrix-calibrated quants, one precision per role across all layers (no sensitivity profile; imatrix from a 200-chunk slice of the same corpus, clef-flash-Pollard-Uniform.imatrix). Shipped as a second option and as the baseline the measured set is compared against.

In both sets the decision head (dec.*, the stack that produces the answers) is held at Q6_K. Both are measured the same way in the tables above.

Multimodal

Vision needs the projector shipped alongside: mmproj-clef-flash-BF16.gguf -- download it too and pass it with --mmproj. It is not quantized; it is small and the text ladder is where the size lives.

llama-server -m clef-flash-Pollard-IQ2_XXS.gguf --mmproj mmproj-clef-flash-BF16.gguf -ngl 99

Download a specific file

pip install -U "huggingface_hub[cli]"
hf download PollardWeights/clef-flash-Pollard \
  --include "clef-flash-Pollard-IQ2_XXS.gguf" --local-dir ./

How to run

This model's architecture (clef) needs a llama.cpp new enough to carry it. A decision model is served on /v1/systemone (no chat or completions):

llama-server -m clef-flash-Pollard-IQ2_XXS.gguf -ngl 99 --port 8080
curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d '{
  "state": "Customer message: I was charged twice for my order last week and nobody has replied.",
  "questions": {
    "route":   {"type": "choice", "instructions": "Which team should handle this?",
                "criteria": {"billing": null, "shipping": null, "technical": null}},
    "angry":   {"type": "noul",   "instructions": "Is the customer angry?"},
    "urgency": {"type": "score",  "instructions": "How urgent is this?",
                "criteria": ["can wait", "this week", "today", "right now"]}
  }
}'

Each answer carries the probability of every option. Some decision models evaluate the whole prompt in one batch, so a long state may need a larger --ubatch-size.

imatrix (calibration)

The importance matrix (clef-flash-Pollard.imatrix, included) was computed on a mixed-domain corpus so the matrix sees every register the model serves.

ARM / AVX

llama.cpp repacks weights into an interleaved layout at load time for faster inference on ARM and AVX machines -- no special file needed, online repacking covers these quants. The old Q4_0_4_4/4_8/8_8 variants are not required.

Errata

  • general.architecture is clef, which upstream llama.cpp added recently. A build older than that support refuses these files with unknown model architecture -- update llama.cpp rather than looking for a different quant. Checked with pollard-ggufcheck, which reads the architecture and the tensor types out of the header and asks upstream what it implements.
  • Measured allocation places bits by per-layer sensitivity under a size budget.
  • Single machine; replication invited.

Credits & license

Built with Pollard Weights -- frontier models, small hardware, no compromise.

Downloads last month
-
GGUF
Model size
9B params
Architecture
clef
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PollardWeights/clef-flash-Pollard

Finetuned
Qwen/Qwen3.5-9B
Quantized
(45)
this model