Instructions to use PollardWeights/clef-flash-Pollard with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use PollardWeights/clef-flash-Pollard with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf PollardWeights/clef-flash-Pollard:IQ2_XXS # Run inference directly in the terminal: llama cli -hf PollardWeights/clef-flash-Pollard:IQ2_XXS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf PollardWeights/clef-flash-Pollard:IQ2_XXS # Run inference directly in the terminal: llama cli -hf PollardWeights/clef-flash-Pollard:IQ2_XXS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf PollardWeights/clef-flash-Pollard:IQ2_XXS # Run inference directly in the terminal: ./llama-cli -hf PollardWeights/clef-flash-Pollard:IQ2_XXS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf PollardWeights/clef-flash-Pollard:IQ2_XXS # Run inference directly in the terminal: ./build/bin/llama-cli -hf PollardWeights/clef-flash-Pollard:IQ2_XXS
Use Docker
docker model run hf.co/PollardWeights/clef-flash-Pollard:IQ2_XXS
- LM Studio
- Jan
- Ollama
How to use PollardWeights/clef-flash-Pollard with Ollama:
ollama run hf.co/PollardWeights/clef-flash-Pollard:IQ2_XXS
- Unsloth Desktop
- Pi
How to use PollardWeights/clef-flash-Pollard with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PollardWeights/clef-flash-Pollard:IQ2_XXS
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "PollardWeights/clef-flash-Pollard:IQ2_XXS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use PollardWeights/clef-flash-Pollard with Docker Model Runner:
docker model run hf.co/PollardWeights/clef-flash-Pollard:IQ2_XXS
- Lemonade
How to use PollardWeights/clef-flash-Pollard with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull PollardWeights/clef-flash-Pollard:IQ2_XXS
Run and chat with the model
lemonade run user.clef-flash-Pollard-IQ2_XXS
List all available models
lemonade list
- Hermes Agent
How to use PollardWeights/clef-flash-Pollard with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PollardWeights/clef-flash-Pollard:IQ2_XXS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default PollardWeights/clef-flash-Pollard:IQ2_XXS
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use PollardWeights/clef-flash-Pollard with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PollardWeights/clef-flash-Pollard:IQ2_XXS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "PollardWeights/clef-flash-Pollard:IQ2_XXS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
clef-flash -- Pollard
Pollard shrank this model: 18.20 GB (f16) -> 3.35 GB -- 82% smaller, 5.4x down.
The smallest rung here; larger, higher-fidelity rungs are listed below.
format this model's size f16 18.20 GB Q8_0 ~9.65 GB Q6_K 7.46 GB Q4_K_M ~5.28 GB PollardMix (this repo's IQ2_XXS) 3.35 GB
Pollard builds of Cloudflare/clef-flash made with Pollard Weights -- a ladder of measured-allocation quants (bits placed by per-layer sensitivity, not a uniform crush).
Standard GGUF, but you need a recent llama.cpp. This model's architecture (clef) is implemented upstream, so any build new enough to carry it runs these files -- llama.cpp itself, and Ollama or LM Studio once they ship a runtime with it. An older build will refuse them with unknown model architecture. The quants are ordinary K-quants.
Model details
| Parameter count | ~9.1B |
| Architecture | clef |
| Input support | text, image |
| imatrix | yes -- see calibration |
| Measured | decision fidelity vs f16 through /v1/systemone -- table below (perplexity does not apply: a decision model answers with probabilities, not text) |
Which file should I choose?
Every rung is the same weights, sized to a different RAM budget by the measured allocation. Pick the largest one that fits your machine with room for context:
- ~9.5 GB RAM / VRAM ->
Uniform-Q6_K(7.46 GB). (needs a current llama.cpp) uniform allocation - ~9.5 GB RAM / VRAM ->
Q6_K(7.46 GB). (needs a current llama.cpp) measured allocation - ~7.4 GB RAM / VRAM ->
IQ4_XS(5.43 GB). (needs a current llama.cpp) measured allocation - ~7.4 GB RAM / VRAM ->
Uniform-IQ4_XS(5.40 GB). (needs a current llama.cpp) uniform allocation - ~6.6 GB RAM / VRAM ->
Uniform-IQ3_S(4.58 GB). (needs a current llama.cpp) uniform allocation - ~6.5 GB RAM / VRAM ->
IQ3_S(4.48 GB). (needs a current llama.cpp) measured allocation - ~5.4 GB RAM / VRAM ->
Uniform-IQ2_XXS(3.36 GB). (needs a current llama.cpp) uniform allocation - ~5.3 GB RAM / VRAM ->
IQ2_XXS(3.35 GB). (needs a current llama.cpp) measured allocation
Available files
| file | size | agrees with f16 | option KL | notes |
|---|---|---|---|---|
clef-flash-Pollard-IQ2_XXS.gguf |
3.35 GB | 100% | 0.030 | measured allocation |
clef-flash-Pollard-Uniform-IQ2_XXS.gguf |
3.36 GB | 100% | 0.032 | uniform allocation |
clef-flash-Pollard-IQ3_S.gguf |
4.48 GB | 100% | 0.0082 | measured allocation |
clef-flash-Pollard-Uniform-IQ3_S.gguf |
4.58 GB | 100% | 0.0033 | uniform allocation |
clef-flash-Pollard-Uniform-IQ4_XS.gguf |
5.40 GB | 100% | 0.00048 | uniform allocation |
clef-flash-Pollard-IQ4_XS.gguf |
5.43 GB | 100% | 0.0019 | measured allocation |
clef-flash-Pollard-Uniform-Q6_K.gguf |
7.46 GB | 100% | 4.5e-05 | uniform allocation |
clef-flash-Pollard-Q6_K.gguf |
7.46 GB | 100% | 6.2e-05 | measured allocation |
Decision fidelity
This is a decision model: it answers typed questions (choice / score / yes-no) with a probability per option, so it is measured on its decisions, not on text. Each rung and the f16 answered the same 20 typed questions through llama-server's /v1/systemone, read the same way (pollard-decision).
| file | size | agrees with f16 | option KL vs f16 | mean prob. drift | accuracy |
|---|---|---|---|---|---|
clef-flash-Pollard-Uniform-Q6_K.gguf |
7.46 GB | 100% | 4.5e-05 | 0.0007 | 100% |
clef-flash-Pollard-Q6_K.gguf |
7.46 GB | 100% | 6.2e-05 | 0.0009 | 100% |
clef-flash-Pollard-IQ4_XS.gguf |
5.43 GB | 100% | 0.0019 | 0.0037 | 100% |
clef-flash-Pollard-Uniform-IQ4_XS.gguf |
5.40 GB | 100% | 0.00048 | 0.0023 | 100% |
clef-flash-Pollard-Uniform-IQ3_S.gguf |
4.58 GB | 100% | 0.0033 | 0.0065 | 100% |
clef-flash-Pollard-IQ3_S.gguf |
4.48 GB | 100% | 0.0082 | 0.0114 | 100% |
clef-flash-Pollard-Uniform-IQ2_XXS.gguf |
3.36 GB | 100% | 0.032 | 0.0330 | 100% |
clef-flash-Pollard-IQ2_XXS.gguf |
3.35 GB | 100% | 0.030 | 0.0343 | 100% |
Agrees = the rung picks the same option as the f16. Option KL = how far its option probabilities moved from the f16's (0 = identical); it is the number that separates the rungs.
Two sets: measured and uniform
This repo ships the same four sizes built two ways.
clef-flash-Pollard-<rung>.gguf-- measured allocation (the Pollard build). A sensitivity profile was measured on the original weights (every one of the 32 layers, including the 24 Gated DeltaNet linear-attention layers), and bits were placed by it: the layers the model is most sensitive to keep more precision, the rest carry the compression. imatrix over the full Calib 3.0 corpus. Profile included:clef-flash-Pollard.sensitivity.json, imatrixclef-flash-Pollard.imatrix.clef-flash-Pollard-Uniform-<rung>.gguf-- uniform allocation. Same rung types, same imatrix-calibrated quants, one precision per role across all layers (no sensitivity profile; imatrix from a 200-chunk slice of the same corpus,clef-flash-Pollard-Uniform.imatrix). Shipped as a second option and as the baseline the measured set is compared against.
In both sets the decision head (dec.*, the stack that produces the answers) is held at Q6_K. Both are measured the same way in the tables above.
Multimodal
Vision needs the projector shipped alongside: mmproj-clef-flash-BF16.gguf -- download it too and pass it with --mmproj. It is not quantized; it is small and the text ladder is where the size lives.
llama-server -m clef-flash-Pollard-IQ2_XXS.gguf --mmproj mmproj-clef-flash-BF16.gguf -ngl 99
Download a specific file
pip install -U "huggingface_hub[cli]"
hf download PollardWeights/clef-flash-Pollard \
--include "clef-flash-Pollard-IQ2_XXS.gguf" --local-dir ./
How to run
This model's architecture (clef) needs a llama.cpp new enough to carry it. A decision model is served on /v1/systemone (no chat or completions):
llama-server -m clef-flash-Pollard-IQ2_XXS.gguf -ngl 99 --port 8080
curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d '{
"state": "Customer message: I was charged twice for my order last week and nobody has replied.",
"questions": {
"route": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": null, "shipping": null, "technical": null}},
"angry": {"type": "noul", "instructions": "Is the customer angry?"},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["can wait", "this week", "today", "right now"]}
}
}'
Each answer carries the probability of every option. Some decision models evaluate the whole prompt in one batch, so a long state may need a larger --ubatch-size.
imatrix (calibration)
The importance matrix (clef-flash-Pollard.imatrix, included) was computed on a mixed-domain corpus so the matrix sees every register the model serves.
ARM / AVX
llama.cpp repacks weights into an interleaved layout at load time for faster inference on ARM and AVX machines -- no special file needed, online repacking covers these quants. The old Q4_0_4_4/4_8/8_8 variants are not required.
Errata
general.architectureisclef, which upstream llama.cpp added recently. A build older than that support refuses these files withunknown model architecture-- update llama.cpp rather than looking for a different quant. Checked withpollard-ggufcheck, which reads the architecture and the tensor types out of the header and asks upstream what it implements.- Measured allocation places bits by per-layer sensitivity under a size budget.
- Single machine; replication invited.
Credits & license
- Base model:
Cloudflare/clef-flash - Quantization tooling: llama.cpp (ggml-org)
- Method + tooling: Pollard Weights -- measure first, no claim before a number.
- License:
apache-2.0, inherited from the base model.
Built with Pollard Weights -- frontier models, small hardware, no compromise.
- Downloads last month
- -
2-bit
3-bit
4-bit
6-bit