Text Generation
GGUF
English
Chinese
pollard
llama.cpp
Mixture of Experts
bailingmoe3
measured-sensitivity
imatrix
conversational
Instructions to use PollardWeights/Ling-3.0-tiny-Pollard with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use PollardWeights/Ling-3.0-tiny-Pollard with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S # Run inference directly in the terminal: llama cli -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S # Run inference directly in the terminal: llama cli -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S # Run inference directly in the terminal: ./llama-cli -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Use Docker
docker model run hf.co/PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
- LM Studio
- Jan
- vLLM
How to use PollardWeights/Ling-3.0-tiny-Pollard with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PollardWeights/Ling-3.0-tiny-Pollard" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PollardWeights/Ling-3.0-tiny-Pollard", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
- Ollama
How to use PollardWeights/Ling-3.0-tiny-Pollard with Ollama:
ollama run hf.co/PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
- Unsloth Studio
How to use PollardWeights/Ling-3.0-tiny-Pollard with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for PollardWeights/Ling-3.0-tiny-Pollard to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for PollardWeights/Ling-3.0-tiny-Pollard to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for PollardWeights/Ling-3.0-tiny-Pollard to start chatting
- Pi
How to use PollardWeights/Ling-3.0-tiny-Pollard with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use PollardWeights/Ling-3.0-tiny-Pollard with Docker Model Runner:
docker model run hf.co/PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
- Lemonade
How to use PollardWeights/Ling-3.0-tiny-Pollard with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Run and chat with the model
lemonade run user.Ling-3.0-tiny-Pollard-IQ3_S
List all available models
lemonade list
- Hermes Agent
How to use PollardWeights/Ling-3.0-tiny-Pollard with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use PollardWeights/Ling-3.0-tiny-Pollard with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 8,341 Bytes
5148d15 8f61394 5148d15 8f61394 5148d15 8f61394 5148d15 8f61394 5148d15 afc43dc 5148d15 afc43dc 5148d15 8f61394 5148d15 8f61394 5148d15 afc43dc 5148d15 8f61394 cfe65d6 8f61394 afc43dc 5148d15 afc43dc 8f61394 afc43dc 5148d15 8f61394 5148d15 afc43dc 9c06103 74c35dc 9c06103 74c35dc 9c06103 74c35dc 9c06103 74c35dc 9c06103 afc43dc 8f61394 5148d15 17b9c3d 5148d15 afc43dc 5148d15 8f61394 afc43dc 74c35dc 5148d15 8f61394 afc43dc 74c35dc afc43dc 8f61394 74c35dc 8f61394 afc43dc 8f61394 afc43dc 8f61394 afc43dc 8f61394 afc43dc 8f61394 5148d15 afc43dc 8f61394 5148d15 8f61394 afc43dc 8f61394 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 | ---
quantized_by: PollardWeights
pipeline_tag: text-generation
base_model: inclusionAI/Ling-3.0-tiny
base_model_relation: quantized
license: mit
language:
- en
- zh
tags:
- pollard
- gguf
- llama.cpp
- moe
- bailingmoe3
- measured-sensitivity
- imatrix
- conversational
---
# Pollard measured-sensitivity quantizations of Ling-3.0-tiny by inclusionAI
Built with **[Pollard Weights](https://github.com/WestWaters/pollard-weights)** on
**llama.cpp** build `b10360` (`48d22e295`) β the first build line with `bailingmoe3`
support ([PR #26608](https://github.com/ggml-org/llama.cpp/pull/26608), merged
2026β08β17). Use that build or newer to run these.
Original model: https://huggingface.co/inclusionAI/Ling-3.0-tiny
## Model details
| | |
|---|---|
| Parameter count | ~7.9B total / ~1.7B active (MoE) β listed as 8B |
| Architecture | `bailingmoe3` (128 experts/layer, topβ8 + 1 shared, 24 layers) |
| Input support | text |
| Speculative decoding | no |
| imatrix | **yes** β [details below](#imatrix-calibration), corpus + matrix included in this repo |
| Perplexity / KLD measured | **yes** β this is the whole point (see next section) |
Uniform quants spend the same bits on every layer. Pollard **measures** how much
crushing each tensor group actually costs β KL-divergence, per layer β then a
KL-aware knapsack spends bits where they matter: more on the sensitive layers,
fewer on the ones that don't care. Same weights, smarter bit allocation.
## Why this over a uniform quant
Held-out KL-divergence vs a Q6_K reference (lower = closer to the full model),
measured on the same held-out set for every build:
| build | size | mean KL | vs uniform |
|---|---|---|---|
| **Ling-3.0-tiny Pollard** | **3.83 GB** | **0.1875** | _baseline_ |
| uniform IQ3 (interpolated to 3.83 GB) | 3.83 GB | β 0.204 | **β 8% higher KL** |
| uniform IQ3_S | 3.51 GB | 0.2821 | reference points |
| uniform IQ3_M | 3.56 GB | 0.2469 | (bracket the curve) |
| uniform IQ4_XS | 4.29 GB | 0.1312 | (bracket the curve) |
At matched size the measured allocation sits **below** the uniform sizeβKL curve.
The measured mix: sensitive early layers get `iq4_xs`, most get `iq3_s`, the
least-sensitive get `iq2_s`; every attention block stays `q6_K`/`q5_K`;
embeddings/output stay `q6_K`; imatrix-uncovered MoE tensors are pinned so the
aggressive base can't crash. (ffn sensitivity spread ~6Γ, attn spread ~16Γ across
the 24 layers β that variance is exactly what a uniform quant wastes. The full
per-tensor map is in [`Ling-3.0-tiny-Pollard.tensor-types.txt`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.tensor-types.txt).)
## Prompt format
```
<role>SYSTEM</role>{system_prompt}
detailed thinking on<|role_end|><role>HUMAN</role>{prompt}<|role_end|><role>ASSISTANT</role>
<think>
```
## Which file should I choose?
Pick the rung for your machine β each is the **same weights**, sized to a different
RAM budget by the measured allocation:
- **~8 GB RAM / VRAM** β **`IQ3_S`** (3.83 GB). The value pick: full model with room
for context, and it beats same-size uniform IQ3 (table above). **Recommended.**
- **~9 GB** β **`IQ4_XS`** (4.64 GB). More fidelity β the sensitive layers move up to
`iq4_xs`.
- **~11 GB** β **`Q6_K`** (6.26 GB). Near-lossless; as close to the full model as a
quant gets.
- Want it even smaller than IQ3_S? Pollard *loses* to uniform at the extreme IQ2 floor
for this model (the weights are too crushed for reallocation to help), so we don't
ship one β *measure first, no claim before a number.*
## Available files
MoE speed: only ~1.7B of the 7.9B params are active per token, so even the big rungs
stay fast on an **Apple M4** (`tg`, llama.cpp Metal).
| Filename | Type | Size | M4 tok/s | Description |
|---|---|---|---|---|
| [Ling-3.0-tiny-Pollard-IQ3_S.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-IQ3_S.gguf) | IQ3 measured mix (IQ2_SβIQ4_XS, q6_K embed/attn) | 3.83 GB | **75.1** | Fits an ~8 GB box. Beats same-size uniform IQ3 (table above). **Recommended.** |
| [Ling-3.0-tiny-Pollard-IQ4_XS.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-IQ4_XS.gguf) | IQ4_XS measured mix (q6_K/q5_K attn, q6_K embed) | 4.64 GB | **75.2** | Fits an ~9 GB box. Higher fidelity β sensitive layers pushed to iq4_xs. |
| [Ling-3.0-tiny-Pollard-Q6_K.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-Q6_K.gguf) | Q5/Q6 measured mix (18L q6_K, 6L q5_K) | 6.26 GB | **66.9** | Fits an ~11 GB box. Near-lossless β maximum quality. |
| [Ling-3.0-tiny-Pollard.imatrix](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.imatrix) | importance matrix | 44 MB | β | The imatrix used, for anyone re-quantizing. |
| [Ling-3.0-tiny-Pollard-calibration.txt](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-calibration.txt) | calibration corpus | ~1 MB | β | The exact corpus the imatrix was computed on. |
| [Ling-3.0-tiny-Pollard.tensor-types.txt](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.tensor-types.txt) | allocation map | 3 KB | β | The measured per-tensor bit assignment. |
## Download a specific file
```bash
pip install -U "huggingface_hub[cli]"
hf download PollardWeights/Ling-3.0-tiny-Pollard \
--include "Ling-3.0-tiny-Pollard-IQ3_S.gguf" --local-dir ./
```
## How to run
These are standard GGUF and run with **llama.cpp** β one-line install:
```bash
curl -LsSf https://llama.app/install.sh | sh
llama-server -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
```
or with a local file:
```bash
llama-cli -m Ling-3.0-tiny-Pollard-IQ3_S.gguf -ngl 99 -p "Explain MoE routing simply."
llama-server -m Ling-3.0-tiny-Pollard-IQ3_S.gguf -ngl 99 # OpenAI-compatible API + web UI at :8080
```
They also work in anything built on llama.cpp β **LM Studio, koboldcpp, ramalama,
Jan, Text Generation WebUI, LoLLMs** β provided the build is recent enough to carry
`bailingmoe3` support (see top). If the app ships an older llama.cpp, update it first.
## imatrix (calibration)
The importance matrix ([`Ling-3.0-tiny-Pollard.imatrix`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.imatrix),
included) was computed on a **mixed-domain corpus** (~245K tokens: encyclopedic
prose, narrative prose, and source code) so the matrix sees every register the model
serves. The exact corpus is included as
[`Ling-3.0-tiny-Pollard-calibration.txt`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-calibration.txt).
The imatrix guides IQ-quant *quality*; it does **not** decide the allocation β the
measured KL sensitivity profile does. That two-step separation (imatrix for quality,
measured KL for where the bits go) is what Pollard adds on top of a standard imatrix
quant.
## Embed / output weights
Token-embedding and output tensors stay at **`q6_K`**, and every attention block is
kept at `q6_K`/`q5_K` rather than dropped to the IQ base β measured sensitivity says
those tensors don't tolerate crushing, so the bits are spent there and clawed back
from the least-sensitive FFN experts.
## ARM / AVX
llama.cpp "repacks" weights into an interleaved layout at load time for faster
inference on ARM and AVX machines β no special file needed, online repacking covers
these quants. The old `Q4_0_4_4/4_8/8_8` variants are not required.
## Notes
- **License:** MIT, inherited from the base model.
- KL was measured against a **Q6_K reference** on a held-out set (a memory-fit
reference on a 16 GB machine; the reported number is the *relative* win vs a
same-size uniform quant, which is what matters here).
- **Quantized, not fine-tuned** β identical weights, better bit allocation.
## Credits
- Base model: [`inclusionAI/Ling-3.0-tiny`](https://huggingface.co/inclusionAI/Ling-3.0-tiny) (Ant Group / inclusionAI)
- Quantization tooling: [llama.cpp](https://github.com/ggml-org/llama.cpp) (ggml-org)
- Method + tooling: [Pollard Weights](https://github.com/WestWaters/pollard-weights) β *measure first, no claim before a number.*
|