---
quantized_by: PollardWeights
pipeline_tag: text-generation
base_model: inclusionAI/Ling-3.0-tiny
base_model_relation: quantized
license: mit
language:
- en
- zh
tags:
- pollard
- gguf
- llama.cpp
- moe
- bailingmoe3
- measured-sensitivity
- imatrix
- conversational
---
# Pollard measured-sensitivity quantizations of Ling-3.0-tiny by inclusionAI
Built with **[Pollard Weights](https://github.com/WestWaters/pollard-weights)** on
**llama.cpp** build `b10360` (`48d22e295`) — the first build line with `bailingmoe3`
support ([PR #26608](https://github.com/ggml-org/llama.cpp/pull/26608), merged
2026‑08‑17). Use that build or newer to run these.
Original model: https://huggingface.co/inclusionAI/Ling-3.0-tiny
## Model details
| | |
|---|---|
| Parameter count | ~7.9B total / ~1.7B active (MoE) — listed as 8B |
| Architecture | `bailingmoe3` (128 experts/layer, top‑8 + 1 shared, 24 layers) |
| Input support | text |
| Speculative decoding | no |
| imatrix | **yes** — [details below](#imatrix-calibration), corpus + matrix included in this repo |
| Perplexity / KLD measured | **yes** — this is the whole point (see next section) |
Uniform quants spend the same bits on every layer. Pollard **measures** how much
crushing each tensor group actually costs — KL-divergence, per layer — then a
KL-aware knapsack spends bits where they matter: more on the sensitive layers,
fewer on the ones that don't care. Same weights, smarter bit allocation.
## Why this over a uniform quant
Held-out KL-divergence vs a Q6_K reference (lower = closer to the full model),
measured on the same held-out set for every build:
| build | size | mean KL | vs uniform |
|---|---|---|---|
| **Ling-3.0-tiny Pollard** | **3.83 GB** | **0.1875** | _baseline_ |
| uniform IQ3 (interpolated to 3.83 GB) | 3.83 GB | ≈ 0.204 | **≈ 8% higher KL** |
| uniform IQ3_S | 3.51 GB | 0.2821 | reference points |
| uniform IQ3_M | 3.56 GB | 0.2469 | (bracket the curve) |
| uniform IQ4_XS | 4.29 GB | 0.1312 | (bracket the curve) |
At matched size the measured allocation sits **below** the uniform size↔KL curve.
The measured mix: sensitive early layers get `iq4_xs`, most get `iq3_s`, the
least-sensitive get `iq2_s`; every attention block stays `q6_K`/`q5_K`;
embeddings/output stay `q6_K`; imatrix-uncovered MoE tensors are pinned so the
aggressive base can't crash. (ffn sensitivity spread ~6×, attn spread ~16× across
the 24 layers — that variance is exactly what a uniform quant wastes. The full
per-tensor map is in [`Ling-3.0-tiny-Pollard.tensor-types.txt`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.tensor-types.txt).)
## Prompt format
```
SYSTEM{system_prompt}
detailed thinking on<|role_end|>HUMAN{prompt}<|role_end|>ASSISTANT
```
## Which file should I choose?
Pick the rung for your machine — each is the **same weights**, sized to a different
RAM budget by the measured allocation:
- **~8 GB RAM / VRAM** → **`IQ3_S`** (3.83 GB). The value pick: full model with room
for context, and it beats same-size uniform IQ3 (table above). **Recommended.**
- **~9 GB** → **`IQ4_XS`** (4.64 GB). More fidelity — the sensitive layers move up to
`iq4_xs`.
- **~11 GB** → **`Q6_K`** (6.26 GB). Near-lossless; as close to the full model as a
quant gets.
- Want it even smaller than IQ3_S? Pollard *loses* to uniform at the extreme IQ2 floor
for this model (the weights are too crushed for reallocation to help), so we don't
ship one — *measure first, no claim before a number.*
## Available files
MoE speed: only ~1.7B of the 7.9B params are active per token, so even the big rungs
stay fast on an **Apple M4** (`tg`, llama.cpp Metal).
| Filename | Type | Size | M4 tok/s | Description |
|---|---|---|---|---|
| [Ling-3.0-tiny-Pollard-IQ3_S.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-IQ3_S.gguf) | IQ3 measured mix (IQ2_S→IQ4_XS, q6_K embed/attn) | 3.83 GB | **75.1** | Fits an ~8 GB box. Beats same-size uniform IQ3 (table above). **Recommended.** |
| [Ling-3.0-tiny-Pollard-IQ4_XS.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-IQ4_XS.gguf) | IQ4_XS measured mix (q6_K/q5_K attn, q6_K embed) | 4.64 GB | **75.2** | Fits an ~9 GB box. Higher fidelity — sensitive layers pushed to iq4_xs. |
| [Ling-3.0-tiny-Pollard-Q6_K.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-Q6_K.gguf) | Q5/Q6 measured mix (18L q6_K, 6L q5_K) | 6.26 GB | **66.9** | Fits an ~11 GB box. Near-lossless — maximum quality. |
| [Ling-3.0-tiny-Pollard.imatrix](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.imatrix) | importance matrix | 44 MB | — | The imatrix used, for anyone re-quantizing. |
| [Ling-3.0-tiny-Pollard-calibration.txt](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-calibration.txt) | calibration corpus | ~1 MB | — | The exact corpus the imatrix was computed on. |
| [Ling-3.0-tiny-Pollard.tensor-types.txt](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.tensor-types.txt) | allocation map | 3 KB | — | The measured per-tensor bit assignment. |
## Download a specific file
```bash
pip install -U "huggingface_hub[cli]"
hf download PollardWeights/Ling-3.0-tiny-Pollard \
--include "Ling-3.0-tiny-Pollard-IQ3_S.gguf" --local-dir ./
```
## How to run
These are standard GGUF and run with **llama.cpp** — one-line install:
```bash
curl -LsSf https://llama.app/install.sh | sh
llama-server -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
```
or with a local file:
```bash
llama-cli -m Ling-3.0-tiny-Pollard-IQ3_S.gguf -ngl 99 -p "Explain MoE routing simply."
llama-server -m Ling-3.0-tiny-Pollard-IQ3_S.gguf -ngl 99 # OpenAI-compatible API + web UI at :8080
```
They also work in anything built on llama.cpp — **LM Studio, koboldcpp, ramalama,
Jan, Text Generation WebUI, LoLLMs** — provided the build is recent enough to carry
`bailingmoe3` support (see top). If the app ships an older llama.cpp, update it first.
## imatrix (calibration)
The importance matrix ([`Ling-3.0-tiny-Pollard.imatrix`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.imatrix),
included) was computed on a **mixed-domain corpus** (~245K tokens: encyclopedic
prose, narrative prose, and source code) so the matrix sees every register the model
serves. The exact corpus is included as
[`Ling-3.0-tiny-Pollard-calibration.txt`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-calibration.txt).
The imatrix guides IQ-quant *quality*; it does **not** decide the allocation — the
measured KL sensitivity profile does. That two-step separation (imatrix for quality,
measured KL for where the bits go) is what Pollard adds on top of a standard imatrix
quant.
## Embed / output weights
Token-embedding and output tensors stay at **`q6_K`**, and every attention block is
kept at `q6_K`/`q5_K` rather than dropped to the IQ base — measured sensitivity says
those tensors don't tolerate crushing, so the bits are spent there and clawed back
from the least-sensitive FFN experts.
## ARM / AVX
llama.cpp "repacks" weights into an interleaved layout at load time for faster
inference on ARM and AVX machines — no special file needed, online repacking covers
these quants. The old `Q4_0_4_4/4_8/8_8` variants are not required.
## Notes
- **License:** MIT, inherited from the base model.
- KL was measured against a **Q6_K reference** on a held-out set (a memory-fit
reference on a 16 GB machine; the reported number is the *relative* win vs a
same-size uniform quant, which is what matters here).
- **Quantized, not fine-tuned** — identical weights, better bit allocation.
## Credits
- Base model: [`inclusionAI/Ling-3.0-tiny`](https://huggingface.co/inclusionAI/Ling-3.0-tiny) (Ant Group / inclusionAI)
- Quantization tooling: [llama.cpp](https://github.com/ggml-org/llama.cpp) (ggml-org)
- Method + tooling: [Pollard Weights](https://github.com/WestWaters/pollard-weights) — *measure first, no claim before a number.*