File size: 8,341 Bytes
5148d15
8f61394
 
5148d15
 
8f61394
 
 
 
5148d15
 
 
 
 
8f61394
5148d15
 
8f61394
5148d15
 
afc43dc
5148d15
afc43dc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5148d15
 
8f61394
 
 
5148d15
8f61394
5148d15
afc43dc
 
5148d15
8f61394
 
cfe65d6
8f61394
afc43dc
 
 
5148d15
afc43dc
 
 
8f61394
afc43dc
 
 
5148d15
8f61394
 
 
 
 
 
 
5148d15
afc43dc
 
9c06103
 
 
74c35dc
9c06103
74c35dc
9c06103
74c35dc
9c06103
74c35dc
9c06103
 
afc43dc
8f61394
5148d15
17b9c3d
 
 
 
 
 
 
 
 
 
 
5148d15
afc43dc
5148d15
 
8f61394
afc43dc
74c35dc
5148d15
 
8f61394
 
afc43dc
 
 
 
74c35dc
afc43dc
 
 
8f61394
 
74c35dc
 
8f61394
 
afc43dc
 
 
 
8f61394
 
afc43dc
 
 
 
 
8f61394
afc43dc
 
 
 
8f61394
afc43dc
 
 
 
 
 
 
 
 
 
 
 
8f61394
5148d15
 
 
 
afc43dc
 
8f61394
 
 
5148d15
8f61394
afc43dc
8f61394
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
---
quantized_by: PollardWeights
pipeline_tag: text-generation
base_model: inclusionAI/Ling-3.0-tiny
base_model_relation: quantized
license: mit
language:
- en
- zh
tags:
- pollard
- gguf
- llama.cpp
- moe
- bailingmoe3
- measured-sensitivity
- imatrix
- conversational
---

# Pollard measured-sensitivity quantizations of Ling-3.0-tiny by inclusionAI

Built with **[Pollard Weights](https://github.com/WestWaters/pollard-weights)** on
**llama.cpp** build `b10360` (`48d22e295`) β€” the first build line with `bailingmoe3`
support ([PR #26608](https://github.com/ggml-org/llama.cpp/pull/26608), merged
2026‑08‑17). Use that build or newer to run these.

Original model: https://huggingface.co/inclusionAI/Ling-3.0-tiny

## Model details

| | |
|---|---|
| Parameter count | ~7.9B total / ~1.7B active (MoE) β€” listed as 8B |
| Architecture | `bailingmoe3` (128 experts/layer, top‑8 + 1 shared, 24 layers) |
| Input support | text |
| Speculative decoding | no |
| imatrix | **yes** β€” [details below](#imatrix-calibration), corpus + matrix included in this repo |
| Perplexity / KLD measured | **yes** β€” this is the whole point (see next section) |

Uniform quants spend the same bits on every layer. Pollard **measures** how much
crushing each tensor group actually costs β€” KL-divergence, per layer β€” then a
KL-aware knapsack spends bits where they matter: more on the sensitive layers,
fewer on the ones that don't care. Same weights, smarter bit allocation.

## Why this over a uniform quant

Held-out KL-divergence vs a Q6_K reference (lower = closer to the full model),
measured on the same held-out set for every build:

| build | size | mean KL | vs uniform |
|---|---|---|---|
| **Ling-3.0-tiny Pollard** | **3.83 GB** | **0.1875** | _baseline_ |
| uniform IQ3 (interpolated to 3.83 GB) | 3.83 GB | β‰ˆ 0.204 | **β‰ˆ 8% higher KL** |
| uniform IQ3_S | 3.51 GB | 0.2821 | reference points |
| uniform IQ3_M | 3.56 GB | 0.2469 | (bracket the curve) |
| uniform IQ4_XS | 4.29 GB | 0.1312 | (bracket the curve) |

At matched size the measured allocation sits **below** the uniform size↔KL curve.
The measured mix: sensitive early layers get `iq4_xs`, most get `iq3_s`, the
least-sensitive get `iq2_s`; every attention block stays `q6_K`/`q5_K`;
embeddings/output stay `q6_K`; imatrix-uncovered MoE tensors are pinned so the
aggressive base can't crash. (ffn sensitivity spread ~6Γ—, attn spread ~16Γ— across
the 24 layers β€” that variance is exactly what a uniform quant wastes. The full
per-tensor map is in [`Ling-3.0-tiny-Pollard.tensor-types.txt`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.tensor-types.txt).)

## Prompt format

```
<role>SYSTEM</role>{system_prompt}
detailed thinking on<|role_end|><role>HUMAN</role>{prompt}<|role_end|><role>ASSISTANT</role>
<think>
```

## Which file should I choose?

Pick the rung for your machine β€” each is the **same weights**, sized to a different
RAM budget by the measured allocation:

- **~8 GB RAM / VRAM** β†’ **`IQ3_S`** (3.83 GB). The value pick: full model with room
  for context, and it beats same-size uniform IQ3 (table above). **Recommended.**
- **~9 GB** β†’ **`IQ4_XS`** (4.64 GB). More fidelity β€” the sensitive layers move up to
  `iq4_xs`.
- **~11 GB** β†’ **`Q6_K`** (6.26 GB). Near-lossless; as close to the full model as a
  quant gets.
- Want it even smaller than IQ3_S? Pollard *loses* to uniform at the extreme IQ2 floor
  for this model (the weights are too crushed for reallocation to help), so we don't
  ship one β€” *measure first, no claim before a number.*

## Available files

MoE speed: only ~1.7B of the 7.9B params are active per token, so even the big rungs
stay fast on an **Apple M4** (`tg`, llama.cpp Metal).

| Filename | Type | Size | M4 tok/s | Description |
|---|---|---|---|---|
| [Ling-3.0-tiny-Pollard-IQ3_S.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-IQ3_S.gguf) | IQ3 measured mix (IQ2_S→IQ4_XS, q6_K embed/attn) | 3.83 GB | **75.1** | Fits an ~8 GB box. Beats same-size uniform IQ3 (table above). **Recommended.** |
| [Ling-3.0-tiny-Pollard-IQ4_XS.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-IQ4_XS.gguf) | IQ4_XS measured mix (q6_K/q5_K attn, q6_K embed) | 4.64 GB | **75.2** | Fits an ~9 GB box. Higher fidelity β€” sensitive layers pushed to iq4_xs. |
| [Ling-3.0-tiny-Pollard-Q6_K.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-Q6_K.gguf) | Q5/Q6 measured mix (18L q6_K, 6L q5_K) | 6.26 GB | **66.9** | Fits an ~11 GB box. Near-lossless β€” maximum quality. |
| [Ling-3.0-tiny-Pollard.imatrix](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.imatrix) | importance matrix | 44 MB | β€” | The imatrix used, for anyone re-quantizing. |
| [Ling-3.0-tiny-Pollard-calibration.txt](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-calibration.txt) | calibration corpus | ~1 MB | β€” | The exact corpus the imatrix was computed on. |
| [Ling-3.0-tiny-Pollard.tensor-types.txt](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.tensor-types.txt) | allocation map | 3 KB | β€” | The measured per-tensor bit assignment. |

## Download a specific file

```bash
pip install -U "huggingface_hub[cli]"
hf download PollardWeights/Ling-3.0-tiny-Pollard \
  --include "Ling-3.0-tiny-Pollard-IQ3_S.gguf" --local-dir ./
```

## How to run

These are standard GGUF and run with **llama.cpp** β€” one-line install:

```bash
curl -LsSf https://llama.app/install.sh | sh
llama-server -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
```

or with a local file:

```bash
llama-cli    -m Ling-3.0-tiny-Pollard-IQ3_S.gguf -ngl 99 -p "Explain MoE routing simply."
llama-server -m Ling-3.0-tiny-Pollard-IQ3_S.gguf -ngl 99      # OpenAI-compatible API + web UI at :8080
```

They also work in anything built on llama.cpp β€” **LM Studio, koboldcpp, ramalama,
Jan, Text Generation WebUI, LoLLMs** β€” provided the build is recent enough to carry
`bailingmoe3` support (see top). If the app ships an older llama.cpp, update it first.

## imatrix (calibration)

The importance matrix ([`Ling-3.0-tiny-Pollard.imatrix`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.imatrix),
included) was computed on a **mixed-domain corpus** (~245K tokens: encyclopedic
prose, narrative prose, and source code) so the matrix sees every register the model
serves. The exact corpus is included as
[`Ling-3.0-tiny-Pollard-calibration.txt`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-calibration.txt).

The imatrix guides IQ-quant *quality*; it does **not** decide the allocation β€” the
measured KL sensitivity profile does. That two-step separation (imatrix for quality,
measured KL for where the bits go) is what Pollard adds on top of a standard imatrix
quant.

## Embed / output weights

Token-embedding and output tensors stay at **`q6_K`**, and every attention block is
kept at `q6_K`/`q5_K` rather than dropped to the IQ base β€” measured sensitivity says
those tensors don't tolerate crushing, so the bits are spent there and clawed back
from the least-sensitive FFN experts.

## ARM / AVX

llama.cpp "repacks" weights into an interleaved layout at load time for faster
inference on ARM and AVX machines β€” no special file needed, online repacking covers
these quants. The old `Q4_0_4_4/4_8/8_8` variants are not required.

## Notes

- **License:** MIT, inherited from the base model.
- KL was measured against a **Q6_K reference** on a held-out set (a memory-fit
  reference on a 16 GB machine; the reported number is the *relative* win vs a
  same-size uniform quant, which is what matters here).
- **Quantized, not fine-tuned** β€” identical weights, better bit allocation.

## Credits

- Base model: [`inclusionAI/Ling-3.0-tiny`](https://huggingface.co/inclusionAI/Ling-3.0-tiny) (Ant Group / inclusionAI)
- Quantization tooling: [llama.cpp](https://github.com/ggml-org/llama.cpp) (ggml-org)
- Method + tooling: [Pollard Weights](https://github.com/WestWaters/pollard-weights) β€” *measure first, no claim before a number.*