--- quantized_by: PollardWeights pipeline_tag: text-generation base_model: inclusionAI/Ling-3.0-tiny base_model_relation: quantized license: mit language: - en - zh tags: - pollard - gguf - llama.cpp - moe - bailingmoe3 - measured-sensitivity - imatrix - conversational --- # Pollard measured-sensitivity quantizations of Ling-3.0-tiny by inclusionAI Built with **[Pollard Weights](https://github.com/WestWaters/pollard-weights)** on **llama.cpp** build `b10360` (`48d22e295`) — the first build line with `bailingmoe3` support ([PR #26608](https://github.com/ggml-org/llama.cpp/pull/26608), merged 2026‑08‑17). Use that build or newer to run these. Original model: https://huggingface.co/inclusionAI/Ling-3.0-tiny ## Model details | | | |---|---| | Parameter count | ~7.9B total / ~1.7B active (MoE) — listed as 8B | | Architecture | `bailingmoe3` (128 experts/layer, top‑8 + 1 shared, 24 layers) | | Input support | text | | Speculative decoding | no | | imatrix | **yes** — [details below](#imatrix-calibration), corpus + matrix included in this repo | | Perplexity / KLD measured | **yes** — this is the whole point (see next section) | Uniform quants spend the same bits on every layer. Pollard **measures** how much crushing each tensor group actually costs — KL-divergence, per layer — then a KL-aware knapsack spends bits where they matter: more on the sensitive layers, fewer on the ones that don't care. Same weights, smarter bit allocation. ## Why this over a uniform quant Held-out KL-divergence vs a Q6_K reference (lower = closer to the full model), measured on the same held-out set for every build: | build | size | mean KL | vs uniform | |---|---|---|---| | **Ling-3.0-tiny Pollard** | **3.83 GB** | **0.1875** | _baseline_ | | uniform IQ3 (interpolated to 3.83 GB) | 3.83 GB | ≈ 0.204 | **≈ 8% higher KL** | | uniform IQ3_S | 3.51 GB | 0.2821 | reference points | | uniform IQ3_M | 3.56 GB | 0.2469 | (bracket the curve) | | uniform IQ4_XS | 4.29 GB | 0.1312 | (bracket the curve) | At matched size the measured allocation sits **below** the uniform size↔KL curve. The measured mix: sensitive early layers get `iq4_xs`, most get `iq3_s`, the least-sensitive get `iq2_s`; every attention block stays `q6_K`/`q5_K`; embeddings/output stay `q6_K`; imatrix-uncovered MoE tensors are pinned so the aggressive base can't crash. (ffn sensitivity spread ~6×, attn spread ~16× across the 24 layers — that variance is exactly what a uniform quant wastes. The full per-tensor map is in [`Ling-3.0-tiny-Pollard.tensor-types.txt`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.tensor-types.txt).) ## Prompt format ``` SYSTEM{system_prompt} detailed thinking on<|role_end|>HUMAN{prompt}<|role_end|>ASSISTANT ``` ## Which file should I choose? Pick the rung for your machine — each is the **same weights**, sized to a different RAM budget by the measured allocation: - **~8 GB RAM / VRAM** → **`IQ3_S`** (3.83 GB). The value pick: full model with room for context, and it beats same-size uniform IQ3 (table above). **Recommended.** - **~9 GB** → **`IQ4_XS`** (4.64 GB). More fidelity — the sensitive layers move up to `iq4_xs`. - **~11 GB** → **`Q6_K`** (6.26 GB). Near-lossless; as close to the full model as a quant gets. - Want it even smaller than IQ3_S? Pollard *loses* to uniform at the extreme IQ2 floor for this model (the weights are too crushed for reallocation to help), so we don't ship one — *measure first, no claim before a number.* ## Available files MoE speed: only ~1.7B of the 7.9B params are active per token, so even the big rungs stay fast on an **Apple M4** (`tg`, llama.cpp Metal). | Filename | Type | Size | M4 tok/s | Description | |---|---|---|---|---| | [Ling-3.0-tiny-Pollard-IQ3_S.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-IQ3_S.gguf) | IQ3 measured mix (IQ2_S→IQ4_XS, q6_K embed/attn) | 3.83 GB | **75.1** | Fits an ~8 GB box. Beats same-size uniform IQ3 (table above). **Recommended.** | | [Ling-3.0-tiny-Pollard-IQ4_XS.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-IQ4_XS.gguf) | IQ4_XS measured mix (q6_K/q5_K attn, q6_K embed) | 4.64 GB | **75.2** | Fits an ~9 GB box. Higher fidelity — sensitive layers pushed to iq4_xs. | | [Ling-3.0-tiny-Pollard-Q6_K.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-Q6_K.gguf) | Q5/Q6 measured mix (18L q6_K, 6L q5_K) | 6.26 GB | **66.9** | Fits an ~11 GB box. Near-lossless — maximum quality. | | [Ling-3.0-tiny-Pollard.imatrix](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.imatrix) | importance matrix | 44 MB | — | The imatrix used, for anyone re-quantizing. | | [Ling-3.0-tiny-Pollard-calibration.txt](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-calibration.txt) | calibration corpus | ~1 MB | — | The exact corpus the imatrix was computed on. | | [Ling-3.0-tiny-Pollard.tensor-types.txt](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.tensor-types.txt) | allocation map | 3 KB | — | The measured per-tensor bit assignment. | ## Download a specific file ```bash pip install -U "huggingface_hub[cli]" hf download PollardWeights/Ling-3.0-tiny-Pollard \ --include "Ling-3.0-tiny-Pollard-IQ3_S.gguf" --local-dir ./ ``` ## How to run These are standard GGUF and run with **llama.cpp** — one-line install: ```bash curl -LsSf https://llama.app/install.sh | sh llama-server -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S ``` or with a local file: ```bash llama-cli -m Ling-3.0-tiny-Pollard-IQ3_S.gguf -ngl 99 -p "Explain MoE routing simply." llama-server -m Ling-3.0-tiny-Pollard-IQ3_S.gguf -ngl 99 # OpenAI-compatible API + web UI at :8080 ``` They also work in anything built on llama.cpp — **LM Studio, koboldcpp, ramalama, Jan, Text Generation WebUI, LoLLMs** — provided the build is recent enough to carry `bailingmoe3` support (see top). If the app ships an older llama.cpp, update it first. ## imatrix (calibration) The importance matrix ([`Ling-3.0-tiny-Pollard.imatrix`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.imatrix), included) was computed on a **mixed-domain corpus** (~245K tokens: encyclopedic prose, narrative prose, and source code) so the matrix sees every register the model serves. The exact corpus is included as [`Ling-3.0-tiny-Pollard-calibration.txt`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-calibration.txt). The imatrix guides IQ-quant *quality*; it does **not** decide the allocation — the measured KL sensitivity profile does. That two-step separation (imatrix for quality, measured KL for where the bits go) is what Pollard adds on top of a standard imatrix quant. ## Embed / output weights Token-embedding and output tensors stay at **`q6_K`**, and every attention block is kept at `q6_K`/`q5_K` rather than dropped to the IQ base — measured sensitivity says those tensors don't tolerate crushing, so the bits are spent there and clawed back from the least-sensitive FFN experts. ## ARM / AVX llama.cpp "repacks" weights into an interleaved layout at load time for faster inference on ARM and AVX machines — no special file needed, online repacking covers these quants. The old `Q4_0_4_4/4_8/8_8` variants are not required. ## Notes - **License:** MIT, inherited from the base model. - KL was measured against a **Q6_K reference** on a held-out set (a memory-fit reference on a 16 GB machine; the reported number is the *relative* win vs a same-size uniform quant, which is what matters here). - **Quantized, not fine-tuned** — identical weights, better bit allocation. ## Credits - Base model: [`inclusionAI/Ling-3.0-tiny`](https://huggingface.co/inclusionAI/Ling-3.0-tiny) (Ant Group / inclusionAI) - Quantization tooling: [llama.cpp](https://github.com/ggml-org/llama.cpp) (ggml-org) - Method + tooling: [Pollard Weights](https://github.com/WestWaters/pollard-weights) — *measure first, no claim before a number.*