Text Generation
GGUF
English
Chinese
pollard
llama.cpp
Mixture of Experts
bailingmoe3
measured-sensitivity
imatrix
conversational
Instructions to use PollardWeights/Ling-3.0-tiny-Pollard with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use PollardWeights/Ling-3.0-tiny-Pollard with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S # Run inference directly in the terminal: llama cli -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S # Run inference directly in the terminal: llama cli -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S # Run inference directly in the terminal: ./llama-cli -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Use Docker
docker model run hf.co/PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
- LM Studio
- Jan
- vLLM
How to use PollardWeights/Ling-3.0-tiny-Pollard with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PollardWeights/Ling-3.0-tiny-Pollard" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PollardWeights/Ling-3.0-tiny-Pollard", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
- Ollama
How to use PollardWeights/Ling-3.0-tiny-Pollard with Ollama:
ollama run hf.co/PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
- Unsloth Studio
How to use PollardWeights/Ling-3.0-tiny-Pollard with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for PollardWeights/Ling-3.0-tiny-Pollard to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for PollardWeights/Ling-3.0-tiny-Pollard to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for PollardWeights/Ling-3.0-tiny-Pollard to start chatting
- Pi
How to use PollardWeights/Ling-3.0-tiny-Pollard with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use PollardWeights/Ling-3.0-tiny-Pollard with Docker Model Runner:
docker model run hf.co/PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
- Lemonade
How to use PollardWeights/Ling-3.0-tiny-Pollard with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Run and chat with the model
lemonade run user.Ling-3.0-tiny-Pollard-IQ3_S
List all available models
lemonade list
- Hermes Agent
How to use PollardWeights/Ling-3.0-tiny-Pollard with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use PollardWeights/Ling-3.0-tiny-Pollard with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| quantized_by: PollardWeights | |
| pipeline_tag: text-generation | |
| base_model: inclusionAI/Ling-3.0-tiny | |
| base_model_relation: quantized | |
| license: mit | |
| language: | |
| - en | |
| - zh | |
| tags: | |
| - pollard | |
| - gguf | |
| - llama.cpp | |
| - moe | |
| - bailingmoe3 | |
| - measured-sensitivity | |
| - imatrix | |
| - conversational | |
| # Pollard measured-sensitivity quantizations of Ling-3.0-tiny by inclusionAI | |
| Built with **[Pollard Weights](https://github.com/WestWaters/pollard-weights)** on | |
| **llama.cpp** build `b10360` (`48d22e295`) β the first build line with `bailingmoe3` | |
| support ([PR #26608](https://github.com/ggml-org/llama.cpp/pull/26608), merged | |
| 2026β08β17). Use that build or newer to run these. | |
| Original model: https://huggingface.co/inclusionAI/Ling-3.0-tiny | |
| ## Model details | |
| | | | | |
| |---|---| | |
| | Parameter count | ~7.9B total / ~1.7B active (MoE) β listed as 8B | | |
| | Architecture | `bailingmoe3` (128 experts/layer, topβ8 + 1 shared, 24 layers) | | |
| | Input support | text | | |
| | Speculative decoding | no | | |
| | imatrix | **yes** β [details below](#imatrix-calibration), corpus + matrix included in this repo | | |
| | Perplexity / KLD measured | **yes** β this is the whole point (see next section) | | |
| Uniform quants spend the same bits on every layer. Pollard **measures** how much | |
| crushing each tensor group actually costs β KL-divergence, per layer β then a | |
| KL-aware knapsack spends bits where they matter: more on the sensitive layers, | |
| fewer on the ones that don't care. Same weights, smarter bit allocation. | |
| ## Why this over a uniform quant | |
| Held-out KL-divergence vs a Q6_K reference (lower = closer to the full model), | |
| measured on the same held-out set for every build: | |
| | build | size | mean KL | vs uniform | | |
| |---|---|---|---| | |
| | **Ling-3.0-tiny Pollard** | **3.83 GB** | **0.1875** | _baseline_ | | |
| | uniform IQ3 (interpolated to 3.83 GB) | 3.83 GB | β 0.204 | **β 8% higher KL** | | |
| | uniform IQ3_S | 3.51 GB | 0.2821 | reference points | | |
| | uniform IQ3_M | 3.56 GB | 0.2469 | (bracket the curve) | | |
| | uniform IQ4_XS | 4.29 GB | 0.1312 | (bracket the curve) | | |
| At matched size the measured allocation sits **below** the uniform sizeβKL curve. | |
| The measured mix: sensitive early layers get `iq4_xs`, most get `iq3_s`, the | |
| least-sensitive get `iq2_s`; every attention block stays `q6_K`/`q5_K`; | |
| embeddings/output stay `q6_K`; imatrix-uncovered MoE tensors are pinned so the | |
| aggressive base can't crash. (ffn sensitivity spread ~6Γ, attn spread ~16Γ across | |
| the 24 layers β that variance is exactly what a uniform quant wastes. The full | |
| per-tensor map is in [`Ling-3.0-tiny-Pollard.tensor-types.txt`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.tensor-types.txt).) | |
| ## Prompt format | |
| ``` | |
| <role>SYSTEM</role>{system_prompt} | |
| detailed thinking on<|role_end|><role>HUMAN</role>{prompt}<|role_end|><role>ASSISTANT</role> | |
| <think> | |
| ``` | |
| ## Which file should I choose? | |
| Pick the rung for your machine β each is the **same weights**, sized to a different | |
| RAM budget by the measured allocation: | |
| - **~8 GB RAM / VRAM** β **`IQ3_S`** (3.83 GB). The value pick: full model with room | |
| for context, and it beats same-size uniform IQ3 (table above). **Recommended.** | |
| - **~9 GB** β **`IQ4_XS`** (4.64 GB). More fidelity β the sensitive layers move up to | |
| `iq4_xs`. | |
| - **~11 GB** β **`Q6_K`** (6.26 GB). Near-lossless; as close to the full model as a | |
| quant gets. | |
| - Want it even smaller than IQ3_S? Pollard *loses* to uniform at the extreme IQ2 floor | |
| for this model (the weights are too crushed for reallocation to help), so we don't | |
| ship one β *measure first, no claim before a number.* | |
| ## Available files | |
| MoE speed: only ~1.7B of the 7.9B params are active per token, so even the big rungs | |
| stay fast on an **Apple M4** (`tg`, llama.cpp Metal). | |
| | Filename | Type | Size | M4 tok/s | Description | | |
| |---|---|---|---|---| | |
| | [Ling-3.0-tiny-Pollard-IQ3_S.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-IQ3_S.gguf) | IQ3 measured mix (IQ2_SβIQ4_XS, q6_K embed/attn) | 3.83 GB | **75.1** | Fits an ~8 GB box. Beats same-size uniform IQ3 (table above). **Recommended.** | | |
| | [Ling-3.0-tiny-Pollard-IQ4_XS.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-IQ4_XS.gguf) | IQ4_XS measured mix (q6_K/q5_K attn, q6_K embed) | 4.64 GB | **75.2** | Fits an ~9 GB box. Higher fidelity β sensitive layers pushed to iq4_xs. | | |
| | [Ling-3.0-tiny-Pollard-Q6_K.gguf](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-Q6_K.gguf) | Q5/Q6 measured mix (18L q6_K, 6L q5_K) | 6.26 GB | **66.9** | Fits an ~11 GB box. Near-lossless β maximum quality. | | |
| | [Ling-3.0-tiny-Pollard.imatrix](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.imatrix) | importance matrix | 44 MB | β | The imatrix used, for anyone re-quantizing. | | |
| | [Ling-3.0-tiny-Pollard-calibration.txt](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-calibration.txt) | calibration corpus | ~1 MB | β | The exact corpus the imatrix was computed on. | | |
| | [Ling-3.0-tiny-Pollard.tensor-types.txt](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.tensor-types.txt) | allocation map | 3 KB | β | The measured per-tensor bit assignment. | | |
| ## Download a specific file | |
| ```bash | |
| pip install -U "huggingface_hub[cli]" | |
| hf download PollardWeights/Ling-3.0-tiny-Pollard \ | |
| --include "Ling-3.0-tiny-Pollard-IQ3_S.gguf" --local-dir ./ | |
| ``` | |
| ## How to run | |
| These are standard GGUF and run with **llama.cpp** β one-line install: | |
| ```bash | |
| curl -LsSf https://llama.app/install.sh | sh | |
| llama-server -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S | |
| ``` | |
| or with a local file: | |
| ```bash | |
| llama-cli -m Ling-3.0-tiny-Pollard-IQ3_S.gguf -ngl 99 -p "Explain MoE routing simply." | |
| llama-server -m Ling-3.0-tiny-Pollard-IQ3_S.gguf -ngl 99 # OpenAI-compatible API + web UI at :8080 | |
| ``` | |
| They also work in anything built on llama.cpp β **LM Studio, koboldcpp, ramalama, | |
| Jan, Text Generation WebUI, LoLLMs** β provided the build is recent enough to carry | |
| `bailingmoe3` support (see top). If the app ships an older llama.cpp, update it first. | |
| ## imatrix (calibration) | |
| The importance matrix ([`Ling-3.0-tiny-Pollard.imatrix`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard.imatrix), | |
| included) was computed on a **mixed-domain corpus** (~245K tokens: encyclopedic | |
| prose, narrative prose, and source code) so the matrix sees every register the model | |
| serves. The exact corpus is included as | |
| [`Ling-3.0-tiny-Pollard-calibration.txt`](https://huggingface.co/PollardWeights/Ling-3.0-tiny-Pollard/blob/main/Ling-3.0-tiny-Pollard-calibration.txt). | |
| The imatrix guides IQ-quant *quality*; it does **not** decide the allocation β the | |
| measured KL sensitivity profile does. That two-step separation (imatrix for quality, | |
| measured KL for where the bits go) is what Pollard adds on top of a standard imatrix | |
| quant. | |
| ## Embed / output weights | |
| Token-embedding and output tensors stay at **`q6_K`**, and every attention block is | |
| kept at `q6_K`/`q5_K` rather than dropped to the IQ base β measured sensitivity says | |
| those tensors don't tolerate crushing, so the bits are spent there and clawed back | |
| from the least-sensitive FFN experts. | |
| ## ARM / AVX | |
| llama.cpp "repacks" weights into an interleaved layout at load time for faster | |
| inference on ARM and AVX machines β no special file needed, online repacking covers | |
| these quants. The old `Q4_0_4_4/4_8/8_8` variants are not required. | |
| ## Notes | |
| - **License:** MIT, inherited from the base model. | |
| - KL was measured against a **Q6_K reference** on a held-out set (a memory-fit | |
| reference on a 16 GB machine; the reported number is the *relative* win vs a | |
| same-size uniform quant, which is what matters here). | |
| - **Quantized, not fine-tuned** β identical weights, better bit allocation. | |
| ## Credits | |
| - Base model: [`inclusionAI/Ling-3.0-tiny`](https://huggingface.co/inclusionAI/Ling-3.0-tiny) (Ant Group / inclusionAI) | |
| - Quantization tooling: [llama.cpp](https://github.com/ggml-org/llama.cpp) (ggml-org) | |
| - Method + tooling: [Pollard Weights](https://github.com/WestWaters/pollard-weights) β *measure first, no claim before a number.* | |