Instructions to use Mike0021/Ling-3.0-tiny-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Mike0021/Ling-3.0-tiny-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Mike0021/Ling-3.0-tiny-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Mike0021/Ling-3.0-tiny-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mike0021/Ling-3.0-tiny-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
- Ollama
How to use Mike0021/Ling-3.0-tiny-GGUF with Ollama:
ollama run hf.co/Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
- Unsloth Studio
How to use Mike0021/Ling-3.0-tiny-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Mike0021/Ling-3.0-tiny-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Mike0021/Ling-3.0-tiny-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Mike0021/Ling-3.0-tiny-GGUF to start chatting
- Pi
How to use Mike0021/Ling-3.0-tiny-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use Mike0021/Ling-3.0-tiny-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use Mike0021/Ling-3.0-tiny-GGUF with Docker Model Runner:
docker model run hf.co/Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
- Lemonade
How to use Mike0021/Ling-3.0-tiny-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Ling-3.0-tiny-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Mike0021/Ling-3.0-tiny-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
| # Reproducing this GGUF release | |
| The release is pinned to immutable source, converter, calibration, and | |
| validation revisions. BailingMoE3 support is not yet merged in upstream | |
| `llama.cpp`; use the exact revision below. | |
| ## 1. Build the converter and runtime | |
| ```bash | |
| git clone https://github.com/aetherbird/llama.cpp.git | |
| git -C llama.cpp checkout d8d862521e9ad842f2b47f3b392b039317782aa0 | |
| python3 -m venv .venv | |
| . .venv/bin/activate | |
| python -m pip install -r llama.cpp/requirements.txt | |
| python -m pip install --extra-index-url https://download.pytorch.org/whl/cu128 \ | |
| gguf==0.19.0 huggingface_hub==0.36.2 numpy==1.26.4 \ | |
| protobuf==4.25.9 safetensors==0.8.0 sentencepiece==0.2.2 \ | |
| tokenizers==0.22.2 'torch==2.8.0+cu128' tqdm==4.70.0 \ | |
| transformers==4.57.6 | |
| cmake -S llama.cpp -B llama.cpp/build \ | |
| -DCMAKE_BUILD_TYPE=Release \ | |
| -DGGML_CUDA=ON \ | |
| -DGGML_NATIVE=OFF \ | |
| -DCMAKE_CUDA_ARCHITECTURES=native | |
| cmake --build llama.cpp/build --config Release --parallel \ | |
| --target llama-cli llama-completion llama-server llama-tokenize llama-imatrix \ | |
| llama-quantize llama-perplexity | |
| ``` | |
| For a CPU-only conversion build, omit `-DGGML_CUDA=ON` and | |
| `-DCMAKE_CUDA_ARCHITECTURES=native`. | |
| ## 2. Download and convert the immutable source revision | |
| ```bash | |
| hf download inclusionAI/Ling-3.0-tiny \ | |
| --revision a2ee06c0f2de5b171701aee7f73f70a1da75483b \ | |
| --local-dir Ling-3.0-tiny | |
| (cd Ling-3.0-tiny && sha256sum -c ../source-safetensors.sha256) | |
| python llama.cpp/convert_hf_to_gguf.py Ling-3.0-tiny \ | |
| --outtype bf16 \ | |
| --outfile Ling-3.0-tiny-BF16.gguf \ | |
| --verbose | |
| ``` | |
| The expected BF16 result is 526 tensors, 15,803,475,232 bytes, with SHA-256 | |
| `2020d58d44887c4078c310dad0363b2af0898a9ed6cc46675c0d23201833952e`. | |
| ## 3. Collect the importance matrix | |
| ```bash | |
| python -m pip install duckdb==1.3.2 | |
| hf download lemon07r/bartowski-imatrix-v5-semantic \ | |
| bartowski-imatrix-v5-semantic.txt \ | |
| --repo-type dataset \ | |
| --revision a306f203ee4323e0afe846ae02c2daafe17384d9 \ | |
| --local-dir calibration | |
| hf download eaddario/imatrix-calibration \ | |
| combined_all_micro.parquet \ | |
| --repo-type dataset \ | |
| --revision e87ed55dcba9d9c3a3e41539f3e728e981b1daa4 \ | |
| --local-dir calibration/eaddario | |
| python - <<'PY' | |
| import duckdb | |
| source = "calibration/eaddario/combined_all_micro.parquet" | |
| output = "calibration/eaddario/combined_all_micro.txt" | |
| rows = duckdb.connect().execute( | |
| "SELECT content FROM read_parquet(?)", [source] | |
| ).fetchall() | |
| with open(output, "w", encoding="utf-8", newline="\n") as handle: | |
| for (content,) in rows: | |
| if content: | |
| handle.write(content.replace("\x00", "")) | |
| handle.write("\n") | |
| PY | |
| printf '%s\n' \ | |
| 'ff879b5a748f822ef539e43c596a3f44ab922f0295ee209d4220d9f86e86a063 calibration/bartowski-imatrix-v5-semantic.txt' \ | |
| '94389921e1f67b180a99de28c3090b41ce6f1960eb13abad21b7eba7cbe11b26 calibration/eaddario/combined_all_micro.parquet' \ | |
| 'fdb2d41abf04a2fb207502741a561a5a9ab385eb0c44a450eae676c410955946 calibration/eaddario/combined_all_micro.txt' \ | |
| | sha256sum -c - | |
| llama.cpp/build/bin/llama-imatrix \ | |
| --model Ling-3.0-tiny-BF16.gguf \ | |
| --file calibration/bartowski-imatrix-v5-semantic.txt \ | |
| --output-file Ling-3.0-tiny-imatrix-primary.gguf \ | |
| --output-format gguf \ | |
| --gpu-layers all --ctx-size 4096 --batch-size 4096 --ubatch-size 512 \ | |
| --threads 16 --threads-batch 16 --device CUDA0 \ | |
| --split-mode none --main-gpu 0 --fit off --flash-attn off \ | |
| --offline --no-ppl \ | |
| --output-frequency 25 | |
| llama.cpp/build/bin/llama-imatrix \ | |
| --model Ling-3.0-tiny-BF16.gguf \ | |
| --file calibration/eaddario/combined_all_micro.txt \ | |
| --in-file Ling-3.0-tiny-imatrix-primary.gguf \ | |
| --output-file Ling-3.0-tiny-imatrix.gguf \ | |
| --output-format gguf \ | |
| --gpu-layers all --ctx-size 4096 --batch-size 4096 --ubatch-size 512 \ | |
| --threads 16 --threads-batch 16 --device CUDA0 \ | |
| --split-mode none --main-gpu 0 --fit off --flash-attn off \ | |
| --offline --no-ppl \ | |
| --output-frequency 25 | |
| ``` | |
| The primary text SHA-256 is | |
| `ff879b5a748f822ef539e43c596a3f44ab922f0295ee209d4220d9f86e86a063`. | |
| The supplement parquet and extracted text SHA-256 values are respectively | |
| `94389921e1f67b180a99de28c3090b41ce6f1960eb13abad21b7eba7cbe11b26` | |
| and `fdb2d41abf04a2fb207502741a561a5a9ab385eb0c44a450eae676c410955946`. | |
| They are used only for activation statistics, not training or evaluation. | |
| `--process-output` is intentionally omitted. The pinned tool documentation | |
| states that it is typically better not to use the importance matrix for | |
| `output.weight`, which is why collection for that tensor defaults to false. | |
| The commands above are tensor-equivalent reproduction commands. The released | |
| matrix and importance-aware model files also embed path strings. Byte-for-byte | |
| SHA-256 reproduction requires the primary dataset at | |
| `/workspace/ling3/calibration/bartowski-imatrix-v5-semantic.txt`, the extracted | |
| supplement at `/workspace/ling3/calibration/eaddario/combined_all_micro.txt`, | |
| and the final matrix at `/tmp/Ling-3.0-tiny-imatrix.gguf`. The source directory | |
| basename must be `Ling-3.0-tiny`. Different path spellings change metadata | |
| bytes without changing the collected statistics or quantized tensor values. | |
| The released final matrix is 44,016,768 bytes with SHA-256 | |
| `e8b15d131f9ce294f922c5c387f7a69829c12100d6a35bb1635a2b859083c3f0`. | |
| It contains 332 entries from 162 complete chunks, and all 8,832 routed-expert | |
| count slots are nonzero. | |
| ## 4. Quantize directly from BF16 | |
| The matrix is deliberately used for K-quants as well as IQ quants. No output | |
| is requantized from Q8 or another reduced-precision artifact. | |
| ```bash | |
| llama.cpp/build/bin/llama-quantize \ | |
| Ling-3.0-tiny-BF16.gguf Ling-3.0-tiny-Q8_0.gguf Q8_0 32 | |
| for quant in Q6_K Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do | |
| llama.cpp/build/bin/llama-quantize \ | |
| --imatrix Ling-3.0-tiny-imatrix.gguf \ | |
| Ling-3.0-tiny-BF16.gguf "Ling-3.0-tiny-${quant}.gguf" "$quant" 32 | |
| done | |
| ``` | |
| Do not add `--allow-requantize` or `--pure` when reproducing these files. | |
| IQ4_NL and MXFP4_MOE candidates were measured but intentionally rejected; they | |
| are not part of the published artifact set. | |
| ## 5. Held-out PPL and KLD comparison | |
| Validation uses WikiText-2 from `ggml-org/ci` at revision | |
| `927b3642933080f1b0e811e2f916e14c292992f9`. The extracted test file SHA-256 | |
| is `173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08`. | |
| ```bash | |
| hf download ggml-org/ci wikitext-2-raw-v1.zip \ | |
| --repo-type dataset \ | |
| --revision 927b3642933080f1b0e811e2f916e14c292992f9 \ | |
| --local-dir validation | |
| unzip -q validation/wikitext-2-raw-v1.zip -d validation | |
| printf '%s\n' \ | |
| 'ef7edb566e3e2b2d31b29c1fdb0c89a4cc683597484c3dc2517919c615435a11 validation/wikitext-2-raw-v1.zip' \ | |
| '173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08 validation/wikitext-2-raw/wiki.test.raw' \ | |
| | sha256sum -c - | |
| ``` | |
| The archive SHA-256 is | |
| `ef7edb566e3e2b2d31b29c1fdb0c89a4cc683597484c3dc2517919c615435a11`. | |
| First create the BF16 log-probability reference. In this pinned build, | |
| `--kl-divergence-base` is an alias for `--save-all-logits`. | |
| ```bash | |
| llama.cpp/build/bin/llama-perplexity \ | |
| --model Ling-3.0-tiny-BF16.gguf \ | |
| --file validation/wikitext-2-raw/wiki.test.raw \ | |
| --gpu-layers all --ctx-size 512 --batch-size 512 --ubatch-size 512 \ | |
| --chunks 32 --threads 16 --threads-batch 16 --device CUDA0 \ | |
| --split-mode none --main-gpu 0 --fit off --flash-attn off \ | |
| --kv-offload --op-offload --no-repack \ | |
| --cache-type-k f16 --cache-type-v f16 \ | |
| --no-warmup --offline \ | |
| --kl-divergence-base validation/bf16-c512-chunks32.kld | |
| ``` | |
| Then compare each quant using the same context and batch sizes. The comparator | |
| reads the exact tokens and chunk count from the reference file. | |
| ```bash | |
| llama.cpp/build/bin/llama-perplexity \ | |
| --model Ling-3.0-tiny-BF16.gguf \ | |
| --gpu-layers all --ctx-size 512 --batch-size 512 --ubatch-size 512 \ | |
| --threads 16 --threads-batch 16 --device CUDA0 \ | |
| --split-mode none --main-gpu 0 --fit off --flash-attn off \ | |
| --kv-offload --op-offload --no-repack \ | |
| --cache-type-k f16 --cache-type-v f16 \ | |
| --no-warmup --offline \ | |
| --kl-divergence-base validation/bf16-c512-chunks32.kld \ | |
| --kl-divergence \ | |
| > validation/kld-BF16.log 2>&1 | |
| for quant in Q8_0 Q6_K Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do | |
| llama.cpp/build/bin/llama-perplexity \ | |
| --model "Ling-3.0-tiny-${quant}.gguf" \ | |
| --gpu-layers all --ctx-size 512 --batch-size 512 --ubatch-size 512 \ | |
| --threads 16 --threads-batch 16 --device CUDA0 \ | |
| --split-mode none --main-gpu 0 --fit off --flash-attn off \ | |
| --kv-offload --op-offload --no-repack \ | |
| --cache-type-k f16 --cache-type-v f16 \ | |
| --no-warmup --offline \ | |
| --kl-divergence-base validation/bf16-c512-chunks32.kld \ | |
| --kl-divergence \ | |
| > "validation/kld-${quant}.log" 2>&1 | |
| done | |
| python validation/parse_kld.py --output validation/kld-results.json \ | |
| BF16=validation/kld-BF16.log \ | |
| Q8_0=validation/kld-Q8_0.log \ | |
| Q6_K=validation/kld-Q6_K.log \ | |
| Q5_K_M=validation/kld-Q5_K_M.log \ | |
| Q4_K_M=validation/kld-Q4_K_M.log \ | |
| Q4_K_S=validation/kld-Q4_K_S.log \ | |
| IQ4_XS=validation/kld-IQ4_XS.log \ | |
| Q3_K_M=validation/kld-Q3_K_M.log \ | |
| IQ3_M=validation/kld-IQ3_M.log \ | |
| IQ2_M=validation/kld-IQ2_M.log | |
| ``` | |
| Do not combine this KLD workflow with `--ppl-stride`; it uses a different | |
| evaluation path. Comparison logs must be newly truncated, never appended. The | |
| parser requires a complete terminal statistics block because error paths in | |
| this pinned `llama-perplexity` can still return process status zero. | |
| The stored BF16 reference is 2,565,373,716 bytes with SHA-256 | |
| `afa4dc9bd2d995dd85d8614494dc9cccec8e0e4bd6b2c24efc81d3bba96638f7`. | |
| ## 6. Structure, tokenizer, and deterministic generation | |
| Generate a canonical shape map from BF16, then compare every quant against it. | |
| The validator also checks the architecture metadata, Q-LoRA tensors, exact | |
| embedded chat template, NEXTN/MTP absence, matrix provenance, and the exact | |
| filename-bound tensor-type inventory of every published model. Its separate | |
| MXFP4 whitelist is retained for auditing rejected candidates. | |
| ```bash | |
| python validation/check_gguf_structure.py \ | |
| --gguf-dump .venv/bin/gguf-dump \ | |
| --model Ling-3.0-tiny-BF16.gguf \ | |
| --source-template Ling-3.0-tiny/chat_template.jinja \ | |
| --write-reference-shapes validation/bf16-tensor-shapes.json \ | |
| --output validation/structure-BF16.json | |
| for quant in Q8_0 Q6_K Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do | |
| python validation/check_gguf_structure.py \ | |
| --gguf-dump .venv/bin/gguf-dump \ | |
| --model "Ling-3.0-tiny-${quant}.gguf" \ | |
| --source-template Ling-3.0-tiny/chat_template.jinja \ | |
| --reference-shapes validation/bf16-tensor-shapes.json \ | |
| --output "validation/structure-${quant}.json" | |
| done | |
| ``` | |
| Generate the independent Transformers reference only after verifying all 32 | |
| source shards and the pinned control files: | |
| ```bash | |
| python3 -m venv .hfref-venv | |
| .hfref-venv/bin/python -m pip install \ | |
| --extra-index-url https://download.pytorch.org/whl/cu128 \ | |
| accelerate==1.10.1 einops==0.8.1 fla-core==0.5.1 \ | |
| huggingface_hub==0.36.2 numpy==2.1.2 safetensors==0.8.0 \ | |
| tokenizers==0.22.2 'torch==2.8.0+cu128' tqdm==4.70.0 \ | |
| transformers==4.57.6 | |
| .hfref-venv/bin/python validation/hf_reference.py \ | |
| Ling-3.0-tiny validation/hf-reference.json \ | |
| --source-weight-manifest source-safetensors.sha256 | |
| python validation/compare_tokenizer.py \ | |
| --llama-tokenize llama.cpp/build/bin/llama-tokenize \ | |
| --model Ling-3.0-tiny-BF16.gguf \ | |
| --hf-reference validation/hf-reference.json \ | |
| --output validation/tokenizer-comparison.json | |
| ``` | |
| The controlled raw-generation smoke uses the exact same full prompt in both | |
| runtimes. `compare_generation.py` now fails if the supplied prompt differs | |
| from the reference, even when the generated suffix happens to match. | |
| ```bash | |
| prompt='The capital of France is Paris. The capital of Germany is Berlin. The capital of Japan is' | |
| for quant in BF16 Q8_0 Q6_K Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do | |
| llama.cpp/build/bin/llama-completion \ | |
| --model "Ling-3.0-tiny-${quant}.gguf" \ | |
| --prompt "$prompt" --predict 12 \ | |
| --ctx-size 512 --batch-size 512 --ubatch-size 128 \ | |
| --threads 16 --threads-batch 16 --device CUDA0 --gpu-layers all \ | |
| --split-mode none --main-gpu 0 --fit off --flash-attn off \ | |
| --kv-offload --op-offload --no-repack \ | |
| --cache-type-k f16 --cache-type-v f16 \ | |
| --seed 1 --samplers temperature --temperature 0 \ | |
| -no-cnv --no-context-shift --no-warmup --no-display-prompt \ | |
| --color off --offline --log-colors off --no-log-timestamps \ | |
| --check-tensors \ | |
| > "validation/gguf-greedy-${quant}.txt" \ | |
| 2> "validation/load-${quant}.log" | |
| python validation/compare_generation.py \ | |
| --artifact "$quant" \ | |
| --hf-reference validation/hf-reference.json \ | |
| --gguf-output "validation/gguf-greedy-${quant}.txt" \ | |
| --case greedy_capitals_12 --prompt-text "$prompt" \ | |
| --mismatch-policy fail \ | |
| --output "validation/generation-${quant}.json" | |
| done | |
| ``` | |
| ## 7. Multiple-choice collapse and 32K execution screens | |
| ```bash | |
| hf download ikawrakow/validation-datasets-for-llama.cpp \ | |
| mmlu-validation.bin --repo-type dataset \ | |
| --revision 37884b81b4957f1950a53b6ff48d77c8dd5e430c \ | |
| --local-dir validation | |
| printf '%s\n' \ | |
| '470af3a74eccacfaf6f43b08aabf510f61e6c92fe20d17241ded934151e225fa validation/mmlu-validation.bin' \ | |
| | sha256sum -c - | |
| for quant in BF16 Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do | |
| llama.cpp/build/bin/llama-perplexity \ | |
| --model "Ling-3.0-tiny-${quant}.gguf" \ | |
| --file validation/mmlu-validation.bin \ | |
| --multiple-choice --multiple-choice-tasks 500 --seed 1 \ | |
| --gpu-layers all --ctx-size 512 --batch-size 512 --ubatch-size 512 \ | |
| --threads 16 --threads-batch 16 --device CUDA0 \ | |
| --split-mode none --main-gpu 0 --fit off --flash-attn off \ | |
| --kv-offload --op-offload --no-repack \ | |
| --cache-type-k f16 --cache-type-v f16 \ | |
| --no-warmup --offline --no-log-timestamps \ | |
| > "validation/mmlu-${quant}.log" 2>&1 | |
| done | |
| ``` | |
| The pinned tool prints “TruthfulQA” internally, but the supplied binary is the | |
| pinned MMLU validation file. Results are summarized in | |
| `validation/multiple-choice-results.json`. | |
| For BF16, Q4_K_M, and IQ2_M, run the 32K execution/prefill screen as follows: | |
| ```bash | |
| for quant in BF16 Q4_K_M IQ2_M; do | |
| llama.cpp/build/bin/llama-perplexity \ | |
| --model "Ling-3.0-tiny-${quant}.gguf" \ | |
| --file validation/wikitext-2-raw/wiki.test.raw \ | |
| --chunks 1 --ctx-size 32768 --batch-size 4096 --ubatch-size 512 \ | |
| --gpu-layers all --threads 16 --threads-batch 16 --device CUDA0 \ | |
| --split-mode none --main-gpu 0 --fit off --flash-attn off \ | |
| --kv-offload --op-offload --no-repack \ | |
| --cache-type-k f16 --cache-type-v f16 \ | |
| --no-warmup --offline --no-log-timestamps \ | |
| > "validation/context-32768-${quant}.log" 2>&1 | |
| done | |
| ``` | |
| This is not exhaustive validation of the native 131K limit. | |
| ## 8. Server/Jinja and final integrity checks | |
| The exact four OpenAI-compatible server requests and their expected semantic | |
| checks are stored in `validation/server-requests.json` and | |
| `validation/server-results.json`. Start Q4_K_M with `llama-server --jinja`, | |
| submit each request to `/v1/chat/completions`, and verify thinking on/off, | |
| Chinese generation, normal EOS stops, and the two-argument required tool call. | |
| ```bash | |
| llama.cpp/build/bin/llama-server \ | |
| --model Ling-3.0-tiny-Q4_K_M.gguf --alias ling-3.0-tiny \ | |
| --host 127.0.0.1 --port 8080 --jinja --ctx-size 8192 \ | |
| --gpu-layers 999 --split-mode none --main-gpu 0 --device CUDA0 \ | |
| --flash-attn off --no-warmup --offline | |
| # In another terminal after the server reports that it is listening: | |
| python validation/run_server_requests.py \ | |
| --requests validation/server-requests.json \ | |
| --response-dir validation | |
| python validation/validate_server_responses.py \ | |
| --requests validation/server-requests.json \ | |
| --response-dir validation \ | |
| --output validation/server-results.json | |
| ``` | |
| Finally, verify every downloaded release artifact from the repository root: | |
| ```bash | |
| sha256sum -c SHA256SUMS | |
| ``` | |
| `conversion_manifest.json` records the exact source/control hashes, build and | |
| binary hashes, package versions, hardware, calibration order, artifact tensor | |
| inventories, Hub introduction commits, and validation-report bindings. | |