Atom-350M GGUF

GGUF quantizations of TinyModels/Atom-350M, for use with llama.cpp and compatible runtimes.

About

Quantizations were produced by using llama.cpp build b10090, via convert_hf_to_gguf.py followed by llama-quantize. Q8_0 and Q6_K are static (non-imatrix) quantizations; Q5_K_M and Q4_K_M use an importance matrix computed with llama-imatrix over 200 chunks of wikitext-2 (wiki.train.raw).

Files

Filename Quant Type Size Description
Atom-350M-Q4_K_M.gguf Q4_K_M 256 MB Balanced, recommended default (imatrix)
Atom-350M-Q5_K_M.gguf Q5_K_M 275 MB Low quality loss (imatrix)
Atom-350M-Q6_K.gguf Q6_K 349 MB Very low quality loss (see note below โ€” many tensors fall back to Q8_0)
Atom-350M-Q8_0.gguf Q8_0 367 MB Near-lossless
Atom-350M-F16.gguf F16 692 MB Full precision reference

Note on K-quants: this model's tensor dimensions (960, 320, 2560) are not evenly divisible by the 256-element superblock size K-quants require. As a result, 192 of 290 tensors in the Q6_K file fall back to Q8_0 quantization, which is why Q6_K is barely smaller than Q8_0. Expect similar (lesser) fallback behavior in the upcoming Q5_K_M/Q4_K_M builds.

Quantization Comparison

Perplexity and KL-divergence measured on wikitext-2 (wiki.test.raw), 20 chunks, n_ctx=512, using llama-perplexity --kl-divergence against saved F16 reference logits. Tokens/sec measured with llama-bench (pp512/tg128, 4 threads) on Apple Macbook Air M2.

Perplexity & speed

Quant Size Perplexity ฮ”PPL vs F16 Prompt (t/s) Gen (t/s) Notes
F16 692 MB 17.5203 โ€” 1625 51.9 reference
Q8_0 367 MB 17.5818 +0.35% 1664 63.5 near-lossless
Q6_K 349 MB 17.6377 +0.67% 1560 64.0 mostly Q8_0 under the hood, see note above
Q5_K_M 275 MB 17.6334 +0.65% 1359 73.1 imatrix-guided
Q4_K_M 256 MB 17.8882 +2.10% 1555 77.0 imatrix-guided, recommended

KL-divergence vs F16

Mean PPL can mask token-level disagreement โ€” two quants can tie on PPL while differing in the full predicted distribution. KLD/ฮ”p are measured per-token; "Same top token" is the share of tokens where the quant agrees with F16 on the single most-likely next token.

Quant Mean KLD Median KLD 99th %ile KLD RMS ฮ”p Same top token
Q8_0 0.00118 0.00082 0.00753 0.891% 97.94%
Q6_K 0.00472 0.00281 0.03898 1.706% 96.22%
Q5_K_M 0.01250 0.00709 0.10824 2.815% 93.96%
Q4_K_M 0.02842 0.01685 0.24259 4.131% 90.92%

Notably, Q5_K_M and Q6_K nearly tie on mean PPL (17.63 vs 17.64 above) but not on KLD โ€” Q6_K is ~2.6x lower and has better top-token agreement, so it's the better pick when distributional fidelity matters more than the near-tie in PPL suggests.

Q8_0/Q6_K are static quantizations; Q5_K_M/Q4_K_M use an imatrix computed from 200 chunks of wikitext-2 wiki.train.raw. The same tensor-shape fallback noted above applies to all four.

Which Quant Should I Choose?

  • Q4_K_M โ€” default pick. Smallest, fastest generation, +2.1% PPL vs F16. Highest KLD (0.028 mean); skip if you need tight distributional fidelity.
  • Q5_K_M โ€” 275MB middle ground if Q4_K_M's KLD is too high, but Q6_K generally wins.
  • Q6_K โ€” best top-token agreement short of Q8_0 (96.2%), though the tensor-shape fallback puts it near Q8_0's size. Worth the ~75MB over Q5_K_M when accuracy matters.
  • Q8_0 โ€” safest, closest to FP16 (97.9% top-token, lowest KLD) if you have ~400MB to spare.
  • F16 โ€” reference baseline, or if you'll re-quantize/fine-tune.

How to Run

llama.cpp

llama-cli -m Atom-350M-Q8_0.gguf -p "prompt" -n 128

LM Studio

Search "Atom-350M GGUF", download your preferred quant, load it.

Ollama

ollama run hf.co/syaffers/Atom-350M-GGUF:Q8_0

text-generation-webui

Drop the .gguf file into your models/ folder and select the llama.cpp (GGUF) loader.

Prompt Format

ChatML:

<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
Downloads last month
239
GGUF
Model size
0.4B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for syaffers/Atom-350M-GGUF

Quantized
(2)
this model