How to Run d1-3B Locally

Built from Liquid AI's published weights with our own importance matrix, and measured on decisions, not just text.

Atomic Chat Discord GitHub

d1-3B is Liquid AI's decision model. It takes a state and answers typed questions in one forward pass: yes or no, a pick from named options, or a score. The answer is read from the model's distribution over the options' tokens, so nothing is generated.

Pick a file

Every number below is measured on one machine. The raw results and logs are in the metrics repo.

  • same answer: across 1,285 decisions, how often the file gives the same answer as the BF16 weights published in LiquidAI/d1-3B. The questions are yes/no, choice and score questions over held-out texts in 30 languages and source code. The number of changed answers is in brackets.
  • option drift: the mean total variation distance between the file's option probabilities and the original's. 0 means identical.
  • KLD and top-1: full-vocabulary agreement with the original on held-out text.
File Size same answer option drift KLD top-1
BF16 5,403 MB on disk reference 0 0 100%
Q8_0 2,875 MB on disk 99.3% (9) 0.0033 0.0010 98.28%
AD-Q6_K 2,348 MB on disk 99.1% (11) 0.0053 0.0022 97.41%
AD-Q5_K_M 1,953 MB on disk 98.1% (25) 0.0112 0.0080 95.09%
AD-Q4_K_M 1,658 MB on disk 97.1% (37) 0.0178 0.0244 91.54%
AD-IQ4_XS 1,570 MB on disk 95.6% (57) 0.0203 0.0283 90.92%

AD- means Atomic Dynamic: the type is chosen per tensor instead of taken from a llama.cpp preset.

  • The token table, which d1-3B also uses as its output head, never goes below Q6_K.
  • Attention stays at Q8_0 or Q6_K.
  • ffn_down in the first and last four blocks takes one step more than the middle.

The projector for images comes as mmproj-d1-3B-BF16 and mmproj-d1-3B-Q8_0. We did not measure image decisions.

How these compare to other GGUFs

  • Liquid's own GGUFs are built from the current weights and carry the type lfm2-d1. Their BF16 has the same tensors as ours.
  • prithivMLmods/d1-3B-GGUF has llama.cpp's stock presets, also from the current weights.

We measured all of them on one machine (an NVIDIA A10, llama.cpp 88dcc46) against our BF16. This run is separate from the table above, so our own rows differ from it slightly.

File Size same answer option drift option KL KLD top-1
AtomicChat Q8_0 2,875 MB 99.2% (10) 0.0034 0.0001 0.0009 98.30%
Liquid Q8_0 2,875 MB 99.4% (8) 0.0034 0.0001 0.0009 98.30%
AtomicChat AD-Q6_K 2,348 MB 99.3% (9) 0.0054 0.0002 0.0022 97.41%
prithivMLmods Q6_K 2,222 MB 98.7% (17) 0.0077 0.0004 0.0041 96.34%
AtomicChat AD-Q5_K_M 1,953 MB 97.7% (29) 0.0112 0.0009 0.0081 95.11%
prithivMLmods Q5_K_M 1,940 MB 97.7% (29) 0.0138 0.0014 0.0121 93.89%
AtomicChat AD-Q4_K_M 1,658 MB 97.0% (38) 0.0174 0.0023 0.0243 91.55%
Liquid Q4_K_M 1,674 MB 95.7% (55) 0.0270 0.0053 0.0392 89.42%
prithivMLmods Q4_K_M 1,674 MB 95.9% (53) 0.0272 0.0053 0.0389 89.44%
  • 4 bits. Our AD-Q4_K_M is 16 MB smaller than the two stock Q4_K_M files.
    • It changes 38 answers instead of 55.
    • Its option KL is less than half of theirs, and its KLD is 38% lower.
  • 5 bits. At about the same size as prithivMLmods' Q5_K_M, ours changes the same number of answers, with about a third less option KL and KLD.
  • 6 bits. Our AD-Q6_K is 126 MB larger than prithivMLmods' Q6_K, so the two are not a same-size comparison.
  • Q8_0. The two Q8_0 files measure the same.
    • A few tensors round differently, because Liquid's converter writes Q8_0 itself.
    • 8 against 10 changed answers is within run-to-run noise. Our own Q8_0 changed 9 in the run above.

Liquid's first GGUFs, published on 6 October, were built from the previous version of the weights. Liquid replaced them on 7 October. Our measurement of those first files is in the metrics repo.

Running it

This needs a llama.cpp build that includes #30110: commit 88dcc46, merged on 7 October, or newer. Older builds do not know the type lfm2-d1 and refuse to load the file.

llama-server -m d1-3B-AD-Q4_K_M.gguf --mmproj mmproj-d1-3B-Q8_0.gguf -ngl 99 -c 8192

Ask through the /v1/systemone endpoint. It takes a state and named, typed questions, and returns each answer with its option probabilities:

curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d '{
  "state": "I was charged twice this month, please refund one of them.",
  "questions": {
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
                          "fraud": "Suspected unauthorised use"}},
    "angry": {"type": "noul", "instructions": "Is the customer angry?"}
  }
}'

Response from AD-Q4_K_M, numbers rounded:

{
  "answers": {
    "team": {"type": "choice", "choice": "billing",
             "probabilities": {"billing": 0.984, "technical": 0.011, "fraud": 0.005}, "confidence": 0.977},
    "angry": {"type": "noul", "noul": 0.099}
  },
  "usage": {"input_tokens": 97, "output_tokens": 0}
}
  • noul questions are yes or no, and score questions take a list of 2 to 10 levels.
  • Images go in an images array as data URLs.
  • The full request format is in the server documentation, and the question schema in the original card.

Checked on the merged endpoint. We ran the PR's final commit against Liquid's own PyTorch code in FP32, on the same 1,285 text decisions:

  • Q8_0 changes 6 answers, with an option KL of 0.0001.
  • AD-Q4_K_M changes 44, with an option KL of 0.0022.
  • These numbers use a different reference and a different readout from the table above, so they do not match it exactly.

How we made and measured them

  • BF16: converted from LiquidAI/d1-3B (revision da1fe36) with llama.cpp 18b5f8b. It matches Liquid's PyTorch code on 300 decisions with a mean option KL of 0.0002. The 9 changed answers there are near-ties, with a largest option difference of 0.035. The projector is tensor-for-tensor identical to Liquid's.
  • Importance matrix: 1.5M tokens of decision prompts, 3,206 in all. The states come from the calibration corpora pool (Wikipedia in 30 languages, code, structured files). Each carries yes/no, choice and score questions and is rendered with the model's own prompt code.
  • Decisions: 1,285 questions over the held-out calib-corpora eval/neutral and eval/code texts, none of them in the calibration set. Before the endpoint existed, they were answered from the option tokens' log-probabilities on /completion: the prompt is rendered with Liquid's prompt.py, and the readout is a softmax over the options' tokens. The scripts are in the metrics repo.
  • Text: llama-perplexity --kl-divergence over eval/neutral, 93 chunks of 4,096 tokens.
  • Setup: an NVIDIA H100 and llama.cpp 18b5f8b for every file.
  • Metadata: on 7 October, after #30110 was merged, the files were re-stamped. The text files got lfm2.decision.type = lfm2-d1 and the systemone template, both taken from a BF16 file the PR's converter wrote. The projectors got clip.vision.image_resize_algo = bicubic. No tensor changed.

Model details

  • Base: LiquidAI/d1-3B, a decision model post-trained from LFM2.5-VL-3B: 3.1B parameters, 32K context, a 128K vocabulary and a SigLIP2 vision encoder.
  • License: LFM Open License v1.0; see LICENSE.
Downloads last month
-
GGUF
Model size
3B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AtomicChat/d1-3B-GGUF

Finetuned
LiquidAI/d1-3B
Quantized
(13)
this model

Collection including AtomicChat/d1-3B-GGUF