Image-Text-to-Text
Transformers
Safetensors
lfm2_vl
liquid
lfm2.5
edge
decision
classification
calibration
system-one
multimodal
decision-model
conversational
custom_code
Instructions to use LiquidAI/d1-3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use LiquidAI/d1-3B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="LiquidAI/d1-3B", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True) model = AutoModelForMultimodalLM.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True, device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use LiquidAI/d1-3B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "LiquidAI/d1-3B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LiquidAI/d1-3B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/LiquidAI/d1-3B
- SGLang
How to use LiquidAI/d1-3B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "LiquidAI/d1-3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LiquidAI/d1-3B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "LiquidAI/d1-3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LiquidAI/d1-3B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use LiquidAI/d1-3B with Docker Model Runner:
docker model run hf.co/LiquidAI/d1-3B
Update README.md
#2
by iamleonie - opened
- README.md +75 -101
- assets/d1-3b-smash.mp4 +3 -0
- assets/di_pareto.png +3 -0
- assets/ood_tasks.mp4 +3 -0
- config.json +1 -1
- lfm2_vl.py +4 -4
- model.safetensors +1 -1
- modeling_d1.py → modeling_lfm_jev.py +1 -1
- prompt.py +48 -9
- runner.py +13 -7
README.md
CHANGED
|
@@ -31,7 +31,6 @@ tags:
|
|
| 31 |
- calibration
|
| 32 |
- system-one
|
| 33 |
- multimodal
|
| 34 |
-
- decision-model
|
| 35 |
---
|
| 36 |
|
| 37 |
<div align="center">
|
|
@@ -50,22 +49,21 @@ tags:
|
|
| 50 |
# d1-3B
|
| 51 |
|
| 52 |
d1-3B is a 3B parameter **decision model** built on [LFM2.5-VL-3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B).
|
| 53 |
-
You give it a state (text, JSON, images, or a mix) and a set of questions. It returns calibrated,
|
| 54 |
-
typed answers in **one forward pass with zero output tokens**
|
|
|
|
| 55 |
|
| 56 |
-
- **Best decision model under 10B on the Decision Index 0.2.1
|
| 57 |
-
and
|
| 58 |
-
- **Multimodal**: images and text in the same state. It scores
|
| 59 |
(LFM2.5-VL-3B: 73.9).
|
| 60 |
- **Fast**: 8 ms a decision on an NVIDIA RTX 4090, 9 ms on an AMD MI325X, 30 ms on an Apple M5 Pro.
|
| 61 |
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-

|
| 65 |
|
| 66 |
> [!NOTE]
|
| 67 |
> 💻 **Demos**: Try d1-3B in a Hugging Face space without any setup:
|
| 68 |
-
>
|
| 69 |
|
| 70 |
|
| 71 |
## 🗒️ Model Details
|
|
@@ -73,14 +71,16 @@ Find more information about open d1 in our [blog post](https://www.liquid.ai/blo
|
|
| 73 |
| Model | Parameters | Description |
|
| 74 |
|---|---|---|
|
| 75 |
| [LFM2.5-VL-3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B) | 3.1B | General-purpose vision-language model (base) |
|
| 76 |
-
| **[d1-3B](https://huggingface.co/LiquidAI/d1-
|
| 77 |
|
| 78 |
d1-3B is a multimodal decision model with the following features:
|
| 79 |
|
| 80 |
- **Total parameters**: 3.12B
|
|
|
|
| 81 |
- **Vision encoder**: SigLIP2 NaFlex shape-optimized 400M
|
| 82 |
- **Context length**: 32,768 tokens
|
| 83 |
- **Vocabulary size**: 128,000
|
|
|
|
| 84 |
|
| 85 |
We recommend d1-3B wherever a pipeline needs a yes/no, a pick from named options, or a rating:
|
| 86 |
routing and triage, moderation, intent and topic classification, extraction checks, reranking, LLM-as-a-judge
|
|
@@ -97,47 +97,36 @@ pip install "transformers>=5.14" torch torchvision pillow
|
|
| 97 |
The model ships its own code, so load it with `trust_remote_code=True`:
|
| 98 |
|
| 99 |
```python
|
|
|
|
|
|
|
|
|
|
| 100 |
import torch
|
|
|
|
| 101 |
from transformers import AutoModel
|
| 102 |
-
from transformers.image_utils import load_image
|
| 103 |
|
| 104 |
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
|
| 105 |
-
|
| 106 |
-
|
| 107 |
|
| 108 |
-
#
|
| 109 |
questions = {
|
| 110 |
-
"refund": {
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
| 114 |
-
"
|
| 115 |
-
|
| 116 |
-
"instructions": "Which team should handle this?",
|
| 117 |
-
"criteria": {
|
| 118 |
-
"billing": "Charges, refunds, invoices",
|
| 119 |
-
"technical": "App or site faults",
|
| 120 |
-
"fraud": "Suspected unauthorised use",
|
| 121 |
-
},
|
| 122 |
-
},
|
| 123 |
-
"urgency": {
|
| 124 |
-
"type": "score",
|
| 125 |
-
"instructions": "How urgent is this?",
|
| 126 |
-
"criteria": ["Can wait", "Today", "Blocking the customer now"],
|
| 127 |
-
},
|
| 128 |
}
|
| 129 |
print(model.system_one("I was charged twice this month, please refund one of them.", questions))
|
| 130 |
|
| 131 |
-
#
|
| 132 |
-
|
| 133 |
-
|
| 134 |
-
|
| 135 |
-
|
| 136 |
-
|
| 137 |
-
}
|
| 138 |
-
print(model.system_one(None, {"cats": cats}, images=[image]))
|
| 139 |
|
| 140 |
-
#
|
| 141 |
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
|
| 142 |
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))
|
| 143 |
```
|
|
@@ -163,32 +152,16 @@ Each call returns `{"answers": {name: answer}, "usage": {"input_tokens": n, "out
|
|
| 163 |
|
| 164 |
## ⚡ Speed
|
| 165 |
|
| 166 |
-
Warm calls
|
| 167 |
-
state and a 384 px image. The last column is throughput with 64 states packed into one pass.
|
| 168 |
-
|
| 169 |
-
### Edge Inference
|
| 170 |
-
|
| 171 |
-
We measure latency on an Apple M5 Pro and, in collaboration with NVIDIA, on an NVIDIA Jetson AGX Thor,
|
| 172 |
-
a Jetson AGX Orin 64 GB and a Jetson Orin Nano.
|
| 173 |
|
| 174 |
| | one question | 3 questions, one pass | 3.4k-token state | 384 px image | 64 states, packed |
|
| 175 |
|---|---:|---:|---:|---:|---:|
|
| 176 |
-
|
|
| 177 |
-
|
|
| 178 |
-
|
|
| 179 |
-
| NVIDIA Jetson Orin Nano | 50 ms | 73 ms | 1640 ms | 202 ms | 38 / s |
|
| 180 |
|
| 181 |
-
|
| 182 |
-
|
| 183 |
-
We measure latency on an NVIDIA RTX 4090 and an AMD MI325X, in bf16, median of 20 runs.
|
| 184 |
-
|
| 185 |
-
| | one question | 3 questions, one pass | 3.4k-token state | 384 px image | 64 states, packed |
|
| 186 |
-
|---|---:|---:|---:|---:|---:|
|
| 187 |
-
| NVIDIA RTX 4090 | 8 ms | 21 ms | 102 ms | 17 ms | 475 / s |
|
| 188 |
-
| AMD MI325X | 9 ms | 14 ms | 44 ms | 18 ms | 1,106 / s |
|
| 189 |
-
|
| 190 |
-
On NVIDIA GPUs, `model.compile(mode="reduce-overhead")` runs single questions as CUDA graphs (the RTX 4090
|
| 191 |
-
row uses it). Without it, a single question takes 16 ms. The first call with a new shape pays for kernel
|
| 192 |
selection or compilation, so warm up the shapes you serve.
|
| 193 |
|
| 194 |
## 📊 Performance
|
|
@@ -197,12 +170,15 @@ All results are on public benchmarks.
|
|
| 197 |
|
| 198 |
### Decision Index 0.2.1
|
| 199 |
|
| 200 |
-
|
|
|
|
|
|
|
|
|
|
| 201 |
|
| 202 |
| Model | Size | Decision Index | Knowledge | Language | Retrieval | Tools | Arts |
|
| 203 |
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 204 |
| Winnow-12B | 12B | 50.02 | 33.8 | 56.0 | 54.0 | 71.0 | 30.0 |
|
| 205 |
-
| **d1-3B** | **3B** | **
|
| 206 |
| Decider 35B-A3B | 36B | 47.11 | 31.8 | 55.5 | 54.7 | 56.5 | 32.6 |
|
| 207 |
| JPT-9B | 9.7B | 46.89 | 31.7 | 56.7 | 44.6 | 67.0 | 28.6 |
|
| 208 |
| Decision 1.0 Lux | 9.7B | 43.49 | 30.9 | 48.0 | 50.0 | 57.2 | 26.4 |
|
|
@@ -214,22 +190,23 @@ We scored d1-3B with the official scorer (not a leaderboard submission). All oth
|
|
| 214 |
|
| 215 |
### Benchmarks as decisions
|
| 216 |
|
| 217 |
-
|
|
|
|
| 218 |
|
| 219 |
| Benchmark | d1-3B | Decider 4B | Decider 2B |
|
| 220 |
|---|---:|---:|---:|
|
| 221 |
-
| SQuAD 2.0 | **
|
| 222 |
-
| Civil Comments | 93.
|
| 223 |
-
| MASSIVE intent |
|
| 224 |
-
| HelpSteer2 |
|
| 225 |
-
| PubMedQA | **
|
| 226 |
-
| BoolQ | 86.
|
| 227 |
-
| XNLI | 85.
|
| 228 |
-
| PAWS-X | **76.
|
| 229 |
-
| **Mean** | **77.
|
| 230 |
-
|
| 231 |
-
d1-3B also scores
|
| 232 |
-
all 23,900 rows) and
|
| 233 |
(dev split).
|
| 234 |
|
| 235 |
### Vision
|
|
@@ -239,22 +216,29 @@ compared with the base model:
|
|
| 239 |
|
| 240 |
| Benchmark | d1-3B | LFM2.5-VL-3B |
|
| 241 |
|---|---:|---:|
|
| 242 |
-
| AI2D |
|
| 243 |
-
| BLINK |
|
| 244 |
-
| CV-Bench | 82.
|
| 245 |
-
| HallusionBench |
|
| 246 |
-
| MMBench | 84.
|
| 247 |
-
| MME |
|
| 248 |
-
| MMStar | 59.
|
| 249 |
-
| MMVP |
|
| 250 |
| POPE | 88.5 | 90.1 |
|
| 251 |
-
| VisualWebBench | 71.
|
| 252 |
-
| VL-RewardBench |
|
| 253 |
-
| **Mean** | **
|
| 254 |
-
| [ImajevBench](https://huggingface.co/datasets/mohit67890/imajev-bench) (dev and calibration, 253 rows) |
|
| 255 |
|
| 256 |
With the images removed, the same questions score 45.1, so the answers come from the images.
|
| 257 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 258 |
## 📬 Contact
|
| 259 |
|
| 260 |
- Got questions or want to connect? [Join our Discord community](https://discord.com/invite/liquid-ai)
|
|
@@ -262,16 +246,6 @@ With the images removed, the same questions score 45.1, so the answers come from
|
|
| 262 |
|
| 263 |
## Citation
|
| 264 |
|
| 265 |
-
```bibtex
|
| 266 |
-
@article{liquidAI2026opend1,
|
| 267 |
-
author = {Liquid AI},
|
| 268 |
-
title = {Open d1: Edge decision models for text, vision, and audio},
|
| 269 |
-
journal = {Liquid AI Blog},
|
| 270 |
-
year = {2026},
|
| 271 |
-
note = {https://www.liquid.ai/blog/d1-open},
|
| 272 |
-
}
|
| 273 |
-
```
|
| 274 |
-
|
| 275 |
```bibtex
|
| 276 |
@article{liquidai2025lfm2,
|
| 277 |
title = {LFM2 Technical Report},
|
|
|
|
| 31 |
- calibration
|
| 32 |
- system-one
|
| 33 |
- multimodal
|
|
|
|
| 34 |
---
|
| 35 |
|
| 36 |
<div align="center">
|
|
|
|
| 49 |
# d1-3B
|
| 50 |
|
| 51 |
d1-3B is a 3B parameter **decision model** built on [LFM2.5-VL-3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B).
|
| 52 |
+
You give it a state (text, JSON, images, or a mix) and a set of named questions. It returns calibrated,
|
| 53 |
+
typed answers in **one forward pass with zero output tokens**: every answer is read directly from the
|
| 54 |
+
model's distribution over the options, with no generation and no parsing.
|
| 55 |
|
| 56 |
+
- **Best decision model under 10B**: 47.12 on the Decision Index 0.2.1. That is ahead of every 4B and 9B
|
| 57 |
+
model on the public leaderboard and level with Decider 35B-A3B (47.11).
|
| 58 |
+
- **Multimodal**: images and text in the same state. It scores 73.7 on 11 public image benchmarks
|
| 59 |
(LFM2.5-VL-3B: 73.9).
|
| 60 |
- **Fast**: 8 ms a decision on an NVIDIA RTX 4090, 9 ms on an AMD MI325X, 30 ms on an Apple M5 Pro.
|
| 61 |
|
| 62 |
+

|
|
|
|
|
|
|
| 63 |
|
| 64 |
> [!NOTE]
|
| 65 |
> 💻 **Demos**: Try d1-3B in a Hugging Face space without any setup:
|
| 66 |
+
> **TODO: add demo**
|
| 67 |
|
| 68 |
|
| 69 |
## 🗒️ Model Details
|
|
|
|
| 71 |
| Model | Parameters | Description |
|
| 72 |
|---|---|---|
|
| 73 |
| [LFM2.5-VL-3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B) | 3.1B | General-purpose vision-language model (base) |
|
| 74 |
+
| **[d1-3B](https://huggingface.co/LiquidAI/d1-3b-RC)** | 3.1B | Post-trained for single-pass, calibrated decisions |
|
| 75 |
|
| 76 |
d1-3B is a multimodal decision model with the following features:
|
| 77 |
|
| 78 |
- **Total parameters**: 3.12B
|
| 79 |
+
- **LM backbone**: LFM2.5-2.6B, 30 layers (22 double-gated short convolution blocks + 8 GQA)
|
| 80 |
- **Vision encoder**: SigLIP2 NaFlex shape-optimized 400M
|
| 81 |
- **Context length**: 32,768 tokens
|
| 82 |
- **Vocabulary size**: 128,000
|
| 83 |
+
- **Output**: no generated tokens. Each answer is a probability distribution over the question's options.
|
| 84 |
|
| 85 |
We recommend d1-3B wherever a pipeline needs a yes/no, a pick from named options, or a rating:
|
| 86 |
routing and triage, moderation, intent and topic classification, extraction checks, reranking, LLM-as-a-judge
|
|
|
|
| 97 |
The model ships its own code, so load it with `trust_remote_code=True`:
|
| 98 |
|
| 99 |
```python
|
| 100 |
+
import io
|
| 101 |
+
import urllib.request
|
| 102 |
+
|
| 103 |
import torch
|
| 104 |
+
from PIL import Image
|
| 105 |
from transformers import AutoModel
|
|
|
|
| 106 |
|
| 107 |
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
|
| 108 |
+
model = AutoModel.from_pretrained("LiquidAI/d1-3b-RC", trust_remote_code=True,
|
| 109 |
+
dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)
|
| 110 |
|
| 111 |
+
# Several named questions over one text state, answered in one pass
|
| 112 |
questions = {
|
| 113 |
+
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
|
| 114 |
+
"team": {"type": "choice", "instructions": "Which team should handle this?",
|
| 115 |
+
"criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
|
| 116 |
+
"fraud": "Suspected unauthorised use"}},
|
| 117 |
+
"urgency": {"type": "score", "instructions": "How urgent is this?",
|
| 118 |
+
"criteria": ["Can wait", "Today", "Blocking the customer now"]},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 119 |
}
|
| 120 |
print(model.system_one("I was charged twice this month, please refund one of them.", questions))
|
| 121 |
|
| 122 |
+
# An image as the whole state
|
| 123 |
+
url = "http://images.cocodataset.org/val2017/000000039769.jpg" # two cats on a sofa
|
| 124 |
+
photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))
|
| 125 |
+
print(model.system_one(None, {"cats": {"type": "choice", "instructions": "How many cats are there?",
|
| 126 |
+
"criteria": {"one": "One", "two": "Two", "more": "Three or more"}}},
|
| 127 |
+
images=[photo]))
|
|
|
|
|
|
|
| 128 |
|
| 129 |
+
# Many requests, packed together with no padding
|
| 130 |
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
|
| 131 |
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))
|
| 132 |
```
|
|
|
|
| 152 |
|
| 153 |
## ⚡ Speed
|
| 154 |
|
| 155 |
+
Warm calls in bf16, median of 20 runs:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 156 |
|
| 157 |
| | one question | 3 questions, one pass | 3.4k-token state | 384 px image | 64 states, packed |
|
| 158 |
|---|---:|---:|---:|---:|---:|
|
| 159 |
+
| NVIDIA RTX 4090 | **8.0 ms** | 21 ms | 102 ms | 17 ms | 475 / s |
|
| 160 |
+
| AMD MI325X | 9.2 ms | 14 ms | 44 ms | 18 ms | 1,106 / s |
|
| 161 |
+
| Apple M5 Pro (`mps`) | 30 ms | 41 ms | 0.64 s | 62 ms | 78 / s |
|
|
|
|
| 162 |
|
| 163 |
+
On NVIDIA GPUs, `model.compile(mode="reduce-overhead")` runs single questions as CUDA graphs. The RTX 4090
|
| 164 |
+
row uses it; without it a single question takes 16 ms. The first call with a new shape pays for kernel
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 165 |
selection or compilation, so warm up the shapes you serve.
|
| 166 |
|
| 167 |
## 📊 Performance
|
|
|
|
| 170 |
|
| 171 |
### Decision Index 0.2.1
|
| 172 |
|
| 173 |
+
The [Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index) casts 38 public
|
| 174 |
+
benchmarks as 120,226 typed decisions, scored as chance-corrected skill. We scored d1-3B with the official
|
| 175 |
+
scorer; it is not a leaderboard submission. All other rows come from the public leaderboard
|
| 176 |
+
(snapshot 2026-09-28). Sizes are served parameters.
|
| 177 |
|
| 178 |
| Model | Size | Decision Index | Knowledge | Language | Retrieval | Tools | Arts |
|
| 179 |
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 180 |
| Winnow-12B | 12B | 50.02 | 33.8 | 56.0 | 54.0 | 71.0 | 30.0 |
|
| 181 |
+
| **d1-3B** | **3B** | **47.12** | 27.8 | 53.5 | 52.2 | 66.7 | **34.6** |
|
| 182 |
| Decider 35B-A3B | 36B | 47.11 | 31.8 | 55.5 | 54.7 | 56.5 | 32.6 |
|
| 183 |
| JPT-9B | 9.7B | 46.89 | 31.7 | 56.7 | 44.6 | 67.0 | 28.6 |
|
| 184 |
| Decision 1.0 Lux | 9.7B | 43.49 | 30.9 | 48.0 | 50.0 | 57.2 | 26.4 |
|
|
|
|
| 190 |
|
| 191 |
### Benchmarks as decisions
|
| 192 |
|
| 193 |
+
Evaluation splits are cast as typed decisions, with the same rows for every model. These scores follow the
|
| 194 |
+
decision protocol, not each benchmark's own.
|
| 195 |
|
| 196 |
| Benchmark | d1-3B | Decider 4B | Decider 2B |
|
| 197 |
|---|---:|---:|---:|
|
| 198 |
+
| SQuAD 2.0 | **83.3** | 76.0 | 67.7 |
|
| 199 |
+
| Civil Comments | 93.3 | 92.8 | **93.6** |
|
| 200 |
+
| MASSIVE intent | 86.9 | **88.3** | 81.1 |
|
| 201 |
+
| HelpSteer2 | 38.0 | **42.0** | 32.0 |
|
| 202 |
+
| PubMedQA | **68.3** | 63.3 | 65.7 |
|
| 203 |
+
| BoolQ | 86.3 | **89.0** | 87.3 |
|
| 204 |
+
| XNLI | 85.6 | **88.6** | 85.0 |
|
| 205 |
+
| PAWS-X | **76.4** | 69.8 | 59.5 |
|
| 206 |
+
| **Mean** | **77.3** | 76.2 | 71.5 |
|
| 207 |
+
|
| 208 |
+
d1-3B also scores 69.5 on [DecisionBench](https://huggingface.co/datasets/Hanno-Labs/decision-bench) (eng v1,
|
| 209 |
+
all 23,900 rows) and 68.8 on [Fast Decisions](https://huggingface.co/datasets/fastino/fast-decisions)
|
| 210 |
(dev split).
|
| 211 |
|
| 212 |
### Vision
|
|
|
|
| 216 |
|
| 217 |
| Benchmark | d1-3B | LFM2.5-VL-3B |
|
| 218 |
|---|---:|---:|
|
| 219 |
+
| AI2D | 80.1 | 80.9 |
|
| 220 |
+
| BLINK | 58.2 | 58.7 |
|
| 221 |
+
| CV-Bench | 82.7 | 87.6 |
|
| 222 |
+
| HallusionBench | 63.8 | 65.0 |
|
| 223 |
+
| MMBench | 84.3 | 84.3 |
|
| 224 |
+
| MME | 81.7 | 82.4 |
|
| 225 |
+
| MMStar | 59.4 | 61.2 |
|
| 226 |
+
| MMVP | 76.3 | 73.7 |
|
| 227 |
| POPE | 88.5 | 90.1 |
|
| 228 |
+
| VisualWebBench | 71.6 | 78.3 |
|
| 229 |
+
| VL-RewardBench | 63.6 | 50.9 |
|
| 230 |
+
| **Mean** | **73.7** | 73.9 |
|
| 231 |
+
| [ImajevBench](https://huggingface.co/datasets/mohit67890/imajev-bench) (dev and calibration, 253 rows) | 66.8 | 66.8 |
|
| 232 |
|
| 233 |
With the images removed, the same questions score 45.1, so the answers come from the images.
|
| 234 |
|
| 235 |
+
## ⚠️ Limitations
|
| 236 |
+
|
| 237 |
+
- **Knowledge-heavy questions**: larger models stay ahead on hard knowledge (for example MMLU-Pro and GPQA).
|
| 238 |
+
- **Thresholds**: calibrate any threshold on data from your own deployment before gating on it.
|
| 239 |
+
- **No text output**: d1-3B answers questions over fixed options. Use [LFM2.5-VL-3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B)
|
| 240 |
+
for open-ended generation.
|
| 241 |
+
|
| 242 |
## 📬 Contact
|
| 243 |
|
| 244 |
- Got questions or want to connect? [Join our Discord community](https://discord.com/invite/liquid-ai)
|
|
|
|
| 246 |
|
| 247 |
## Citation
|
| 248 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 249 |
```bibtex
|
| 250 |
@article{liquidai2025lfm2,
|
| 251 |
title = {LFM2 Technical Report},
|
assets/d1-3b-smash.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:08f69b1a67bf32cd18fad32b5872519b580be9e94bafbe26a13d5cf0aaa49c1f
|
| 3 |
+
size 53062187
|
assets/di_pareto.png
ADDED
|
Git LFS Details
|
assets/ood_tasks.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:b491dfa851157e12ba2bee33a394c9565408247cf0004e493730622c62878e42
|
| 3 |
+
size 8931625
|
config.json
CHANGED
|
@@ -3,7 +3,7 @@
|
|
| 3 |
"Lfm2VlForConditionalGeneration"
|
| 4 |
],
|
| 5 |
"auto_map": {
|
| 6 |
-
"AutoModel": "
|
| 7 |
},
|
| 8 |
"bos_token_id": 124894,
|
| 9 |
"do_image_splitting": true,
|
|
|
|
| 3 |
"Lfm2VlForConditionalGeneration"
|
| 4 |
],
|
| 5 |
"auto_map": {
|
| 6 |
+
"AutoModel": "modeling_lfm_jev.LfmJevModel"
|
| 7 |
},
|
| 8 |
"bos_token_id": 124894,
|
| 9 |
"do_image_splitting": true,
|
lfm2_vl.py
CHANGED
|
@@ -1,9 +1,9 @@
|
|
| 1 |
"""LFM2-VL and LFM2 on the hybrid stack.
|
| 2 |
|
| 3 |
HF's classes stay the frame (the SigLIP2 tower, the projector, the image
|
| 4 |
-
features, loading and saving); only the language model is
|
| 5 |
-
|
| 6 |
-
of questions runs as one `hybrid.Tree`,
|
| 7 |
implementation.
|
| 8 |
"""
|
| 9 |
|
|
@@ -46,7 +46,7 @@ class Attention(nn.Module):
|
|
| 46 |
|
| 47 |
|
| 48 |
class ShortConv(nn.Module):
|
| 49 |
-
"""`out(C * conv(B * x))`; HF's `nn.Conv1d` holds the taps,
|
| 50 |
|
| 51 |
def __init__(self, cfg):
|
| 52 |
super().__init__()
|
|
|
|
| 1 |
"""LFM2-VL and LFM2 on the hybrid stack.
|
| 2 |
|
| 3 |
HF's classes stay the frame (the SigLIP2 tower, the projector, the image
|
| 4 |
+
features, loading and saving); only the language model is ours, with HF's module
|
| 5 |
+
names, so checkpoints load and save unchanged, and a state
|
| 6 |
+
with any number of questions runs as one `hybrid.Tree`, which matches HF's own
|
| 7 |
implementation.
|
| 8 |
"""
|
| 9 |
|
|
|
|
| 46 |
|
| 47 |
|
| 48 |
class ShortConv(nn.Module):
|
| 49 |
+
"""`out(C * conv(B * x))`; HF's `nn.Conv1d` holds the taps, our causal conv runs them."""
|
| 50 |
|
| 51 |
def __init__(self, cfg):
|
| 52 |
super().__init__()
|
model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 6247065504
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:042be2b86f2140cf7e696d9d26c0dd8d2f0f8f943b783eb9ffa34903b41c98bb
|
| 3 |
size 6247065504
|
modeling_d1.py → modeling_lfm_jev.py
RENAMED
|
@@ -14,7 +14,7 @@ from .lfm2_vl import Lfm2VlForConditionalGeneration
|
|
| 14 |
from .runner import SystemOne
|
| 15 |
|
| 16 |
|
| 17 |
-
class
|
| 18 |
@cached_property
|
| 19 |
def engine(self) -> SystemOne:
|
| 20 |
from transformers import AutoTokenizer
|
|
|
|
| 14 |
from .runner import SystemOne
|
| 15 |
|
| 16 |
|
| 17 |
+
class LfmJevModel(Lfm2VlForConditionalGeneration):
|
| 18 |
@cached_property
|
| 19 |
def engine(self) -> SystemOne:
|
| 20 |
from transformers import AutoTokenizer
|
prompt.py
CHANGED
|
@@ -20,6 +20,7 @@ SYSTEMS: dict[str, str | None] = {
|
|
| 20 |
),
|
| 21 |
}
|
| 22 |
DEFAULT_SYSTEM = "none"
|
|
|
|
| 23 |
|
| 24 |
IM_START = "<|im_start|>"
|
| 25 |
IM_END = "<|im_end|>"
|
|
@@ -54,6 +55,20 @@ class Score:
|
|
| 54 |
Question = Choice | Noul | Score
|
| 55 |
|
| 56 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 57 |
# --------------------------------------------------------------------------- #
|
| 58 |
# verbalizer
|
| 59 |
# --------------------------------------------------------------------------- #
|
|
@@ -109,9 +124,12 @@ _TOKENIZERS: dict[int, object] = {}
|
|
| 109 |
def aliases(tokenizer, labels: Sequence[str]) -> list[tuple[str, int]]:
|
| 110 |
"""Assign every label a distinct single-token code: [(code, token_id)].
|
| 111 |
|
| 112 |
-
Memoised on the codes rather than on the labels
|
| 113 |
-
unless the labels are already letters, so
|
| 114 |
-
|
|
|
|
|
|
|
|
|
|
| 115 |
"""
|
| 116 |
_TOKENIZERS.setdefault(id(tokenizer), tokenizer)
|
| 117 |
codes = tuple(option_codes(labels))
|
|
@@ -191,16 +209,36 @@ def readout(tokenizer, q: Question, logz, calibration=None) -> list[float]:
|
|
| 191 |
return [e / sum(exps) for e in exps]
|
| 192 |
|
| 193 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 194 |
# --------------------------------------------------------------------------- #
|
| 195 |
# state and question rendering
|
| 196 |
# --------------------------------------------------------------------------- #
|
| 197 |
|
| 198 |
|
| 199 |
-
DEFAULT_MODEL = "LiquidAI/LFM2.5-VL-3B"
|
| 200 |
-
|
| 201 |
# How a state is rendered: `json_only`, the default, writes every state as the object it is; `json` keeps
|
| 202 |
# three shortcuts (`Message:`, `Passage:` / `Asked:`, a lone question's text); `sections` writes nested
|
| 203 |
# states as labelled blocks.
|
|
|
|
| 204 |
DEFAULT_STATE_STYLE = "json_only"
|
| 205 |
|
| 206 |
|
|
@@ -211,8 +249,9 @@ def _is_scalar(v: Any) -> bool:
|
|
| 211 |
def _sections(obj: Any, path: str, out: list[str]) -> None: # noqa: C901
|
| 212 |
"""Flatten a nested state into labelled blocks, keeping real newlines.
|
| 213 |
|
| 214 |
-
`json.dumps` escapes every newline inside a
|
| 215 |
-
|
|
|
|
| 216 |
"""
|
| 217 |
head = f"[{path}]\n" if path else ""
|
| 218 |
if isinstance(obj, dict):
|
|
@@ -319,8 +358,8 @@ def prefix_text(
|
|
| 319 |
return f"{bos}{turn}{IM_START}user\n{images}{body}"
|
| 320 |
|
| 321 |
|
| 322 |
-
# What sits between the assistant header and the answer slot
|
| 323 |
-
#
|
| 324 |
DEFAULT_LEAD = ""
|
| 325 |
LEADS: dict[str, str] = {}
|
| 326 |
|
|
|
|
| 20 |
),
|
| 21 |
}
|
| 22 |
DEFAULT_SYSTEM = "none"
|
| 23 |
+
SYSTEM = SYSTEMS[DEFAULT_SYSTEM]
|
| 24 |
|
| 25 |
IM_START = "<|im_start|>"
|
| 26 |
IM_END = "<|im_end|>"
|
|
|
|
| 55 |
Question = Choice | Noul | Score
|
| 56 |
|
| 57 |
|
| 58 |
+
def row_to_question(row: dict) -> Question:
|
| 59 |
+
"""A row's question (`kind`, `instructions`, `criteria`), as the model is asked it."""
|
| 60 |
+
kind = row["kind"]
|
| 61 |
+
if kind == "noul":
|
| 62 |
+
return Noul(instructions=row["instructions"], criteria=row.get("criteria"))
|
| 63 |
+
if kind == "score":
|
| 64 |
+
return Score(instructions=row["instructions"], criteria=list(row["criteria"]))
|
| 65 |
+
return Choice(instructions=row["instructions"], criteria=row["criteria"])
|
| 66 |
+
|
| 67 |
+
|
| 68 |
+
def cardinality(q: Question) -> int:
|
| 69 |
+
return 2 if isinstance(q, Noul) else len(q.criteria)
|
| 70 |
+
|
| 71 |
+
|
| 72 |
# --------------------------------------------------------------------------- #
|
| 73 |
# verbalizer
|
| 74 |
# --------------------------------------------------------------------------- #
|
|
|
|
| 124 |
def aliases(tokenizer, labels: Sequence[str]) -> list[tuple[str, int]]:
|
| 125 |
"""Assign every label a distinct single-token code: [(code, token_id)].
|
| 126 |
|
| 127 |
+
Memoised on the codes rather than on the labels. The codes are positional
|
| 128 |
+
unless the labels are already letters, so two orderings of one option list
|
| 129 |
+
share an entry, and a 151-option inventory question does not re-derive its
|
| 130 |
+
codes once per permuted row. Keying on the label tuple instead misses on
|
| 131 |
+
every such row. Measured against a contended node it was not the cost it
|
| 132 |
+
first looked like, but the entry is the codes and the key should say so.
|
| 133 |
"""
|
| 134 |
_TOKENIZERS.setdefault(id(tokenizer), tokenizer)
|
| 135 |
codes = tuple(option_codes(labels))
|
|
|
|
| 209 |
return [e / sum(exps) for e in exps]
|
| 210 |
|
| 211 |
|
| 212 |
+
def option_labels(q: Question) -> list[str]:
|
| 213 |
+
if isinstance(q, Noul):
|
| 214 |
+
return ["yes", "no"]
|
| 215 |
+
if isinstance(q, Score):
|
| 216 |
+
return [str(i) for i in range(len(q.criteria))]
|
| 217 |
+
return list(q.criteria.keys())
|
| 218 |
+
|
| 219 |
+
|
| 220 |
+
def gold_code(tokenizer, q: Question, gold: str | int) -> str:
|
| 221 |
+
"""The exact string the model must produce in the answer slot."""
|
| 222 |
+
if isinstance(q, Noul):
|
| 223 |
+
if isinstance(gold, str):
|
| 224 |
+
return "yes" if gold.strip().lower() in {"yes", "true", "1"} else "no"
|
| 225 |
+
return "yes" if int(gold) == 1 else "no"
|
| 226 |
+
if isinstance(q, Score):
|
| 227 |
+
return str(int(gold))
|
| 228 |
+
labels = list(q.criteria.keys())
|
| 229 |
+
idx = labels.index(gold) if gold in labels else int(gold)
|
| 230 |
+
return aliases(tokenizer, labels)[idx][0]
|
| 231 |
+
|
| 232 |
+
|
| 233 |
# --------------------------------------------------------------------------- #
|
| 234 |
# state and question rendering
|
| 235 |
# --------------------------------------------------------------------------- #
|
| 236 |
|
| 237 |
|
|
|
|
|
|
|
| 238 |
# How a state is rendered: `json_only`, the default, writes every state as the object it is; `json` keeps
|
| 239 |
# three shortcuts (`Message:`, `Passage:` / `Asked:`, a lone question's text); `sections` writes nested
|
| 240 |
# states as labelled blocks.
|
| 241 |
+
DEFAULT_MODEL = "LiquidAI/LFM2.5-VL-3B"
|
| 242 |
DEFAULT_STATE_STYLE = "json_only"
|
| 243 |
|
| 244 |
|
|
|
|
| 249 |
def _sections(obj: Any, path: str, out: list[str]) -> None: # noqa: C901
|
| 250 |
"""Flatten a nested state into labelled blocks, keeping real newlines.
|
| 251 |
|
| 252 |
+
A caller's state is JSON, but `json.dumps` escapes every newline inside a
|
| 253 |
+
log line or a record, so a 40-line CrowdStrike host record arrives as one
|
| 254 |
+
unreadable string of `\n`. The model reads a case file, not a payload.
|
| 255 |
"""
|
| 256 |
head = f"[{path}]\n" if path else ""
|
| 257 |
if isinstance(obj, dict):
|
|
|
|
| 358 |
return f"{bos}{turn}{IM_START}user\n{images}{body}"
|
| 359 |
|
| 360 |
|
| 361 |
+
# What sits between the assistant header and the answer slot: nothing on LFM2-VL, whose template opens
|
| 362 |
+
# no reasoning block, so the slot is already clean. A family module can add its own (`LEADS`).
|
| 363 |
DEFAULT_LEAD = ""
|
| 364 |
LEADS: dict[str, str] = {}
|
| 365 |
|
runner.py
CHANGED
|
@@ -39,7 +39,7 @@ MODELS = {
|
|
| 39 |
|
| 40 |
|
| 41 |
def load_backbone(model_id: str = DEFAULT_MODEL, dtype=torch.bfloat16):
|
| 42 |
-
"""The checkpoint and its tokenizer,
|
| 43 |
from transformers import AutoModelForImageTextToText, AutoTokenizer, PretrainedConfig
|
| 44 |
|
| 45 |
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
|
|
@@ -82,8 +82,8 @@ class SystemOne(SystemOneApi):
|
|
| 82 |
model=None,
|
| 83 |
tokenizer=None,
|
| 84 |
):
|
| 85 |
-
"""`model` and `tokenizer`, when given, are a backbone already loaded (
|
| 86 |
-
it stays on its device unless `device` says otherwise."""
|
| 87 |
if model is None:
|
| 88 |
model, tokenizer = load_backbone(model_id)
|
| 89 |
else:
|
|
@@ -103,7 +103,8 @@ class SystemOne(SystemOneApi):
|
|
| 103 |
self.option_style = option_style
|
| 104 |
self.token_budget = token_budget
|
| 105 |
self.processor = None
|
| 106 |
-
# CUDA graphs for single questions on NVIDIA
|
|
|
|
| 107 |
if compile and torch.version.hip:
|
| 108 |
raise ValueError("compile=True needs CUDA: on ROCm the CUDA graphs fault after a few dozen calls")
|
| 109 |
self._one_pass = (torch.compile(self.model.forward, mode="reduce-overhead")
|
|
@@ -119,14 +120,19 @@ class SystemOne(SystemOneApi):
|
|
| 119 |
|
| 120 |
# --------------------------------------------------------------- forward
|
| 121 |
|
| 122 |
-
|
| 123 |
-
|
|
|
|
|
|
|
| 124 |
|
| 125 |
The rows are one tree (`hybrid.py`): their common start is its trunk and
|
| 126 |
is read once; the rest of each row is a branch, packed with no padding.
|
| 127 |
Mathematically each row alone; in bf16 the kernels differ by batch shape.
|
| 128 |
"""
|
| 129 |
-
|
|
|
|
|
|
|
|
|
|
| 130 |
row = self._one_pass(input_ids=torch.tensor(rows, device=self.device), logits_to_keep=1).logits[0, -1]
|
| 131 |
return [row.float() - torch.logsumexp(row.float(), dim=-1)]
|
| 132 |
shared = 0 # every row keeps at least its last token
|
|
|
|
| 39 |
|
| 40 |
|
| 41 |
def load_backbone(model_id: str = DEFAULT_MODEL, dtype=torch.bfloat16):
|
| 42 |
+
"""The checkpoint and its tokenizer. SigLIP2 has no FA2 kernel here, so SDPA."""
|
| 43 |
from transformers import AutoModelForImageTextToText, AutoTokenizer, PretrainedConfig
|
| 44 |
|
| 45 |
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
|
|
|
|
| 82 |
model=None,
|
| 83 |
tokenizer=None,
|
| 84 |
):
|
| 85 |
+
"""`model` and `tokenizer`, when given, are a backbone already loaded (the Hub's remote-code class
|
| 86 |
+
passes itself); it stays on its device unless `device` says otherwise."""
|
| 87 |
if model is None:
|
| 88 |
model, tokenizer = load_backbone(model_id)
|
| 89 |
else:
|
|
|
|
| 103 |
self.option_style = option_style
|
| 104 |
self.token_budget = token_budget
|
| 105 |
self.processor = None
|
| 106 |
+
# CUDA graphs for single questions, on NVIDIA: where eager time is kernel launches, 16 to 8 ms a
|
| 107 |
+
# decision on an RTX 4090. The first two prompt lengths compile; later ones reuse the dynamic graph.
|
| 108 |
if compile and torch.version.hip:
|
| 109 |
raise ValueError("compile=True needs CUDA: on ROCm the CUDA graphs fault after a few dozen calls")
|
| 110 |
self._one_pass = (torch.compile(self.model.forward, mode="reduce-overhead")
|
|
|
|
| 120 |
|
| 121 |
# --------------------------------------------------------------- forward
|
| 122 |
|
| 123 |
+
@torch.inference_mode()
|
| 124 |
+
def _logz(self, texts: Sequence[str], questions: Sequence[Question] | None = None) -> list[torch.Tensor]:
|
| 125 |
+
"""Log-softmax at the answer slot, one row per text, in one pass; the slot needs only the texts, and
|
| 126 |
+
`questions` is accepted for the scoring interface.
|
| 127 |
|
| 128 |
The rows are one tree (`hybrid.py`): their common start is its trunk and
|
| 129 |
is read once; the rest of each row is a branch, packed with no padding.
|
| 130 |
Mathematically each row alone; in bf16 the kernels differ by batch shape.
|
| 131 |
"""
|
| 132 |
+
return self._logz_ids([self.tokenizer.encode(t, add_special_tokens=False) for t in texts])
|
| 133 |
+
|
| 134 |
+
def _logz_ids(self, rows: list[list[int]]) -> list[torch.Tensor]:
|
| 135 |
+
if len(rows) == 1: # nothing to share: a plain chain, 6-12% faster than a tree of one
|
| 136 |
row = self._one_pass(input_ids=torch.tensor(rows, device=self.device), logits_to_keep=1).logits[0, -1]
|
| 137 |
return [row.float() - torch.logsumexp(row.float(), dim=-1)]
|
| 138 |
shared = 0 # every row keeps at least its last token
|