README.md CHANGED
@@ -31,7 +31,6 @@ tags:
31
  - calibration
32
  - system-one
33
  - multimodal
34
- - decision-model
35
  ---
36
 
37
  <div align="center">
@@ -50,22 +49,21 @@ tags:
50
  # d1-3B
51
 
52
  d1-3B is a 3B parameter **decision model** built on [LFM2.5-VL-3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B).
53
- You give it a state (text, JSON, images, or a mix) and a set of questions. It returns calibrated,
54
- typed answers in **one forward pass with zero output tokens**.
 
55
 
56
- - **Best decision model under 10B on the Decision Index 0.2.1**: 48.57, ahead of every 4B and 9B model
57
- and of Decider 35B-A3B (47.11).
58
- - **Multimodal**: images and text in the same state. It scores 74.1 on 11 public image benchmarks
59
  (LFM2.5-VL-3B: 73.9).
60
  - **Fast**: 8 ms a decision on an NVIDIA RTX 4090, 9 ms on an AMD MI325X, 30 ms on an Apple M5 Pro.
61
 
62
- Find more information about open d1 in our [blog post](https://www.liquid.ai/blog/open-d1).
63
-
64
- ![image](https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/r2H7UlZ_m48SYAhZWIHiu.png)
65
 
66
  > [!NOTE]
67
  > 💻 **Demos**: Try d1-3B in a Hugging Face space without any setup:
68
- > [**Open d1 Arcade**](https://huggingface.co/spaces/LiquidAI/system-one-arcade): Collection of 10 demos using d1-3B
69
 
70
 
71
  ## 🗒️ Model Details
@@ -73,14 +71,16 @@ Find more information about open d1 in our [blog post](https://www.liquid.ai/blo
73
  | Model | Parameters | Description |
74
  |---|---|---|
75
  | [LFM2.5-VL-3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B) | 3.1B | General-purpose vision-language model (base) |
76
- | **[d1-3B](https://huggingface.co/LiquidAI/d1-3B)** | 3.1B | Post-trained for single-pass, calibrated decisions |
77
 
78
  d1-3B is a multimodal decision model with the following features:
79
 
80
  - **Total parameters**: 3.12B
 
81
  - **Vision encoder**: SigLIP2 NaFlex shape-optimized 400M
82
  - **Context length**: 32,768 tokens
83
  - **Vocabulary size**: 128,000
 
84
 
85
  We recommend d1-3B wherever a pipeline needs a yes/no, a pick from named options, or a rating:
86
  routing and triage, moderation, intent and topic classification, extraction checks, reranking, LLM-as-a-judge
@@ -97,47 +97,36 @@ pip install "transformers>=5.14" torch torchvision pillow
97
  The model ships its own code, so load it with `trust_remote_code=True`:
98
 
99
  ```python
 
 
 
100
  import torch
 
101
  from transformers import AutoModel
102
- from transformers.image_utils import load_image
103
 
104
  device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
105
- dtype = torch.float32 if device == "cpu" else torch.bfloat16
106
- model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True, dtype=dtype).to(device)
107
 
108
- # Text: several named questions over one state, answered in one pass
109
  questions = {
110
- "refund": {
111
- "type": "noul",
112
- "instructions": "Is the customer asking for a refund?",
113
- },
114
- "team": {
115
- "type": "choice",
116
- "instructions": "Which team should handle this?",
117
- "criteria": {
118
- "billing": "Charges, refunds, invoices",
119
- "technical": "App or site faults",
120
- "fraud": "Suspected unauthorised use",
121
- },
122
- },
123
- "urgency": {
124
- "type": "score",
125
- "instructions": "How urgent is this?",
126
- "criteria": ["Can wait", "Today", "Blocking the customer now"],
127
- },
128
  }
129
  print(model.system_one("I was charged twice this month, please refund one of them.", questions))
130
 
131
- # Image: the photo is the whole state
132
- image = load_image("http://images.cocodataset.org/val2017/000000039769.jpg") # two cats on a sofa
133
- cats = {
134
- "type": "choice",
135
- "instructions": "How many cats are there?",
136
- "criteria": {"one": "One", "two": "Two", "more": "Three or more"},
137
- }
138
- print(model.system_one(None, {"cats": cats}, images=[image]))
139
 
140
- # Batch: many requests, packed together with no padding
141
  tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
142
  print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))
143
  ```
@@ -163,32 +152,16 @@ Each call returns `{"answers": {name: answer}, "usage": {"input_tokens": n, "out
163
 
164
  ## ⚡ Speed
165
 
166
- Warm calls, one request at a time: a single question, three questions over one state, a 3.4k-token
167
- state and a 384 px image. The last column is throughput with 64 states packed into one pass.
168
-
169
- ### Edge Inference
170
-
171
- We measure latency on an Apple M5 Pro and, in collaboration with NVIDIA, on an NVIDIA Jetson AGX Thor,
172
- a Jetson AGX Orin 64 GB and a Jetson Orin Nano.
173
 
174
  | | one question | 3 questions, one pass | 3.4k-token state | 384 px image | 64 states, packed |
175
  |---|---:|---:|---:|---:|---:|
176
- | Apple M5 Pro (`mps`) | 30 ms | 41 ms | 640 ms | 62 ms | 78 / s |
177
- | NVIDIA Jetson AGX Thor | 16 ms | 20 ms | 220 ms | 35 ms | 262 / s |
178
- | NVIDIA Jetson AGX Orin 64 GB | 26 ms | 35 ms | 560 ms | 83 ms | 110 / s |
179
- | NVIDIA Jetson Orin Nano | 50 ms | 73 ms | 1640 ms | 202 ms | 38 / s |
180
 
181
- ### GPU Inference
182
-
183
- We measure latency on an NVIDIA RTX 4090 and an AMD MI325X, in bf16, median of 20 runs.
184
-
185
- | | one question | 3 questions, one pass | 3.4k-token state | 384 px image | 64 states, packed |
186
- |---|---:|---:|---:|---:|---:|
187
- | NVIDIA RTX 4090 | 8 ms | 21 ms | 102 ms | 17 ms | 475 / s |
188
- | AMD MI325X | 9 ms | 14 ms | 44 ms | 18 ms | 1,106 / s |
189
-
190
- On NVIDIA GPUs, `model.compile(mode="reduce-overhead")` runs single questions as CUDA graphs (the RTX 4090
191
- row uses it). Without it, a single question takes 16 ms. The first call with a new shape pays for kernel
192
  selection or compilation, so warm up the shapes you serve.
193
 
194
  ## 📊 Performance
@@ -197,12 +170,15 @@ All results are on public benchmarks.
197
 
198
  ### Decision Index 0.2.1
199
 
200
- We scored d1-3B with the official scorer (not a leaderboard submission). All other rows come from the public leaderboard v0.2.1.
 
 
 
201
 
202
  | Model | Size | Decision Index | Knowledge | Language | Retrieval | Tools | Arts |
203
  |---|---:|---:|---:|---:|---:|---:|---:|
204
  | Winnow-12B | 12B | 50.02 | 33.8 | 56.0 | 54.0 | 71.0 | 30.0 |
205
- | **d1-3B** | **3B** | **48.57** | 23.8 | 56.4 | 52.8 | **74.5** | **36.3** |
206
  | Decider 35B-A3B | 36B | 47.11 | 31.8 | 55.5 | 54.7 | 56.5 | 32.6 |
207
  | JPT-9B | 9.7B | 46.89 | 31.7 | 56.7 | 44.6 | 67.0 | 28.6 |
208
  | Decision 1.0 Lux | 9.7B | 43.49 | 30.9 | 48.0 | 50.0 | 57.2 | 26.4 |
@@ -214,22 +190,23 @@ We scored d1-3B with the official scorer (not a leaderboard submission). All oth
214
 
215
  ### Benchmarks as decisions
216
 
217
- Besides the Decision Index, we added a few other internal evaluations based on public benchmarks.
 
218
 
219
  | Benchmark | d1-3B | Decider 4B | Decider 2B |
220
  |---|---:|---:|---:|
221
- | SQuAD 2.0 | **85.3** | 76.0 | 67.7 |
222
- | Civil Comments | 93.0 | 92.8 | **93.6** |
223
- | MASSIVE intent | 87.3 | **88.3** | 81.1 |
224
- | HelpSteer2 | 36.7 | **42.0** | 32.0 |
225
- | PubMedQA | **66.0** | 63.3 | 65.7 |
226
- | BoolQ | 86.7 | **89.0** | 87.3 |
227
- | XNLI | 85.0 | **88.6** | 85.0 |
228
- | PAWS-X | **76.9** | 69.8 | 59.5 |
229
- | **Mean** | **77.1** | 76.2 | 71.5 |
230
-
231
- d1-3B also scores 71.8 on [DecisionBench](https://huggingface.co/datasets/Hanno-Labs/decision-bench) (eng v1,
232
- all 23,900 rows) and 69.3 on [Fast Decisions](https://huggingface.co/datasets/fastino/fast-decisions)
233
  (dev split).
234
 
235
  ### Vision
@@ -239,22 +216,29 @@ compared with the base model:
239
 
240
  | Benchmark | d1-3B | LFM2.5-VL-3B |
241
  |---|---:|---:|
242
- | AI2D | 79.9 | 80.9 |
243
- | BLINK | 59.2 | 58.7 |
244
- | CV-Bench | 82.1 | 87.6 |
245
- | HallusionBench | 65.3 | 65.0 |
246
- | MMBench | 84.9 | 84.3 |
247
- | MME | 82.1 | 82.4 |
248
- | MMStar | 59.9 | 61.2 |
249
- | MMVP | 77.0 | 73.7 |
250
  | POPE | 88.5 | 90.1 |
251
- | VisualWebBench | 71.4 | 78.3 |
252
- | VL-RewardBench | 65.0 | 50.9 |
253
- | **Mean** | **74.1** | 73.9 |
254
- | [ImajevBench](https://huggingface.co/datasets/mohit67890/imajev-bench) (dev and calibration, 253 rows) | 64.0 | 66.8 |
255
 
256
  With the images removed, the same questions score 45.1, so the answers come from the images.
257
 
 
 
 
 
 
 
 
258
  ## 📬 Contact
259
 
260
  - Got questions or want to connect? [Join our Discord community](https://discord.com/invite/liquid-ai)
@@ -262,16 +246,6 @@ With the images removed, the same questions score 45.1, so the answers come from
262
 
263
  ## Citation
264
 
265
- ```bibtex
266
- @article{liquidAI2026opend1,
267
- author = {Liquid AI},
268
- title = {Open d1: Edge decision models for text, vision, and audio},
269
- journal = {Liquid AI Blog},
270
- year = {2026},
271
- note = {https://www.liquid.ai/blog/d1-open},
272
- }
273
- ```
274
-
275
  ```bibtex
276
  @article{liquidai2025lfm2,
277
  title = {LFM2 Technical Report},
 
31
  - calibration
32
  - system-one
33
  - multimodal
 
34
  ---
35
 
36
  <div align="center">
 
49
  # d1-3B
50
 
51
  d1-3B is a 3B parameter **decision model** built on [LFM2.5-VL-3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B).
52
+ You give it a state (text, JSON, images, or a mix) and a set of named questions. It returns calibrated,
53
+ typed answers in **one forward pass with zero output tokens**: every answer is read directly from the
54
+ model's distribution over the options, with no generation and no parsing.
55
 
56
+ - **Best decision model under 10B**: 47.12 on the Decision Index 0.2.1. That is ahead of every 4B and 9B
57
+ model on the public leaderboard and level with Decider 35B-A3B (47.11).
58
+ - **Multimodal**: images and text in the same state. It scores 73.7 on 11 public image benchmarks
59
  (LFM2.5-VL-3B: 73.9).
60
  - **Fast**: 8 ms a decision on an NVIDIA RTX 4090, 9 ms on an AMD MI325X, 30 ms on an Apple M5 Pro.
61
 
62
+ ![Decision Index against model size](assets/di_pareto.png)
 
 
63
 
64
  > [!NOTE]
65
  > 💻 **Demos**: Try d1-3B in a Hugging Face space without any setup:
66
+ > **TODO: add demo**
67
 
68
 
69
  ## 🗒️ Model Details
 
71
  | Model | Parameters | Description |
72
  |---|---|---|
73
  | [LFM2.5-VL-3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B) | 3.1B | General-purpose vision-language model (base) |
74
+ | **[d1-3B](https://huggingface.co/LiquidAI/d1-3b-RC)** | 3.1B | Post-trained for single-pass, calibrated decisions |
75
 
76
  d1-3B is a multimodal decision model with the following features:
77
 
78
  - **Total parameters**: 3.12B
79
+ - **LM backbone**: LFM2.5-2.6B, 30 layers (22 double-gated short convolution blocks + 8 GQA)
80
  - **Vision encoder**: SigLIP2 NaFlex shape-optimized 400M
81
  - **Context length**: 32,768 tokens
82
  - **Vocabulary size**: 128,000
83
+ - **Output**: no generated tokens. Each answer is a probability distribution over the question's options.
84
 
85
  We recommend d1-3B wherever a pipeline needs a yes/no, a pick from named options, or a rating:
86
  routing and triage, moderation, intent and topic classification, extraction checks, reranking, LLM-as-a-judge
 
97
  The model ships its own code, so load it with `trust_remote_code=True`:
98
 
99
  ```python
100
+ import io
101
+ import urllib.request
102
+
103
  import torch
104
+ from PIL import Image
105
  from transformers import AutoModel
 
106
 
107
  device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
108
+ model = AutoModel.from_pretrained("LiquidAI/d1-3b-RC", trust_remote_code=True,
109
+ dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)
110
 
111
+ # Several named questions over one text state, answered in one pass
112
  questions = {
113
+ "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
114
+ "team": {"type": "choice", "instructions": "Which team should handle this?",
115
+ "criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
116
+ "fraud": "Suspected unauthorised use"}},
117
+ "urgency": {"type": "score", "instructions": "How urgent is this?",
118
+ "criteria": ["Can wait", "Today", "Blocking the customer now"]},
 
 
 
 
 
 
 
 
 
 
 
 
119
  }
120
  print(model.system_one("I was charged twice this month, please refund one of them.", questions))
121
 
122
+ # An image as the whole state
123
+ url = "http://images.cocodataset.org/val2017/000000039769.jpg" # two cats on a sofa
124
+ photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))
125
+ print(model.system_one(None, {"cats": {"type": "choice", "instructions": "How many cats are there?",
126
+ "criteria": {"one": "One", "two": "Two", "more": "Three or more"}}},
127
+ images=[photo]))
 
 
128
 
129
+ # Many requests, packed together with no padding
130
  tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
131
  print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))
132
  ```
 
152
 
153
  ## ⚡ Speed
154
 
155
+ Warm calls in bf16, median of 20 runs:
 
 
 
 
 
 
156
 
157
  | | one question | 3 questions, one pass | 3.4k-token state | 384 px image | 64 states, packed |
158
  |---|---:|---:|---:|---:|---:|
159
+ | NVIDIA RTX 4090 | **8.0 ms** | 21 ms | 102 ms | 17 ms | 475 / s |
160
+ | AMD MI325X | 9.2 ms | 14 ms | 44 ms | 18 ms | 1,106 / s |
161
+ | Apple M5 Pro (`mps`) | 30 ms | 41 ms | 0.64 s | 62 ms | 78 / s |
 
162
 
163
+ On NVIDIA GPUs, `model.compile(mode="reduce-overhead")` runs single questions as CUDA graphs. The RTX 4090
164
+ row uses it; without it a single question takes 16 ms. The first call with a new shape pays for kernel
 
 
 
 
 
 
 
 
 
165
  selection or compilation, so warm up the shapes you serve.
166
 
167
  ## 📊 Performance
 
170
 
171
  ### Decision Index 0.2.1
172
 
173
+ The [Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index) casts 38 public
174
+ benchmarks as 120,226 typed decisions, scored as chance-corrected skill. We scored d1-3B with the official
175
+ scorer; it is not a leaderboard submission. All other rows come from the public leaderboard
176
+ (snapshot 2026-09-28). Sizes are served parameters.
177
 
178
  | Model | Size | Decision Index | Knowledge | Language | Retrieval | Tools | Arts |
179
  |---|---:|---:|---:|---:|---:|---:|---:|
180
  | Winnow-12B | 12B | 50.02 | 33.8 | 56.0 | 54.0 | 71.0 | 30.0 |
181
+ | **d1-3B** | **3B** | **47.12** | 27.8 | 53.5 | 52.2 | 66.7 | **34.6** |
182
  | Decider 35B-A3B | 36B | 47.11 | 31.8 | 55.5 | 54.7 | 56.5 | 32.6 |
183
  | JPT-9B | 9.7B | 46.89 | 31.7 | 56.7 | 44.6 | 67.0 | 28.6 |
184
  | Decision 1.0 Lux | 9.7B | 43.49 | 30.9 | 48.0 | 50.0 | 57.2 | 26.4 |
 
190
 
191
  ### Benchmarks as decisions
192
 
193
+ Evaluation splits are cast as typed decisions, with the same rows for every model. These scores follow the
194
+ decision protocol, not each benchmark's own.
195
 
196
  | Benchmark | d1-3B | Decider 4B | Decider 2B |
197
  |---|---:|---:|---:|
198
+ | SQuAD 2.0 | **83.3** | 76.0 | 67.7 |
199
+ | Civil Comments | 93.3 | 92.8 | **93.6** |
200
+ | MASSIVE intent | 86.9 | **88.3** | 81.1 |
201
+ | HelpSteer2 | 38.0 | **42.0** | 32.0 |
202
+ | PubMedQA | **68.3** | 63.3 | 65.7 |
203
+ | BoolQ | 86.3 | **89.0** | 87.3 |
204
+ | XNLI | 85.6 | **88.6** | 85.0 |
205
+ | PAWS-X | **76.4** | 69.8 | 59.5 |
206
+ | **Mean** | **77.3** | 76.2 | 71.5 |
207
+
208
+ d1-3B also scores 69.5 on [DecisionBench](https://huggingface.co/datasets/Hanno-Labs/decision-bench) (eng v1,
209
+ all 23,900 rows) and 68.8 on [Fast Decisions](https://huggingface.co/datasets/fastino/fast-decisions)
210
  (dev split).
211
 
212
  ### Vision
 
216
 
217
  | Benchmark | d1-3B | LFM2.5-VL-3B |
218
  |---|---:|---:|
219
+ | AI2D | 80.1 | 80.9 |
220
+ | BLINK | 58.2 | 58.7 |
221
+ | CV-Bench | 82.7 | 87.6 |
222
+ | HallusionBench | 63.8 | 65.0 |
223
+ | MMBench | 84.3 | 84.3 |
224
+ | MME | 81.7 | 82.4 |
225
+ | MMStar | 59.4 | 61.2 |
226
+ | MMVP | 76.3 | 73.7 |
227
  | POPE | 88.5 | 90.1 |
228
+ | VisualWebBench | 71.6 | 78.3 |
229
+ | VL-RewardBench | 63.6 | 50.9 |
230
+ | **Mean** | **73.7** | 73.9 |
231
+ | [ImajevBench](https://huggingface.co/datasets/mohit67890/imajev-bench) (dev and calibration, 253 rows) | 66.8 | 66.8 |
232
 
233
  With the images removed, the same questions score 45.1, so the answers come from the images.
234
 
235
+ ## ⚠️ Limitations
236
+
237
+ - **Knowledge-heavy questions**: larger models stay ahead on hard knowledge (for example MMLU-Pro and GPQA).
238
+ - **Thresholds**: calibrate any threshold on data from your own deployment before gating on it.
239
+ - **No text output**: d1-3B answers questions over fixed options. Use [LFM2.5-VL-3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B)
240
+ for open-ended generation.
241
+
242
  ## 📬 Contact
243
 
244
  - Got questions or want to connect? [Join our Discord community](https://discord.com/invite/liquid-ai)
 
246
 
247
  ## Citation
248
 
 
 
 
 
 
 
 
 
 
 
249
  ```bibtex
250
  @article{liquidai2025lfm2,
251
  title = {LFM2 Technical Report},
assets/d1-3b-smash.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:08f69b1a67bf32cd18fad32b5872519b580be9e94bafbe26a13d5cf0aaa49c1f
3
+ size 53062187
assets/di_pareto.png ADDED

Git LFS Details

  • SHA256: 62e7f5907b765063a710291d273f33d622d9bcff3efc74f3343d4dbdb469d73f
  • Pointer size: 131 Bytes
  • Size of remote file: 146 kB
assets/ood_tasks.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b491dfa851157e12ba2bee33a394c9565408247cf0004e493730622c62878e42
3
+ size 8931625
config.json CHANGED
@@ -3,7 +3,7 @@
3
  "Lfm2VlForConditionalGeneration"
4
  ],
5
  "auto_map": {
6
- "AutoModel": "modeling_d1.D1Model"
7
  },
8
  "bos_token_id": 124894,
9
  "do_image_splitting": true,
 
3
  "Lfm2VlForConditionalGeneration"
4
  ],
5
  "auto_map": {
6
+ "AutoModel": "modeling_lfm_jev.LfmJevModel"
7
  },
8
  "bos_token_id": 124894,
9
  "do_image_splitting": true,
lfm2_vl.py CHANGED
@@ -1,9 +1,9 @@
1
  """LFM2-VL and LFM2 on the hybrid stack.
2
 
3
  HF's classes stay the frame (the SigLIP2 tower, the projector, the image
4
- features, loading and saving); only the language model is replaced, with HF's
5
- module names, so checkpoints load and save unchanged. A state with any number
6
- of questions runs as one `hybrid.Tree`, with the same outputs as HF's own
7
  implementation.
8
  """
9
 
@@ -46,7 +46,7 @@ class Attention(nn.Module):
46
 
47
 
48
  class ShortConv(nn.Module):
49
- """`out(C * conv(B * x))`; HF's `nn.Conv1d` holds the taps, `causal_conv` runs them."""
50
 
51
  def __init__(self, cfg):
52
  super().__init__()
 
1
  """LFM2-VL and LFM2 on the hybrid stack.
2
 
3
  HF's classes stay the frame (the SigLIP2 tower, the projector, the image
4
+ features, loading and saving); only the language model is ours, with HF's module
5
+ names, so checkpoints load and save unchanged, and a state
6
+ with any number of questions runs as one `hybrid.Tree`, which matches HF's own
7
  implementation.
8
  """
9
 
 
46
 
47
 
48
  class ShortConv(nn.Module):
49
+ """`out(C * conv(B * x))`; HF's `nn.Conv1d` holds the taps, our causal conv runs them."""
50
 
51
  def __init__(self, cfg):
52
  super().__init__()
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:50e03317847caf6df9a9aee27ed40f20554a86a21e60d1d47ba41a422b546c0c
3
  size 6247065504
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:042be2b86f2140cf7e696d9d26c0dd8d2f0f8f943b783eb9ffa34903b41c98bb
3
  size 6247065504
modeling_d1.py → modeling_lfm_jev.py RENAMED
@@ -14,7 +14,7 @@ from .lfm2_vl import Lfm2VlForConditionalGeneration
14
  from .runner import SystemOne
15
 
16
 
17
- class D1Model(Lfm2VlForConditionalGeneration):
18
  @cached_property
19
  def engine(self) -> SystemOne:
20
  from transformers import AutoTokenizer
 
14
  from .runner import SystemOne
15
 
16
 
17
+ class LfmJevModel(Lfm2VlForConditionalGeneration):
18
  @cached_property
19
  def engine(self) -> SystemOne:
20
  from transformers import AutoTokenizer
prompt.py CHANGED
@@ -20,6 +20,7 @@ SYSTEMS: dict[str, str | None] = {
20
  ),
21
  }
22
  DEFAULT_SYSTEM = "none"
 
23
 
24
  IM_START = "<|im_start|>"
25
  IM_END = "<|im_end|>"
@@ -54,6 +55,20 @@ class Score:
54
  Question = Choice | Noul | Score
55
 
56
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
57
  # --------------------------------------------------------------------------- #
58
  # verbalizer
59
  # --------------------------------------------------------------------------- #
@@ -109,9 +124,12 @@ _TOKENIZERS: dict[int, object] = {}
109
  def aliases(tokenizer, labels: Sequence[str]) -> list[tuple[str, int]]:
110
  """Assign every label a distinct single-token code: [(code, token_id)].
111
 
112
- Memoised on the codes rather than on the labels: the codes are positional
113
- unless the labels are already letters, so every option list of one length
114
- shares an entry.
 
 
 
115
  """
116
  _TOKENIZERS.setdefault(id(tokenizer), tokenizer)
117
  codes = tuple(option_codes(labels))
@@ -191,16 +209,36 @@ def readout(tokenizer, q: Question, logz, calibration=None) -> list[float]:
191
  return [e / sum(exps) for e in exps]
192
 
193
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
194
  # --------------------------------------------------------------------------- #
195
  # state and question rendering
196
  # --------------------------------------------------------------------------- #
197
 
198
 
199
- DEFAULT_MODEL = "LiquidAI/LFM2.5-VL-3B"
200
-
201
  # How a state is rendered: `json_only`, the default, writes every state as the object it is; `json` keeps
202
  # three shortcuts (`Message:`, `Passage:` / `Asked:`, a lone question's text); `sections` writes nested
203
  # states as labelled blocks.
 
204
  DEFAULT_STATE_STYLE = "json_only"
205
 
206
 
@@ -211,8 +249,9 @@ def _is_scalar(v: Any) -> bool:
211
  def _sections(obj: Any, path: str, out: list[str]) -> None: # noqa: C901
212
  """Flatten a nested state into labelled blocks, keeping real newlines.
213
 
214
- `json.dumps` escapes every newline inside a log line or a record, so a
215
- multi-line record would arrive as one string of `\n`; this keeps it readable.
 
216
  """
217
  head = f"[{path}]\n" if path else ""
218
  if isinstance(obj, dict):
@@ -319,8 +358,8 @@ def prefix_text(
319
  return f"{bos}{turn}{IM_START}user\n{images}{body}"
320
 
321
 
322
- # What sits between the assistant header and the answer slot, per model type: nothing on LFM2-VL, whose
323
- # template opens no reasoning block.
324
  DEFAULT_LEAD = ""
325
  LEADS: dict[str, str] = {}
326
 
 
20
  ),
21
  }
22
  DEFAULT_SYSTEM = "none"
23
+ SYSTEM = SYSTEMS[DEFAULT_SYSTEM]
24
 
25
  IM_START = "<|im_start|>"
26
  IM_END = "<|im_end|>"
 
55
  Question = Choice | Noul | Score
56
 
57
 
58
+ def row_to_question(row: dict) -> Question:
59
+ """A row's question (`kind`, `instructions`, `criteria`), as the model is asked it."""
60
+ kind = row["kind"]
61
+ if kind == "noul":
62
+ return Noul(instructions=row["instructions"], criteria=row.get("criteria"))
63
+ if kind == "score":
64
+ return Score(instructions=row["instructions"], criteria=list(row["criteria"]))
65
+ return Choice(instructions=row["instructions"], criteria=row["criteria"])
66
+
67
+
68
+ def cardinality(q: Question) -> int:
69
+ return 2 if isinstance(q, Noul) else len(q.criteria)
70
+
71
+
72
  # --------------------------------------------------------------------------- #
73
  # verbalizer
74
  # --------------------------------------------------------------------------- #
 
124
  def aliases(tokenizer, labels: Sequence[str]) -> list[tuple[str, int]]:
125
  """Assign every label a distinct single-token code: [(code, token_id)].
126
 
127
+ Memoised on the codes rather than on the labels. The codes are positional
128
+ unless the labels are already letters, so two orderings of one option list
129
+ share an entry, and a 151-option inventory question does not re-derive its
130
+ codes once per permuted row. Keying on the label tuple instead misses on
131
+ every such row. Measured against a contended node it was not the cost it
132
+ first looked like, but the entry is the codes and the key should say so.
133
  """
134
  _TOKENIZERS.setdefault(id(tokenizer), tokenizer)
135
  codes = tuple(option_codes(labels))
 
209
  return [e / sum(exps) for e in exps]
210
 
211
 
212
+ def option_labels(q: Question) -> list[str]:
213
+ if isinstance(q, Noul):
214
+ return ["yes", "no"]
215
+ if isinstance(q, Score):
216
+ return [str(i) for i in range(len(q.criteria))]
217
+ return list(q.criteria.keys())
218
+
219
+
220
+ def gold_code(tokenizer, q: Question, gold: str | int) -> str:
221
+ """The exact string the model must produce in the answer slot."""
222
+ if isinstance(q, Noul):
223
+ if isinstance(gold, str):
224
+ return "yes" if gold.strip().lower() in {"yes", "true", "1"} else "no"
225
+ return "yes" if int(gold) == 1 else "no"
226
+ if isinstance(q, Score):
227
+ return str(int(gold))
228
+ labels = list(q.criteria.keys())
229
+ idx = labels.index(gold) if gold in labels else int(gold)
230
+ return aliases(tokenizer, labels)[idx][0]
231
+
232
+
233
  # --------------------------------------------------------------------------- #
234
  # state and question rendering
235
  # --------------------------------------------------------------------------- #
236
 
237
 
 
 
238
  # How a state is rendered: `json_only`, the default, writes every state as the object it is; `json` keeps
239
  # three shortcuts (`Message:`, `Passage:` / `Asked:`, a lone question's text); `sections` writes nested
240
  # states as labelled blocks.
241
+ DEFAULT_MODEL = "LiquidAI/LFM2.5-VL-3B"
242
  DEFAULT_STATE_STYLE = "json_only"
243
 
244
 
 
249
  def _sections(obj: Any, path: str, out: list[str]) -> None: # noqa: C901
250
  """Flatten a nested state into labelled blocks, keeping real newlines.
251
 
252
+ A caller's state is JSON, but `json.dumps` escapes every newline inside a
253
+ log line or a record, so a 40-line CrowdStrike host record arrives as one
254
+ unreadable string of `\n`. The model reads a case file, not a payload.
255
  """
256
  head = f"[{path}]\n" if path else ""
257
  if isinstance(obj, dict):
 
358
  return f"{bos}{turn}{IM_START}user\n{images}{body}"
359
 
360
 
361
+ # What sits between the assistant header and the answer slot: nothing on LFM2-VL, whose template opens
362
+ # no reasoning block, so the slot is already clean. A family module can add its own (`LEADS`).
363
  DEFAULT_LEAD = ""
364
  LEADS: dict[str, str] = {}
365
 
runner.py CHANGED
@@ -39,7 +39,7 @@ MODELS = {
39
 
40
 
41
  def load_backbone(model_id: str = DEFAULT_MODEL, dtype=torch.bfloat16):
42
- """The checkpoint and its tokenizer, with SDPA attention."""
43
  from transformers import AutoModelForImageTextToText, AutoTokenizer, PretrainedConfig
44
 
45
  tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
@@ -82,8 +82,8 @@ class SystemOne(SystemOneApi):
82
  model=None,
83
  tokenizer=None,
84
  ):
85
- """`model` and `tokenizer`, when given, are a backbone already loaded (`D1Model` passes itself);
86
- it stays on its device unless `device` says otherwise."""
87
  if model is None:
88
  model, tokenizer = load_backbone(model_id)
89
  else:
@@ -103,7 +103,8 @@ class SystemOne(SystemOneApi):
103
  self.option_style = option_style
104
  self.token_budget = token_budget
105
  self.processor = None
106
- # CUDA graphs for single questions on NVIDIA, where eager time is mostly kernel launches.
 
107
  if compile and torch.version.hip:
108
  raise ValueError("compile=True needs CUDA: on ROCm the CUDA graphs fault after a few dozen calls")
109
  self._one_pass = (torch.compile(self.model.forward, mode="reduce-overhead")
@@ -119,14 +120,19 @@ class SystemOne(SystemOneApi):
119
 
120
  # --------------------------------------------------------------- forward
121
 
122
- def _logz_ids(self, rows: list[list[int]]) -> list[torch.Tensor]:
123
- """Log-softmax at the answer slot, one row per token list, in one pass.
 
 
124
 
125
  The rows are one tree (`hybrid.py`): their common start is its trunk and
126
  is read once; the rest of each row is a branch, packed with no padding.
127
  Mathematically each row alone; in bf16 the kernels differ by batch shape.
128
  """
129
- if len(rows) == 1: # nothing to share: a plain chain is faster than a tree of one
 
 
 
130
  row = self._one_pass(input_ids=torch.tensor(rows, device=self.device), logits_to_keep=1).logits[0, -1]
131
  return [row.float() - torch.logsumexp(row.float(), dim=-1)]
132
  shared = 0 # every row keeps at least its last token
 
39
 
40
 
41
  def load_backbone(model_id: str = DEFAULT_MODEL, dtype=torch.bfloat16):
42
+ """The checkpoint and its tokenizer. SigLIP2 has no FA2 kernel here, so SDPA."""
43
  from transformers import AutoModelForImageTextToText, AutoTokenizer, PretrainedConfig
44
 
45
  tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
 
82
  model=None,
83
  tokenizer=None,
84
  ):
85
+ """`model` and `tokenizer`, when given, are a backbone already loaded (the Hub's remote-code class
86
+ passes itself); it stays on its device unless `device` says otherwise."""
87
  if model is None:
88
  model, tokenizer = load_backbone(model_id)
89
  else:
 
103
  self.option_style = option_style
104
  self.token_budget = token_budget
105
  self.processor = None
106
+ # CUDA graphs for single questions, on NVIDIA: where eager time is kernel launches, 16 to 8 ms a
107
+ # decision on an RTX 4090. The first two prompt lengths compile; later ones reuse the dynamic graph.
108
  if compile and torch.version.hip:
109
  raise ValueError("compile=True needs CUDA: on ROCm the CUDA graphs fault after a few dozen calls")
110
  self._one_pass = (torch.compile(self.model.forward, mode="reduce-overhead")
 
120
 
121
  # --------------------------------------------------------------- forward
122
 
123
+ @torch.inference_mode()
124
+ def _logz(self, texts: Sequence[str], questions: Sequence[Question] | None = None) -> list[torch.Tensor]:
125
+ """Log-softmax at the answer slot, one row per text, in one pass; the slot needs only the texts, and
126
+ `questions` is accepted for the scoring interface.
127
 
128
  The rows are one tree (`hybrid.py`): their common start is its trunk and
129
  is read once; the rest of each row is a branch, packed with no padding.
130
  Mathematically each row alone; in bf16 the kernels differ by batch shape.
131
  """
132
+ return self._logz_ids([self.tokenizer.encode(t, add_special_tokens=False) for t in texts])
133
+
134
+ def _logz_ids(self, rows: list[list[int]]) -> list[torch.Tensor]:
135
+ if len(rows) == 1: # nothing to share: a plain chain, 6-12% faster than a tree of one
136
  row = self._one_pass(input_ids=torch.tensor(rows, device=self.device), logits_to_keep=1).logits[0, -1]
137
  return [row.float() - torch.logsumexp(row.float(), dim=-1)]
138
  shared = 0 # every row keeps at least its last token