lukbit commited on
Commit
03470f2
·
verified ·
1 Parent(s): d8530d4

deVision v0.2 card (github 458d9974ff8771520d2298926520dc235a398ed2)

Browse files
Files changed (2) hide show
  1. MANIFEST.json +3 -3
  2. README.md +36 -110
MANIFEST.json CHANGED
@@ -8,8 +8,8 @@
8
  "sha256": "cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30"
9
  },
10
  "README.md": {
11
- "size": 13520,
12
- "sha256": "584f165e73880275bb92348107839e7483a012ef3fa23ea7a6717576e97c05c4"
13
  },
14
  "config.json": {
15
  "size": 353,
@@ -33,7 +33,7 @@
33
  },
34
  "provenance.json": {
35
  "size": 1324,
36
- "sha256": "94d064129af62edafbcf05396b0a0b589034218f829070967039e8a3126e152c"
37
  },
38
  "tokenizer/tokenizer.json": {
39
  "size": 3583228,
 
8
  "sha256": "cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30"
9
  },
10
  "README.md": {
11
+ "size": 9556,
12
+ "sha256": "44f335c6cc7e9fa1cf33d6f24aed4269f78dd114e8cf0abfb9598b636f2b9462"
13
  },
14
  "config.json": {
15
  "size": 353,
 
33
  },
34
  "provenance.json": {
35
  "size": 1324,
36
+ "sha256": "99cd0254a1bc35a09b46beef7280ce5243379638eba0125d9a039c222efe654c"
37
  },
38
  "tokenizer/tokenizer.json": {
39
  "size": 3583228,
README.md CHANGED
@@ -145,125 +145,46 @@ model-index:
145
 
146
  # deVision v0.2
147
 
148
- **Image + typed questions → structured answers and calibrated probabilities.** deVision is a non-autoregressive visual decision model built from Laya's ModernBERT-large decision model and a SigLIP2 vision encoder. It scores the supplied options without generating text. Multiple questions about one image share a single image encoding.
149
 
150
- This is release **v0.2**. It supports English, one image per request, yes/no decisions (`noul`) and multiple-choice decisions (`choice`). Weights and code are released under the Apache-2.0 licence; see [Licence and data](#licence-and-data).
151
-
152
- ## Install
153
 
154
  ```bash
155
  pip install "devision @ git+https://github.com/byebyebruce/devision"
156
  ```
157
 
158
- Python 3.11 or newer is required; the install includes the HTTP server and browser demo. If the model repository is private, authenticate first with `hf auth login` using an account that has access.
159
-
160
- ## Python quickstart
161
-
162
  ```python
163
  import devision
164
 
165
- model = devision.load("lukbit/devision", revision="v0.2", device="cpu")
166
  result = model.predict(
167
- image="photo.jpg", # an image path, an http(s) URL, base64, bytes or a PIL image
168
- state="", # text context; empty when there is none
169
  questions={
170
  "has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"},
171
- "room": {
172
- "type": "choice",
173
- "instructions": "Which room is this?",
174
- "criteria": {"kitchen": None, "bathroom": None, "bedroom": None},
175
- },
176
  },
177
  )
178
- print(result["answers"]["has_fork"]["noul"]) # probability of yes
179
- print(result["answers"]["room"]["choice"]) # selected option
180
- print(result["answers"]["room"]["probabilities"]) # probability of each option
181
  ```
182
 
183
- `model.decide(...)` takes the same arguments. Pin `revision="v0.2"` so that later releases do not change your results. The default device is `auto` (CUDA, then MPS, then CPU).
184
-
185
- ### Image input
186
-
187
- The Python `image=` argument accepts:
188
-
189
- | Input | Example |
190
- |---|---|
191
- | HTTP(S) URL | `image="https://example.com/photo.jpg"` |
192
- | Local path | `image="photo.jpg"` or `image=Path("photo.jpg")` |
193
- | Base64 data URI | `image="data:image/png;base64,..."` |
194
- | Plain base64 | `image=base64_string` |
195
- | Encoded image bytes | `image=image_bytes` (PNG / JPEG file contents) |
196
- | PIL image | `image=pil_image` |
197
-
198
- ### Context
199
-
200
- `state` carries text context as a string, object or array, exactly as in Jev:
201
-
202
- ```python
203
- result = model.decide(
204
- image="photo.jpg",
205
- state={"note": "The customer says this item arrived damaged."},
206
- questions={"damaged": {"type": "noul", "instructions": "Is the item visibly damaged?"}},
207
- )
208
- ```
209
-
210
- An `image` key inside an object state stays text: `state={"image": "photo.jpg"}` does not load the file. The Jev-compatible image part `state=[{"type": "image", "url": "https://..."}]` also works; do not combine it with `image=`.
211
-
212
- ### Output
213
-
214
- - `noul`: the probability that the statement holds, between 0 and 1.
215
- - `choice`: the highest-probability option; `probabilities` covers every option and sums to 1.
216
- - `confidence` (choice) follows Jev: `(n * p_max - 1) / (n - 1)`.
217
- - Invalid requests raise `devision.InvalidRequest`.
218
-
219
- ## HTTP API and browser demo
220
 
221
  ```bash
222
- pip install "devision @ git+https://github.com/byebyebruce/devision"
223
- devision-serve --checkpoint lukbit/devision --device cpu --port 8000
 
224
  ```
225
 
226
- Open `http://127.0.0.1:8000` for the demo. The server speaks the Jev-compatible `POST /v1/systemone` format; an image is an element of the `state` array:
227
-
228
- ```json
229
- {"state": [{"type": "image", "url": "https://example.com/photo.jpg"}, "optional text"],
230
- "questions": {"has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"}}}
231
- ```
232
-
233
- Use `{"type": "image", "base64": "..."}` for an uploaded picture. HTTP never reads server-local paths. Invalid requests and `score` questions return HTTP 422.
234
-
235
- ## Architecture
236
-
237
- | Component | |
238
- |---|---|
239
- | Vision encoder | Frozen SigLIP2-B/16, 256 × 256 input, about 93M parameters |
240
- | Image preprocessing | Aspect-preserving resize and letterbox padding |
241
- | Connector | 2 × 2 patch grouping, then LayerNorm → Linear → GELU → Linear |
242
- | Visual sequence | 64 visual tokens, inserted after `[CLS]` |
243
- | Text encoder | Laya-initialised ModernBERT-large, about 395M parameters |
244
- | Decision head | Laya's two Transformer layers and option scorer, about 27M parameters |
245
- | Checkpoint | About 518M parameters, FP32, 2.07 GB of weights (LoRA merged) |
246
-
247
- Each option is scored at its `[MASK]` position; a temperature-scaled softmax turns the scores into probabilities.
248
-
249
- ## Training and calibration
250
 
251
- deVision is trained in two stages. Stage 1 aligns the projector on COCO captions (image-conditioned masked words). Stage 2 trains the decision on yes / no and multiple-choice questions with Laya's RLCD objective (projector, decision head and a LoRA on ModernBERT; the LoRA is merged for release). Earlier training covered COCO-derived existence / position / size questions, VQAv2, GQA, A-OKVQA, ScienceQA, AI2D, TQA, VSR, Visual7W telling, mirrored left/right pairs and counting. The last training step continued from that checkpoint for one pass over 353,830 questions: an extension pack of 315,342 questions from thirteen public sources (Objects365, TallyQA, CLEVR, CLEVR-Math, Super-CLEVR, FigureQA, MapQA, IconQA, SNLI-VE, Vision-Flan, VisOnlyQA, SpatialSense, PixMo-Count) plus 38,488 replayed questions of the earlier abilities; 44,229 steps on Apple MPS in FP32, learning rates `5e-5` (projector, head) and `1e-4` (LoRA), 500 warm-up steps. Full configurations and results: the [training log](https://github.com/byebyebruce/devision/blob/master/docs/training-log.md).
252
-
253
- Temperatures fitted on the project's calibration set (a question's type / option-count bucket first, else its type):
254
-
255
- | Bucket | Temperature |
256
- |---|---:|
257
- | Choice, 2 options | 1.8836 |
258
- | Choice, 3–5 options | 1.9790 |
259
- | Noul (yes / no) | 1.6382 |
260
- | Choice, any other count | 1.9568 |
261
-
262
- Calibration changes probabilities, not visual ability. Validate confidence thresholds on your own data.
263
 
264
  ## Evaluation
265
 
266
- Results of this checkpoint through `decide`, with the fitted temperatures; ECE uses 15 bins. *Mismatched* is the accuracy when every picture is swapped for an unrelated one (how much the answer depends on the picture). The project test sets use held-out pictures; they are project-specific subsets or generated questions, not official leaderboard scores, and several have informed decisions across training rounds.
267
 
268
  | Set | Questions | Accuracy | Mismatched | ECE |
269
  |---|---:|---:|---:|---:|
@@ -279,9 +200,7 @@ Results of this checkpoint through `decide`, with the fitted temperatures; ECE u
279
  | Visual7W, project held-out split | 1,000 | 0.721 | 0.422 | 0.030 |
280
  | Fresh counting test (unseen pictures) | 600 | 0.740 | 0.507 | 0.065 |
281
 
282
- ### Comparison with Laya Vision 201M
283
-
284
- Laya Vision's published per-question predictions on the same questions; differences in percentage points with 95% intervals from paired resampling by picture. We did not rerun Laya Vision.
285
 
286
  | Set | Questions | Laya Vision | deVision | Difference |
287
  |---|---:|---:|---:|---|
@@ -292,23 +211,30 @@ Laya Vision's published per-question predictions on the same questions; differen
292
 
293
  On the full POPE random / popular / adversarial sets (3,000 questions each) deVision scores **0.891 / 0.868 / 0.791**, against Laya Vision's published **0.836 / 0.819 / 0.777** (aggregate scores only, not paired).
294
 
295
- CPU latency: P50 148 ms, P95 161 ms (Darwin arm64, 4 threads, FP32, one question per request, warm-up excluded). Several questions about one picture in one request share the image encoding, so each extra question costs less.
 
 
 
 
 
 
 
 
 
 
296
 
297
  ## Limitations
298
 
299
- - **Scientific diagrams and charts remain difficult.** On the 323 ScienceQA natural-science questions that need the picture, this release is about 24 points behind Laya Vision (see the comparison table). Comparing two named regions of a picture (two magnet poles, two series of a chart, two regions of a map, two line segments) is the main open weakness; training on tens of thousands of such questions left it near chance.
300
- - **Spatial reasoning is incomplete.** A left/right question is answered right on both a picture and its mirror for about 72% of position pairs and 65% of relative-position choice pairs, but only 19% of relative-position yes / no pairs.
301
- - **Object presence leans towards "yes".** Compared with the previous checkpoint, it more often says an object is present when the annotation says it is not (COCO existence test 95.5% → 93.8%; POPE adversarial 79.1%).
302
- - **Small text and fine detail** are limited by the 256 × 256 input; this is not an OCR model.
303
- - **English, one image, `noul` and `choice` only.** Other languages, several images and `score` questions are not supported.
304
- - **Probabilities can be miscalibrated out of domain.** A low ECE on these benchmarks does not guarantee reliable confidence on your data.
305
 
306
- ## Licence and data
307
 
308
- The weights and the code are released under the **Apache-2.0** licence, the licence of the three base models ([Laya](https://huggingface.co/convaiinnovations/laya), [SigLIP2](https://huggingface.co/google/siglip2-base-patch16-256), [ModernBERT](https://huggingface.co/answerdotai/ModernBERT-large)). The training data come from public datasets with their own terms, some of them non-commercial (for example ScienceQA, CC BY-NC-SA 4.0) or covering the pictures separately (COCO / Flickr images). Whether such terms carry over to trained weights is not settled; check the datasets listed in this card's metadata against your use.
309
 
310
  ## Links
311
 
312
- - [Code, training configurations and experiment records](https://github.com/byebyebruce/devision) — this release is tag `model-v0.2`
313
- - [Evaluation details](evaluation/results.md) · [metrics](evaluation/results.json) · [provenance](provenance.json)
314
- - [Laya decision model](https://huggingface.co/convaiinnovations/laya) · [SigLIP2](https://huggingface.co/google/siglip2-base-patch16-256) · [Laya Vision](https://github.com/r33drichards/laya-vision)
 
145
 
146
  # deVision v0.2
147
 
148
+ **Image + English questions → structured answers with calibrated probabilities.** deVision pairs a SigLIP2 vision encoder with Laya's ModernBERT decision model. It scores the options of yes/no (`noul`) and multiple-choice (`choice`) questions instead of generating text, follows the Jev answer format, and runs without a GPU. Several questions about one image share a single image encoding.
149
 
150
+ ## Usage
 
 
151
 
152
  ```bash
153
  pip install "devision @ git+https://github.com/byebyebruce/devision"
154
  ```
155
 
 
 
 
 
156
  ```python
157
  import devision
158
 
159
+ model = devision.load("lukbit/devision") # latest release; revision="v0.2" pins this one; device defaults to auto
160
  result = model.predict(
161
+ image="photo.jpg", # a path, an http(s) URL, a data URI, base64, bytes or a PIL image
162
+ state="", # text context; "" when there is none
163
  questions={
164
  "has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"},
165
+ "room": {"type": "choice", "instructions": "Which room is this?",
166
+ "criteria": {"kitchen": None, "bathroom": None, "bedroom": None}},
 
 
 
167
  },
168
  )
169
+ print(result["answers"]["has_fork"]["noul"]) # probability of yes
170
+ print(result["answers"]["room"]["probabilities"]) # one probability per option, summing to 1
 
171
  ```
172
 
173
+ HTTP server with a browser demo at `http://127.0.0.1:8000`:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
174
 
175
  ```bash
176
+ devision-serve --checkpoint lukbit/devision --port 8000
177
+ curl -s http://127.0.0.1:8000/v1/systemone -H 'Content-Type: application/json' \
178
+ -d '{"image": "https://example.com/photo.jpg", "state": "", "questions": {"has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"}}}'
179
  ```
180
 
181
+ Over HTTP, `image` is an http(s) URL, a data URI or base64; the server never reads its own files. Responses follow Jev; `confidence` for `choice` is `(n * p_max - 1) / (n - 1)`. Invalid requests return 422 (`devision.InvalidRequest` in Python).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
182
 
183
+ **Good for:** whether something is present and how many (ask counts as a choice), what kind of thing or scene, colours and materials, up/down and left/right in photos (as a choice). **Not for:** comparing two places on a diagram, chart or map; relative-position yes/no questions; small text; other languages or several images. Use the probabilities as thresholds and send uncertain cases to a person or a stronger model, after checking the thresholds on your own data.
 
 
 
 
 
 
 
 
 
 
 
184
 
185
  ## Evaluation
186
 
187
+ Accuracy through `decide` with the fitted temperatures. *Mismatched*: the same questions with every picture swapped for an unrelated one. Test sets use held-out pictures; they are project subsets, not official leaderboard scores.
188
 
189
  | Set | Questions | Accuracy | Mismatched | ECE |
190
  |---|---:|---:|---:|---:|
 
200
  | Visual7W, project held-out split | 1,000 | 0.721 | 0.422 | 0.030 |
201
  | Fresh counting test (unseen pictures) | 600 | 0.740 | 0.507 | 0.065 |
202
 
203
+ Against Laya Vision 201M on the same questions (its published per-question predictions; difference in points, 95% interval from paired resampling by picture):
 
 
204
 
205
  | Set | Questions | Laya Vision | deVision | Difference |
206
  |---|---:|---:|---:|---|
 
211
 
212
  On the full POPE random / popular / adversarial sets (3,000 questions each) deVision scores **0.891 / 0.868 / 0.791**, against Laya Vision's published **0.836 / 0.819 / 0.777** (aggregate scores only, not paired).
213
 
214
+ CPU latency: P50 148 ms, P95 161 ms (Darwin arm64, 4 threads, FP32, one question per request, warm-up excluded).
215
+
216
+ ## Model
217
+
218
+ | | |
219
+ |---|---|
220
+ | Vision | Frozen SigLIP2-B/16, 256 × 256 letterboxed input, 64 visual tokens after a 2 × 2 merge and an MLP projector |
221
+ | Decision | Laya-initialised ModernBERT-large and Laya's decision head; each option is scored at its `[MASK]`, then a temperature-scaled softmax |
222
+ | Size | About 518M parameters, FP32, 2.07 GB (LoRA merged) |
223
+
224
+ Training: the projector is first aligned on COCO captions, then the decision is trained with Laya's RLCD objective on yes/no and multiple-choice questions (projector, decision head and a LoRA on ModernBERT), round after round. This release continued for one pass over 353,830 questions, 315,342 of them from thirteen public datasets (Objects365, TallyQA, CLEVR, CLEVR-Math, Super-CLEVR, FigureQA, MapQA, IconQA, SNLI-VE, Vision-Flan, VisOnlyQA, SpatialSense, PixMo-Count) and the rest replaying earlier data (VQAv2, GQA, A-OKVQA, ScienceQA, VSR, Visual7W, COCO-derived questions). Temperatures are fitted per question type and option count. Details: the [training log](https://github.com/byebyebruce/devision/blob/master/docs/training-log.md).
225
 
226
  ## Limitations
227
 
228
+ - **Comparing two places named in the question** (two magnet poles, two chart series, two map regions) stays near chance; on the ScienceQA questions that need the picture it is about 24 points behind Laya Vision.
229
+ - **Relative-position yes/no questions** are weak: right on both a picture and its mirror for only 19% of pairs (65% for relative-position choice questions).
230
+ - **Presence leans towards "yes"**: it says an absent object is there more often than the previous release (POPE adversarial 0.791).
231
+ - English only, one image, `noul` and `choice` only; 256 × 256 input, so no small text.
232
+ - Calibration was fitted on this project's data and may not hold on yours.
 
233
 
234
+ ## Licence
235
 
236
+ Weights and code: Apache-2.0, like the three base models ([Laya](https://huggingface.co/convaiinnovations/laya), [SigLIP2](https://huggingface.co/google/siglip2-base-patch16-256), [ModernBERT](https://huggingface.co/answerdotai/ModernBERT-large)). The training datasets (listed in this card's metadata) have their own terms, some non-commercial (e.g. ScienceQA, CC BY-NC-SA 4.0); check them against your use.
237
 
238
  ## Links
239
 
240
+ [Code and training records](https://github.com/byebyebruce/devision) (this release: tag `model-v0.2`) · [evaluation details](evaluation/results.md) · [provenance](provenance.json) · [Laya Vision](https://github.com/r33drichards/laya-vision)