File size: 9,508 Bytes
1681a40
 
 
d8530d4
1681a40
 
 
 
 
d8530d4
 
 
 
 
 
 
 
 
 
 
 
 
1681a40
 
d8530d4
 
 
1681a40
 
d8530d4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1681a40
 
d8530d4
1681a40
03470f2
1681a40
03470f2
1681a40
 
b971e2b
1681a40
 
 
 
 
03470f2
1681a40
03470f2
 
1681a40
d8530d4
03470f2
 
1681a40
 
03470f2
 
1681a40
 
03470f2
1681a40
 
03470f2
 
 
1681a40
 
03470f2
1681a40
03470f2
1681a40
 
 
03470f2
d8530d4
 
 
 
 
 
 
 
 
 
 
 
 
 
1681a40
03470f2
1681a40
d8530d4
1681a40
d8530d4
 
 
 
 
 
1681a40
03470f2
 
 
 
 
 
 
 
 
 
 
1681a40
d8530d4
1681a40
03470f2
 
 
 
 
1681a40
03470f2
1681a40
03470f2
1681a40
d8530d4
1681a40
03470f2
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
---
language:
- en
license: apache-2.0
library_name: devision
pipeline_tag: visual-question-answering
base_model:
- convaiinnovations/laya
- google/siglip2-base-patch16-256
datasets:
- HuggingFaceM4/VQAv2
- lmms-lab/GQA
- HuggingFaceM4/A-OKVQA
- derek-thomas/ScienceQA
- HuggingFaceM4/the_cauldron
- cambridgeltl/vsr_random
- jxu124/objects365
- J1mb0o/e-SNLI-VE
- Vision-Flan/vision-flan_191-task_1k
- ryokamoi/VisOnlyQA_Train
- RyanWW/Super-CLEVR
- allenai/pixmo-count
tags:
- devision
- visual-decisions
- calibrated-probabilities
- jev
- system-one
- rlcd
- cpu
model-index:
- name: deVision v0.2
  results:
  - task:
      type: visual-question-answering
    dataset:
      name: COCO object presence
      type: test_exist
    metrics:
    - type: accuracy
      value: 0.9384
    - type: ece
      value: 0.0204
  - task:
      type: visual-question-answering
    dataset:
      name: VQAv2 multiple choice
      type: test_vqa_choice
    metrics:
    - type: accuracy
      value: 0.8979
    - type: ece
      value: 0.0262
  - task:
      type: visual-question-answering
    dataset:
      name: COCO size
      type: test_size
    metrics:
    - type: accuracy
      value: 0.86
    - type: ece
      value: 0.036
  - task:
      type: visual-question-answering
    dataset:
      name: POPE, project filtered set
      type: bench_pope
    metrics:
    - type: accuracy
      value: 0.8511
    - type: ece
      value: 0.0578
  - task:
      type: visual-question-answering
    dataset:
      name: GQA val subset
      type: test_gqa
    metrics:
    - type: accuracy
      value: 0.7843
    - type: ece
      value: 0.0419
  - task:
      type: visual-question-answering
    dataset:
      name: COCO position
      type: test_position
    metrics:
    - type: accuracy
      value: 0.8721
    - type: ece
      value: 0.0251
  - task:
      type: visual-question-answering
    dataset:
      name: VQAv2 yes/no subset
      type: test_vqa_yesno
    metrics:
    - type: accuracy
      value: 0.71
    - type: ece
      value: 0.0223
  - task:
      type: visual-question-answering
    dataset:
      name: COCO relative position
      type: test_relation
    metrics:
    - type: accuracy
      value: 0.7108
    - type: ece
      value: 0.0471
  - task:
      type: visual-question-answering
    dataset:
      name: VSR, project held-out split
      type: test_vsr
    metrics:
    - type: accuracy
      value: 0.6792
    - type: ece
      value: 0.0478
  - task:
      type: visual-question-answering
    dataset:
      name: Visual7W, project held-out split
      type: test_v7w
    metrics:
    - type: accuracy
      value: 0.721
    - type: ece
      value: 0.0302
  - task:
      type: visual-question-answering
    dataset:
      name: Fresh counting test (unseen pictures)
      type: test_count_fresh
    metrics:
    - type: accuracy
      value: 0.74
    - type: ece
      value: 0.0648
---

# deVision v0.2

**Image + English questions → structured answers with calibrated probabilities.** deVision pairs a SigLIP2 vision encoder with Laya's ModernBERT decision model. It scores the options of yes/no (`noul`) and multiple-choice (`choice`) questions instead of generating text, follows the Jev answer format, and runs without a GPU. Several questions about one image share a single image encoding.

## Usage

```bash
pip install devision
```

```python
import devision

model = devision.load("lukbit/devision")   # latest release; revision="v0.2" pins this one; device defaults to auto
result = model.predict(
    image="photo.jpg",   # a path, an http(s) URL, a data URI, base64, bytes or a PIL image
    state="",            # text context; "" when there is none
    questions={
        "has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"},
        "room": {"type": "choice", "instructions": "Which room is this?",
                 "criteria": {"kitchen": None, "bathroom": None, "bedroom": None}},
    },
)
print(result["answers"]["has_fork"]["noul"])        # probability of yes
print(result["answers"]["room"]["probabilities"])   # one probability per option, summing to 1
```

HTTP server with a browser demo at `http://127.0.0.1:8000`:

```bash
devision-serve --checkpoint lukbit/devision --port 8000
curl -s http://127.0.0.1:8000/v1/systemone -H 'Content-Type: application/json' \
  -d '{"image": "https://example.com/photo.jpg", "state": "", "questions": {"has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"}}}'
```

Over HTTP, `image` is an http(s) URL, a data URI or base64; the server never reads its own files. Responses follow Jev; `confidence` for `choice` is `(n * p_max - 1) / (n - 1)`. Invalid requests return 422 (`devision.InvalidRequest` in Python).

**Good for:** whether something is present and how many (ask counts as a choice), what kind of thing or scene, colours and materials, up/down and left/right in photos (as a choice). **Not for:** comparing two places on a diagram, chart or map; relative-position yes/no questions; small text; other languages or several images. Use the probabilities as thresholds and send uncertain cases to a person or a stronger model, after checking the thresholds on your own data.

## Evaluation

Accuracy through `decide` with the fitted temperatures. *Mismatched*: the same questions with every picture swapped for an unrelated one. Test sets use held-out pictures; they are project subsets, not official leaderboard scores.

| Set | Questions | Accuracy | Mismatched | ECE |
|---|---:|---:|---:|---:|
| COCO object presence | 1,120 | 0.938 | 0.493 | 0.020 |
| VQAv2 multiple choice | 1,420 | 0.898 | 0.399 | 0.026 |
| COCO size | 1,100 | 0.860 | 0.504 | 0.036 |
| POPE, project filtered set | 8,676 | 0.851 | 0.527 | 0.058 |
| GQA val subset | 992 | 0.784 | 0.532 | 0.042 |
| COCO position | 1,274 | 0.872 | 0.493 | 0.025 |
| VQAv2 yes/no subset | 1,000 | 0.710 | 0.518 | 0.022 |
| COCO relative position | 1,950 | 0.711 | 0.484 | 0.047 |
| VSR, project held-out split | 904 | 0.679 | 0.481 | 0.048 |
| Visual7W, project held-out split | 1,000 | 0.721 | 0.422 | 0.030 |
| Fresh counting test (unseen pictures) | 600 | 0.740 | 0.507 | 0.065 |

Against Laya Vision 201M on the same questions (its published per-question predictions; difference in points, 95% interval from paired resampling by picture):

| Set | Questions | Laya Vision | deVision | Difference |
|---|---:|---:|---:|---|
| VQAv2 yes/no | 4,887 | 0.717 | 0.725 | +0.8 [-0.9, +2.3] |
| A-OKVQA | 1,138 | 0.598 | 0.626 | +2.7 [-0.6, +6.0] |
| ScienceQA with images | 2,097 | 0.824 | 0.766 | -5.8 [-7.9, -3.6] |
| ScienceQA natural science, needs the picture | 323 | 0.700 | 0.455 | -24.5 [-31.6, -17.6] |

On the full POPE random / popular / adversarial sets (3,000 questions each) deVision scores **0.891 / 0.868 / 0.791**, against Laya Vision's published **0.836 / 0.819 / 0.777** (aggregate scores only, not paired).

CPU latency: P50 148 ms, P95 161 ms (Darwin arm64, 4 threads, FP32, one question per request, warm-up excluded).

## Model

| | |
|---|---|
| Vision | Frozen SigLIP2-B/16, 256 × 256 letterboxed input, 64 visual tokens after a 2 × 2 merge and an MLP projector |
| Decision | Laya-initialised ModernBERT-large and Laya's decision head; each option is scored at its `[MASK]`, then a temperature-scaled softmax |
| Size | About 518M parameters, FP32, 2.07 GB (LoRA merged) |

Training: the projector is first aligned on COCO captions, then the decision is trained with Laya's RLCD objective on yes/no and multiple-choice questions (projector, decision head and a LoRA on ModernBERT), round after round. This release continued for one pass over 353,830 questions, 315,342 of them from thirteen public datasets (Objects365, TallyQA, CLEVR, CLEVR-Math, Super-CLEVR, FigureQA, MapQA, IconQA, SNLI-VE, Vision-Flan, VisOnlyQA, SpatialSense, PixMo-Count) and the rest replaying earlier data (VQAv2, GQA, A-OKVQA, ScienceQA, VSR, Visual7W, COCO-derived questions). Temperatures are fitted per question type and option count. Details: the [training log](https://github.com/byebyebruce/devision/blob/master/docs/training-log.md).

## Limitations

- **Comparing two places named in the question** (two magnet poles, two chart series, two map regions) stays near chance; on the ScienceQA questions that need the picture it is about 24 points behind Laya Vision.
- **Relative-position yes/no questions** are weak: right on both a picture and its mirror for only 19% of pairs (65% for relative-position choice questions).
- **Presence leans towards "yes"**: it says an absent object is there more often than the previous release (POPE adversarial 0.791).
- English only, one image, `noul` and `choice` only; 256 × 256 input, so no small text.
- Calibration was fitted on this project's data and may not hold on yours.

## Licence

Weights and code: Apache-2.0, like the three base models ([Laya](https://huggingface.co/convaiinnovations/laya), [SigLIP2](https://huggingface.co/google/siglip2-base-patch16-256), [ModernBERT](https://huggingface.co/answerdotai/ModernBERT-large)). The training datasets (listed in this card's metadata) have their own terms, some non-commercial (e.g. ScienceQA, CC BY-NC-SA 4.0); check them against your use.

## Links

[Code and training records](https://github.com/byebyebruce/devision) (this release: tag `model-v0.2`) · [evaluation details](evaluation/results.md) · [provenance](provenance.json) · [Laya Vision](https://github.com/r33drichards/laya-vision)