File size: 9,508 Bytes
1681a40 d8530d4 1681a40 d8530d4 1681a40 d8530d4 1681a40 d8530d4 1681a40 d8530d4 1681a40 03470f2 1681a40 03470f2 1681a40 b971e2b 1681a40 03470f2 1681a40 03470f2 1681a40 d8530d4 03470f2 1681a40 03470f2 1681a40 03470f2 1681a40 03470f2 1681a40 03470f2 1681a40 03470f2 1681a40 03470f2 d8530d4 1681a40 03470f2 1681a40 d8530d4 1681a40 d8530d4 1681a40 03470f2 1681a40 d8530d4 1681a40 03470f2 1681a40 03470f2 1681a40 03470f2 1681a40 d8530d4 1681a40 03470f2 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 | ---
language:
- en
license: apache-2.0
library_name: devision
pipeline_tag: visual-question-answering
base_model:
- convaiinnovations/laya
- google/siglip2-base-patch16-256
datasets:
- HuggingFaceM4/VQAv2
- lmms-lab/GQA
- HuggingFaceM4/A-OKVQA
- derek-thomas/ScienceQA
- HuggingFaceM4/the_cauldron
- cambridgeltl/vsr_random
- jxu124/objects365
- J1mb0o/e-SNLI-VE
- Vision-Flan/vision-flan_191-task_1k
- ryokamoi/VisOnlyQA_Train
- RyanWW/Super-CLEVR
- allenai/pixmo-count
tags:
- devision
- visual-decisions
- calibrated-probabilities
- jev
- system-one
- rlcd
- cpu
model-index:
- name: deVision v0.2
results:
- task:
type: visual-question-answering
dataset:
name: COCO object presence
type: test_exist
metrics:
- type: accuracy
value: 0.9384
- type: ece
value: 0.0204
- task:
type: visual-question-answering
dataset:
name: VQAv2 multiple choice
type: test_vqa_choice
metrics:
- type: accuracy
value: 0.8979
- type: ece
value: 0.0262
- task:
type: visual-question-answering
dataset:
name: COCO size
type: test_size
metrics:
- type: accuracy
value: 0.86
- type: ece
value: 0.036
- task:
type: visual-question-answering
dataset:
name: POPE, project filtered set
type: bench_pope
metrics:
- type: accuracy
value: 0.8511
- type: ece
value: 0.0578
- task:
type: visual-question-answering
dataset:
name: GQA val subset
type: test_gqa
metrics:
- type: accuracy
value: 0.7843
- type: ece
value: 0.0419
- task:
type: visual-question-answering
dataset:
name: COCO position
type: test_position
metrics:
- type: accuracy
value: 0.8721
- type: ece
value: 0.0251
- task:
type: visual-question-answering
dataset:
name: VQAv2 yes/no subset
type: test_vqa_yesno
metrics:
- type: accuracy
value: 0.71
- type: ece
value: 0.0223
- task:
type: visual-question-answering
dataset:
name: COCO relative position
type: test_relation
metrics:
- type: accuracy
value: 0.7108
- type: ece
value: 0.0471
- task:
type: visual-question-answering
dataset:
name: VSR, project held-out split
type: test_vsr
metrics:
- type: accuracy
value: 0.6792
- type: ece
value: 0.0478
- task:
type: visual-question-answering
dataset:
name: Visual7W, project held-out split
type: test_v7w
metrics:
- type: accuracy
value: 0.721
- type: ece
value: 0.0302
- task:
type: visual-question-answering
dataset:
name: Fresh counting test (unseen pictures)
type: test_count_fresh
metrics:
- type: accuracy
value: 0.74
- type: ece
value: 0.0648
---
# deVision v0.2
**Image + English questions → structured answers with calibrated probabilities.** deVision pairs a SigLIP2 vision encoder with Laya's ModernBERT decision model. It scores the options of yes/no (`noul`) and multiple-choice (`choice`) questions instead of generating text, follows the Jev answer format, and runs without a GPU. Several questions about one image share a single image encoding.
## Usage
```bash
pip install devision
```
```python
import devision
model = devision.load("lukbit/devision") # latest release; revision="v0.2" pins this one; device defaults to auto
result = model.predict(
image="photo.jpg", # a path, an http(s) URL, a data URI, base64, bytes or a PIL image
state="", # text context; "" when there is none
questions={
"has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"},
"room": {"type": "choice", "instructions": "Which room is this?",
"criteria": {"kitchen": None, "bathroom": None, "bedroom": None}},
},
)
print(result["answers"]["has_fork"]["noul"]) # probability of yes
print(result["answers"]["room"]["probabilities"]) # one probability per option, summing to 1
```
HTTP server with a browser demo at `http://127.0.0.1:8000`:
```bash
devision-serve --checkpoint lukbit/devision --port 8000
curl -s http://127.0.0.1:8000/v1/systemone -H 'Content-Type: application/json' \
-d '{"image": "https://example.com/photo.jpg", "state": "", "questions": {"has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"}}}'
```
Over HTTP, `image` is an http(s) URL, a data URI or base64; the server never reads its own files. Responses follow Jev; `confidence` for `choice` is `(n * p_max - 1) / (n - 1)`. Invalid requests return 422 (`devision.InvalidRequest` in Python).
**Good for:** whether something is present and how many (ask counts as a choice), what kind of thing or scene, colours and materials, up/down and left/right in photos (as a choice). **Not for:** comparing two places on a diagram, chart or map; relative-position yes/no questions; small text; other languages or several images. Use the probabilities as thresholds and send uncertain cases to a person or a stronger model, after checking the thresholds on your own data.
## Evaluation
Accuracy through `decide` with the fitted temperatures. *Mismatched*: the same questions with every picture swapped for an unrelated one. Test sets use held-out pictures; they are project subsets, not official leaderboard scores.
| Set | Questions | Accuracy | Mismatched | ECE |
|---|---:|---:|---:|---:|
| COCO object presence | 1,120 | 0.938 | 0.493 | 0.020 |
| VQAv2 multiple choice | 1,420 | 0.898 | 0.399 | 0.026 |
| COCO size | 1,100 | 0.860 | 0.504 | 0.036 |
| POPE, project filtered set | 8,676 | 0.851 | 0.527 | 0.058 |
| GQA val subset | 992 | 0.784 | 0.532 | 0.042 |
| COCO position | 1,274 | 0.872 | 0.493 | 0.025 |
| VQAv2 yes/no subset | 1,000 | 0.710 | 0.518 | 0.022 |
| COCO relative position | 1,950 | 0.711 | 0.484 | 0.047 |
| VSR, project held-out split | 904 | 0.679 | 0.481 | 0.048 |
| Visual7W, project held-out split | 1,000 | 0.721 | 0.422 | 0.030 |
| Fresh counting test (unseen pictures) | 600 | 0.740 | 0.507 | 0.065 |
Against Laya Vision 201M on the same questions (its published per-question predictions; difference in points, 95% interval from paired resampling by picture):
| Set | Questions | Laya Vision | deVision | Difference |
|---|---:|---:|---:|---|
| VQAv2 yes/no | 4,887 | 0.717 | 0.725 | +0.8 [-0.9, +2.3] |
| A-OKVQA | 1,138 | 0.598 | 0.626 | +2.7 [-0.6, +6.0] |
| ScienceQA with images | 2,097 | 0.824 | 0.766 | -5.8 [-7.9, -3.6] |
| ScienceQA natural science, needs the picture | 323 | 0.700 | 0.455 | -24.5 [-31.6, -17.6] |
On the full POPE random / popular / adversarial sets (3,000 questions each) deVision scores **0.891 / 0.868 / 0.791**, against Laya Vision's published **0.836 / 0.819 / 0.777** (aggregate scores only, not paired).
CPU latency: P50 148 ms, P95 161 ms (Darwin arm64, 4 threads, FP32, one question per request, warm-up excluded).
## Model
| | |
|---|---|
| Vision | Frozen SigLIP2-B/16, 256 × 256 letterboxed input, 64 visual tokens after a 2 × 2 merge and an MLP projector |
| Decision | Laya-initialised ModernBERT-large and Laya's decision head; each option is scored at its `[MASK]`, then a temperature-scaled softmax |
| Size | About 518M parameters, FP32, 2.07 GB (LoRA merged) |
Training: the projector is first aligned on COCO captions, then the decision is trained with Laya's RLCD objective on yes/no and multiple-choice questions (projector, decision head and a LoRA on ModernBERT), round after round. This release continued for one pass over 353,830 questions, 315,342 of them from thirteen public datasets (Objects365, TallyQA, CLEVR, CLEVR-Math, Super-CLEVR, FigureQA, MapQA, IconQA, SNLI-VE, Vision-Flan, VisOnlyQA, SpatialSense, PixMo-Count) and the rest replaying earlier data (VQAv2, GQA, A-OKVQA, ScienceQA, VSR, Visual7W, COCO-derived questions). Temperatures are fitted per question type and option count. Details: the [training log](https://github.com/byebyebruce/devision/blob/master/docs/training-log.md).
## Limitations
- **Comparing two places named in the question** (two magnet poles, two chart series, two map regions) stays near chance; on the ScienceQA questions that need the picture it is about 24 points behind Laya Vision.
- **Relative-position yes/no questions** are weak: right on both a picture and its mirror for only 19% of pairs (65% for relative-position choice questions).
- **Presence leans towards "yes"**: it says an absent object is there more often than the previous release (POPE adversarial 0.791).
- English only, one image, `noul` and `choice` only; 256 × 256 input, so no small text.
- Calibration was fitted on this project's data and may not hold on yours.
## Licence
Weights and code: Apache-2.0, like the three base models ([Laya](https://huggingface.co/convaiinnovations/laya), [SigLIP2](https://huggingface.co/google/siglip2-base-patch16-256), [ModernBERT](https://huggingface.co/answerdotai/ModernBERT-large)). The training datasets (listed in this card's metadata) have their own terms, some non-commercial (e.g. ScienceQA, CC BY-NC-SA 4.0); check them against your use.
## Links
[Code and training records](https://github.com/byebyebruce/devision) (this release: tag `model-v0.2`) · [evaluation details](evaluation/results.md) · [provenance](provenance.json) · [Laya Vision](https://github.com/r33drichards/laya-vision)
|