deVision v0.2
Image + English questions → structured answers with calibrated probabilities. deVision pairs a SigLIP2 vision encoder with Laya's ModernBERT decision model. It scores the options of yes/no (noul) and multiple-choice (choice) questions instead of generating text, follows the Jev answer format, and runs without a GPU. Several questions about one image share a single image encoding.
Usage
pip install devision
import devision
model = devision.load("lukbit/devision") # latest release; revision="v0.2" pins this one; device defaults to auto
result = model.predict(
image="photo.jpg", # a path, an http(s) URL, a data URI, base64, bytes or a PIL image
state="", # text context; "" when there is none
questions={
"has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"},
"room": {"type": "choice", "instructions": "Which room is this?",
"criteria": {"kitchen": None, "bathroom": None, "bedroom": None}},
},
)
print(result["answers"]["has_fork"]["noul"]) # probability of yes
print(result["answers"]["room"]["probabilities"]) # one probability per option, summing to 1
HTTP server with a browser demo at http://127.0.0.1:8000:
devision-serve --checkpoint lukbit/devision --port 8000
curl -s http://127.0.0.1:8000/v1/systemone -H 'Content-Type: application/json' \
-d '{"image": "https://example.com/photo.jpg", "state": "", "questions": {"has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"}}}'
Over HTTP, image is an http(s) URL, a data URI or base64; the server never reads its own files. Responses follow Jev; confidence for choice is (n * p_max - 1) / (n - 1). Invalid requests return 422 (devision.InvalidRequest in Python).
Good for: whether something is present and how many (ask counts as a choice), what kind of thing or scene, colours and materials, up/down and left/right in photos (as a choice). Not for: comparing two places on a diagram, chart or map; relative-position yes/no questions; small text; other languages or several images. Use the probabilities as thresholds and send uncertain cases to a person or a stronger model, after checking the thresholds on your own data.
Evaluation
Accuracy through decide with the fitted temperatures. Mismatched: the same questions with every picture swapped for an unrelated one. Test sets use held-out pictures; they are project subsets, not official leaderboard scores.
| Set | Questions | Accuracy | Mismatched | ECE |
|---|---|---|---|---|
| COCO object presence | 1,120 | 0.938 | 0.493 | 0.020 |
| VQAv2 multiple choice | 1,420 | 0.898 | 0.399 | 0.026 |
| COCO size | 1,100 | 0.860 | 0.504 | 0.036 |
| POPE, project filtered set | 8,676 | 0.851 | 0.527 | 0.058 |
| GQA val subset | 992 | 0.784 | 0.532 | 0.042 |
| COCO position | 1,274 | 0.872 | 0.493 | 0.025 |
| VQAv2 yes/no subset | 1,000 | 0.710 | 0.518 | 0.022 |
| COCO relative position | 1,950 | 0.711 | 0.484 | 0.047 |
| VSR, project held-out split | 904 | 0.679 | 0.481 | 0.048 |
| Visual7W, project held-out split | 1,000 | 0.721 | 0.422 | 0.030 |
| Fresh counting test (unseen pictures) | 600 | 0.740 | 0.507 | 0.065 |
Against Laya Vision 201M on the same questions (its published per-question predictions; difference in points, 95% interval from paired resampling by picture):
| Set | Questions | Laya Vision | deVision | Difference |
|---|---|---|---|---|
| VQAv2 yes/no | 4,887 | 0.717 | 0.725 | +0.8 [-0.9, +2.3] |
| A-OKVQA | 1,138 | 0.598 | 0.626 | +2.7 [-0.6, +6.0] |
| ScienceQA with images | 2,097 | 0.824 | 0.766 | -5.8 [-7.9, -3.6] |
| ScienceQA natural science, needs the picture | 323 | 0.700 | 0.455 | -24.5 [-31.6, -17.6] |
On the full POPE random / popular / adversarial sets (3,000 questions each) deVision scores 0.891 / 0.868 / 0.791, against Laya Vision's published 0.836 / 0.819 / 0.777 (aggregate scores only, not paired).
CPU latency: P50 148 ms, P95 161 ms (Darwin arm64, 4 threads, FP32, one question per request, warm-up excluded).
Model
| Vision | Frozen SigLIP2-B/16, 256 × 256 letterboxed input, 64 visual tokens after a 2 × 2 merge and an MLP projector |
| Decision | Laya-initialised ModernBERT-large and Laya's decision head; each option is scored at its [MASK], then a temperature-scaled softmax |
| Size | About 518M parameters, FP32, 2.07 GB (LoRA merged) |
Training: the projector is first aligned on COCO captions, then the decision is trained with Laya's RLCD objective on yes/no and multiple-choice questions (projector, decision head and a LoRA on ModernBERT), round after round. This release continued for one pass over 353,830 questions, 315,342 of them from thirteen public datasets (Objects365, TallyQA, CLEVR, CLEVR-Math, Super-CLEVR, FigureQA, MapQA, IconQA, SNLI-VE, Vision-Flan, VisOnlyQA, SpatialSense, PixMo-Count) and the rest replaying earlier data (VQAv2, GQA, A-OKVQA, ScienceQA, VSR, Visual7W, COCO-derived questions). Temperatures are fitted per question type and option count. Details: the training log.
Limitations
- Comparing two places named in the question (two magnet poles, two chart series, two map regions) stays near chance; on the ScienceQA questions that need the picture it is about 24 points behind Laya Vision.
- Relative-position yes/no questions are weak: right on both a picture and its mirror for only 19% of pairs (65% for relative-position choice questions).
- Presence leans towards "yes": it says an absent object is there more often than the previous release (POPE adversarial 0.791).
- English only, one image,
noulandchoiceonly; 256 × 256 input, so no small text. - Calibration was fitted on this project's data and may not hold on yours.
Licence
Weights and code: Apache-2.0, like the three base models (Laya, SigLIP2, ModernBERT). The training datasets (listed in this card's metadata) have their own terms, some non-commercial (e.g. ScienceQA, CC BY-NC-SA 4.0); check them against your use.
Links
Code and training records (this release: tag model-v0.2) · evaluation details · provenance · Laya Vision
- Downloads last month
- 27
Model tree for lukbit/devision
Base model
convaiinnovations/layaDatasets used to train lukbit/devision
lmms-lab-encoder/GQA
derek-thomas/ScienceQA
Evaluation results
- accuracy on COCO object presenceself-reported0.938
- ece on COCO object presenceself-reported0.020
- accuracy on VQAv2 multiple choiceself-reported0.898
- ece on VQAv2 multiple choiceself-reported0.026
- accuracy on COCO sizeself-reported0.860
- ece on COCO sizeself-reported0.036
- accuracy on POPE, project filtered setself-reported0.851
- ece on POPE, project filtered setself-reported0.058