deVision v0.2

Image + English questions → structured answers with calibrated probabilities. deVision pairs a SigLIP2 vision encoder with Laya's ModernBERT decision model. It scores the options of yes/no (noul) and multiple-choice (choice) questions instead of generating text, follows the Jev answer format, and runs without a GPU. Several questions about one image share a single image encoding.

Usage

pip install devision
import devision

model = devision.load("lukbit/devision")   # latest release; revision="v0.2" pins this one; device defaults to auto
result = model.predict(
    image="photo.jpg",   # a path, an http(s) URL, a data URI, base64, bytes or a PIL image
    state="",            # text context; "" when there is none
    questions={
        "has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"},
        "room": {"type": "choice", "instructions": "Which room is this?",
                 "criteria": {"kitchen": None, "bathroom": None, "bedroom": None}},
    },
)
print(result["answers"]["has_fork"]["noul"])        # probability of yes
print(result["answers"]["room"]["probabilities"])   # one probability per option, summing to 1

HTTP server with a browser demo at http://127.0.0.1:8000:

devision-serve --checkpoint lukbit/devision --port 8000
curl -s http://127.0.0.1:8000/v1/systemone -H 'Content-Type: application/json' \
  -d '{"image": "https://example.com/photo.jpg", "state": "", "questions": {"has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"}}}'

Over HTTP, image is an http(s) URL, a data URI or base64; the server never reads its own files. Responses follow Jev; confidence for choice is (n * p_max - 1) / (n - 1). Invalid requests return 422 (devision.InvalidRequest in Python).

Good for: whether something is present and how many (ask counts as a choice), what kind of thing or scene, colours and materials, up/down and left/right in photos (as a choice). Not for: comparing two places on a diagram, chart or map; relative-position yes/no questions; small text; other languages or several images. Use the probabilities as thresholds and send uncertain cases to a person or a stronger model, after checking the thresholds on your own data.

Evaluation

Accuracy through decide with the fitted temperatures. Mismatched: the same questions with every picture swapped for an unrelated one. Test sets use held-out pictures; they are project subsets, not official leaderboard scores.

Set Questions Accuracy Mismatched ECE
COCO object presence 1,120 0.938 0.493 0.020
VQAv2 multiple choice 1,420 0.898 0.399 0.026
COCO size 1,100 0.860 0.504 0.036
POPE, project filtered set 8,676 0.851 0.527 0.058
GQA val subset 992 0.784 0.532 0.042
COCO position 1,274 0.872 0.493 0.025
VQAv2 yes/no subset 1,000 0.710 0.518 0.022
COCO relative position 1,950 0.711 0.484 0.047
VSR, project held-out split 904 0.679 0.481 0.048
Visual7W, project held-out split 1,000 0.721 0.422 0.030
Fresh counting test (unseen pictures) 600 0.740 0.507 0.065

Against Laya Vision 201M on the same questions (its published per-question predictions; difference in points, 95% interval from paired resampling by picture):

Set Questions Laya Vision deVision Difference
VQAv2 yes/no 4,887 0.717 0.725 +0.8 [-0.9, +2.3]
A-OKVQA 1,138 0.598 0.626 +2.7 [-0.6, +6.0]
ScienceQA with images 2,097 0.824 0.766 -5.8 [-7.9, -3.6]
ScienceQA natural science, needs the picture 323 0.700 0.455 -24.5 [-31.6, -17.6]

On the full POPE random / popular / adversarial sets (3,000 questions each) deVision scores 0.891 / 0.868 / 0.791, against Laya Vision's published 0.836 / 0.819 / 0.777 (aggregate scores only, not paired).

CPU latency: P50 148 ms, P95 161 ms (Darwin arm64, 4 threads, FP32, one question per request, warm-up excluded).

Model

Vision Frozen SigLIP2-B/16, 256 × 256 letterboxed input, 64 visual tokens after a 2 × 2 merge and an MLP projector
Decision Laya-initialised ModernBERT-large and Laya's decision head; each option is scored at its [MASK], then a temperature-scaled softmax
Size About 518M parameters, FP32, 2.07 GB (LoRA merged)

Training: the projector is first aligned on COCO captions, then the decision is trained with Laya's RLCD objective on yes/no and multiple-choice questions (projector, decision head and a LoRA on ModernBERT), round after round. This release continued for one pass over 353,830 questions, 315,342 of them from thirteen public datasets (Objects365, TallyQA, CLEVR, CLEVR-Math, Super-CLEVR, FigureQA, MapQA, IconQA, SNLI-VE, Vision-Flan, VisOnlyQA, SpatialSense, PixMo-Count) and the rest replaying earlier data (VQAv2, GQA, A-OKVQA, ScienceQA, VSR, Visual7W, COCO-derived questions). Temperatures are fitted per question type and option count. Details: the training log.

Limitations

  • Comparing two places named in the question (two magnet poles, two chart series, two map regions) stays near chance; on the ScienceQA questions that need the picture it is about 24 points behind Laya Vision.
  • Relative-position yes/no questions are weak: right on both a picture and its mirror for only 19% of pairs (65% for relative-position choice questions).
  • Presence leans towards "yes": it says an absent object is there more often than the previous release (POPE adversarial 0.791).
  • English only, one image, noul and choice only; 256 × 256 input, so no small text.
  • Calibration was fitted on this project's data and may not hold on yours.

Licence

Weights and code: Apache-2.0, like the three base models (Laya, SigLIP2, ModernBERT). The training datasets (listed in this card's metadata) have their own terms, some non-commercial (e.g. ScienceQA, CC BY-NC-SA 4.0); check them against your use.

Links

Code and training records (this release: tag model-v0.2) · evaluation details · provenance · Laya Vision

Downloads last month
27
Safetensors
Model size
0.5B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lukbit/devision

Finetuned
(152)
this model

Datasets used to train lukbit/devision

Evaluation results