devision / README.md
lukbit's picture
deVision v0.2 card (github 5d2127c99773055ee7266d525d82b1b6b9be5e80)
b971e2b verified
|
Raw History Blame Contribute Delete
9.51 kB
---
language:
- en
license: apache-2.0
library_name: devision
pipeline_tag: visual-question-answering
base_model:
- convaiinnovations/laya
- google/siglip2-base-patch16-256
datasets:
- HuggingFaceM4/VQAv2
- lmms-lab/GQA
- HuggingFaceM4/A-OKVQA
- derek-thomas/ScienceQA
- HuggingFaceM4/the_cauldron
- cambridgeltl/vsr_random
- jxu124/objects365
- J1mb0o/e-SNLI-VE
- Vision-Flan/vision-flan_191-task_1k
- ryokamoi/VisOnlyQA_Train
- RyanWW/Super-CLEVR
- allenai/pixmo-count
tags:
- devision
- visual-decisions
- calibrated-probabilities
- jev
- system-one
- rlcd
- cpu
model-index:
- name: deVision v0.2
results:
- task:
type: visual-question-answering
dataset:
name: COCO object presence
type: test_exist
metrics:
- type: accuracy
value: 0.9384
- type: ece
value: 0.0204
- task:
type: visual-question-answering
dataset:
name: VQAv2 multiple choice
type: test_vqa_choice
metrics:
- type: accuracy
value: 0.8979
- type: ece
value: 0.0262
- task:
type: visual-question-answering
dataset:
name: COCO size
type: test_size
metrics:
- type: accuracy
value: 0.86
- type: ece
value: 0.036
- task:
type: visual-question-answering
dataset:
name: POPE, project filtered set
type: bench_pope
metrics:
- type: accuracy
value: 0.8511
- type: ece
value: 0.0578
- task:
type: visual-question-answering
dataset:
name: GQA val subset
type: test_gqa
metrics:
- type: accuracy
value: 0.7843
- type: ece
value: 0.0419
- task:
type: visual-question-answering
dataset:
name: COCO position
type: test_position
metrics:
- type: accuracy
value: 0.8721
- type: ece
value: 0.0251
- task:
type: visual-question-answering
dataset:
name: VQAv2 yes/no subset
type: test_vqa_yesno
metrics:
- type: accuracy
value: 0.71
- type: ece
value: 0.0223
- task:
type: visual-question-answering
dataset:
name: COCO relative position
type: test_relation
metrics:
- type: accuracy
value: 0.7108
- type: ece
value: 0.0471
- task:
type: visual-question-answering
dataset:
name: VSR, project held-out split
type: test_vsr
metrics:
- type: accuracy
value: 0.6792
- type: ece
value: 0.0478
- task:
type: visual-question-answering
dataset:
name: Visual7W, project held-out split
type: test_v7w
metrics:
- type: accuracy
value: 0.721
- type: ece
value: 0.0302
- task:
type: visual-question-answering
dataset:
name: Fresh counting test (unseen pictures)
type: test_count_fresh
metrics:
- type: accuracy
value: 0.74
- type: ece
value: 0.0648
---
# deVision v0.2
**Image + English questions → structured answers with calibrated probabilities.** deVision pairs a SigLIP2 vision encoder with Laya's ModernBERT decision model. It scores the options of yes/no (`noul`) and multiple-choice (`choice`) questions instead of generating text, follows the Jev answer format, and runs without a GPU. Several questions about one image share a single image encoding.
## Usage
```bash
pip install devision
```
```python
import devision
model = devision.load("lukbit/devision") # latest release; revision="v0.2" pins this one; device defaults to auto
result = model.predict(
image="photo.jpg", # a path, an http(s) URL, a data URI, base64, bytes or a PIL image
state="", # text context; "" when there is none
questions={
"has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"},
"room": {"type": "choice", "instructions": "Which room is this?",
"criteria": {"kitchen": None, "bathroom": None, "bedroom": None}},
},
)
print(result["answers"]["has_fork"]["noul"]) # probability of yes
print(result["answers"]["room"]["probabilities"]) # one probability per option, summing to 1
```
HTTP server with a browser demo at `http://127.0.0.1:8000`:
```bash
devision-serve --checkpoint lukbit/devision --port 8000
curl -s http://127.0.0.1:8000/v1/systemone -H 'Content-Type: application/json' \
-d '{"image": "https://example.com/photo.jpg", "state": "", "questions": {"has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"}}}'
```
Over HTTP, `image` is an http(s) URL, a data URI or base64; the server never reads its own files. Responses follow Jev; `confidence` for `choice` is `(n * p_max - 1) / (n - 1)`. Invalid requests return 422 (`devision.InvalidRequest` in Python).
**Good for:** whether something is present and how many (ask counts as a choice), what kind of thing or scene, colours and materials, up/down and left/right in photos (as a choice). **Not for:** comparing two places on a diagram, chart or map; relative-position yes/no questions; small text; other languages or several images. Use the probabilities as thresholds and send uncertain cases to a person or a stronger model, after checking the thresholds on your own data.
## Evaluation
Accuracy through `decide` with the fitted temperatures. *Mismatched*: the same questions with every picture swapped for an unrelated one. Test sets use held-out pictures; they are project subsets, not official leaderboard scores.
| Set | Questions | Accuracy | Mismatched | ECE |
|---|---:|---:|---:|---:|
| COCO object presence | 1,120 | 0.938 | 0.493 | 0.020 |
| VQAv2 multiple choice | 1,420 | 0.898 | 0.399 | 0.026 |
| COCO size | 1,100 | 0.860 | 0.504 | 0.036 |
| POPE, project filtered set | 8,676 | 0.851 | 0.527 | 0.058 |
| GQA val subset | 992 | 0.784 | 0.532 | 0.042 |
| COCO position | 1,274 | 0.872 | 0.493 | 0.025 |
| VQAv2 yes/no subset | 1,000 | 0.710 | 0.518 | 0.022 |
| COCO relative position | 1,950 | 0.711 | 0.484 | 0.047 |
| VSR, project held-out split | 904 | 0.679 | 0.481 | 0.048 |
| Visual7W, project held-out split | 1,000 | 0.721 | 0.422 | 0.030 |
| Fresh counting test (unseen pictures) | 600 | 0.740 | 0.507 | 0.065 |
Against Laya Vision 201M on the same questions (its published per-question predictions; difference in points, 95% interval from paired resampling by picture):
| Set | Questions | Laya Vision | deVision | Difference |
|---|---:|---:|---:|---|
| VQAv2 yes/no | 4,887 | 0.717 | 0.725 | +0.8 [-0.9, +2.3] |
| A-OKVQA | 1,138 | 0.598 | 0.626 | +2.7 [-0.6, +6.0] |
| ScienceQA with images | 2,097 | 0.824 | 0.766 | -5.8 [-7.9, -3.6] |
| ScienceQA natural science, needs the picture | 323 | 0.700 | 0.455 | -24.5 [-31.6, -17.6] |
On the full POPE random / popular / adversarial sets (3,000 questions each) deVision scores **0.891 / 0.868 / 0.791**, against Laya Vision's published **0.836 / 0.819 / 0.777** (aggregate scores only, not paired).
CPU latency: P50 148 ms, P95 161 ms (Darwin arm64, 4 threads, FP32, one question per request, warm-up excluded).
## Model
| | |
|---|---|
| Vision | Frozen SigLIP2-B/16, 256 × 256 letterboxed input, 64 visual tokens after a 2 × 2 merge and an MLP projector |
| Decision | Laya-initialised ModernBERT-large and Laya's decision head; each option is scored at its `[MASK]`, then a temperature-scaled softmax |
| Size | About 518M parameters, FP32, 2.07 GB (LoRA merged) |
Training: the projector is first aligned on COCO captions, then the decision is trained with Laya's RLCD objective on yes/no and multiple-choice questions (projector, decision head and a LoRA on ModernBERT), round after round. This release continued for one pass over 353,830 questions, 315,342 of them from thirteen public datasets (Objects365, TallyQA, CLEVR, CLEVR-Math, Super-CLEVR, FigureQA, MapQA, IconQA, SNLI-VE, Vision-Flan, VisOnlyQA, SpatialSense, PixMo-Count) and the rest replaying earlier data (VQAv2, GQA, A-OKVQA, ScienceQA, VSR, Visual7W, COCO-derived questions). Temperatures are fitted per question type and option count. Details: the [training log](https://github.com/byebyebruce/devision/blob/master/docs/training-log.md).
## Limitations
- **Comparing two places named in the question** (two magnet poles, two chart series, two map regions) stays near chance; on the ScienceQA questions that need the picture it is about 24 points behind Laya Vision.
- **Relative-position yes/no questions** are weak: right on both a picture and its mirror for only 19% of pairs (65% for relative-position choice questions).
- **Presence leans towards "yes"**: it says an absent object is there more often than the previous release (POPE adversarial 0.791).
- English only, one image, `noul` and `choice` only; 256 × 256 input, so no small text.
- Calibration was fitted on this project's data and may not hold on yours.
## Licence
Weights and code: Apache-2.0, like the three base models ([Laya](https://huggingface.co/convaiinnovations/laya), [SigLIP2](https://huggingface.co/google/siglip2-base-patch16-256), [ModernBERT](https://huggingface.co/answerdotai/ModernBERT-large)). The training datasets (listed in this card's metadata) have their own terms, some non-commercial (e.g. ScienceQA, CC BY-NC-SA 4.0); check them against your use.
## Links
[Code and training records](https://github.com/byebyebruce/devision) (this release: tag `model-v0.2`) · [evaluation details](evaluation/results.md) · [provenance](provenance.json) · [Laya Vision](https://github.com/r33drichards/laya-vision)