|
Download README.md from lukbit/devision: direct link, hf CLI and curl.
- Browser
- Download file 9.51 kB
-
https://huggingface.co/lukbit/devision/resolve/main/README.md
- Command line
-
hf download hf://lukbit/devision/README.md
-
curl -L -o README.md https://huggingface.co/lukbit/devision/resolve/main/README.md
9.51 kB
| language: | |
| - en | |
| license: apache-2.0 | |
| library_name: devision | |
| pipeline_tag: visual-question-answering | |
| base_model: | |
| - convaiinnovations/laya | |
| - google/siglip2-base-patch16-256 | |
| datasets: | |
| - HuggingFaceM4/VQAv2 | |
| - lmms-lab/GQA | |
| - HuggingFaceM4/A-OKVQA | |
| - derek-thomas/ScienceQA | |
| - HuggingFaceM4/the_cauldron | |
| - cambridgeltl/vsr_random | |
| - jxu124/objects365 | |
| - J1mb0o/e-SNLI-VE | |
| - Vision-Flan/vision-flan_191-task_1k | |
| - ryokamoi/VisOnlyQA_Train | |
| - RyanWW/Super-CLEVR | |
| - allenai/pixmo-count | |
| tags: | |
| - devision | |
| - visual-decisions | |
| - calibrated-probabilities | |
| - jev | |
| - system-one | |
| - rlcd | |
| - cpu | |
| model-index: | |
| - name: deVision v0.2 | |
| results: | |
| - task: | |
| type: visual-question-answering | |
| dataset: | |
| name: COCO object presence | |
| type: test_exist | |
| metrics: | |
| - type: accuracy | |
| value: 0.9384 | |
| - type: ece | |
| value: 0.0204 | |
| - task: | |
| type: visual-question-answering | |
| dataset: | |
| name: VQAv2 multiple choice | |
| type: test_vqa_choice | |
| metrics: | |
| - type: accuracy | |
| value: 0.8979 | |
| - type: ece | |
| value: 0.0262 | |
| - task: | |
| type: visual-question-answering | |
| dataset: | |
| name: COCO size | |
| type: test_size | |
| metrics: | |
| - type: accuracy | |
| value: 0.86 | |
| - type: ece | |
| value: 0.036 | |
| - task: | |
| type: visual-question-answering | |
| dataset: | |
| name: POPE, project filtered set | |
| type: bench_pope | |
| metrics: | |
| - type: accuracy | |
| value: 0.8511 | |
| - type: ece | |
| value: 0.0578 | |
| - task: | |
| type: visual-question-answering | |
| dataset: | |
| name: GQA val subset | |
| type: test_gqa | |
| metrics: | |
| - type: accuracy | |
| value: 0.7843 | |
| - type: ece | |
| value: 0.0419 | |
| - task: | |
| type: visual-question-answering | |
| dataset: | |
| name: COCO position | |
| type: test_position | |
| metrics: | |
| - type: accuracy | |
| value: 0.8721 | |
| - type: ece | |
| value: 0.0251 | |
| - task: | |
| type: visual-question-answering | |
| dataset: | |
| name: VQAv2 yes/no subset | |
| type: test_vqa_yesno | |
| metrics: | |
| - type: accuracy | |
| value: 0.71 | |
| - type: ece | |
| value: 0.0223 | |
| - task: | |
| type: visual-question-answering | |
| dataset: | |
| name: COCO relative position | |
| type: test_relation | |
| metrics: | |
| - type: accuracy | |
| value: 0.7108 | |
| - type: ece | |
| value: 0.0471 | |
| - task: | |
| type: visual-question-answering | |
| dataset: | |
| name: VSR, project held-out split | |
| type: test_vsr | |
| metrics: | |
| - type: accuracy | |
| value: 0.6792 | |
| - type: ece | |
| value: 0.0478 | |
| - task: | |
| type: visual-question-answering | |
| dataset: | |
| name: Visual7W, project held-out split | |
| type: test_v7w | |
| metrics: | |
| - type: accuracy | |
| value: 0.721 | |
| - type: ece | |
| value: 0.0302 | |
| - task: | |
| type: visual-question-answering | |
| dataset: | |
| name: Fresh counting test (unseen pictures) | |
| type: test_count_fresh | |
| metrics: | |
| - type: accuracy | |
| value: 0.74 | |
| - type: ece | |
| value: 0.0648 | |
| # deVision v0.2 | |
| **Image + English questions → structured answers with calibrated probabilities.** deVision pairs a SigLIP2 vision encoder with Laya's ModernBERT decision model. It scores the options of yes/no (`noul`) and multiple-choice (`choice`) questions instead of generating text, follows the Jev answer format, and runs without a GPU. Several questions about one image share a single image encoding. | |
| ## Usage | |
| ```bash | |
| pip install devision | |
| ``` | |
| ```python | |
| import devision | |
| model = devision.load("lukbit/devision") # latest release; revision="v0.2" pins this one; device defaults to auto | |
| result = model.predict( | |
| image="photo.jpg", # a path, an http(s) URL, a data URI, base64, bytes or a PIL image | |
| state="", # text context; "" when there is none | |
| questions={ | |
| "has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"}, | |
| "room": {"type": "choice", "instructions": "Which room is this?", | |
| "criteria": {"kitchen": None, "bathroom": None, "bedroom": None}}, | |
| }, | |
| ) | |
| print(result["answers"]["has_fork"]["noul"]) # probability of yes | |
| print(result["answers"]["room"]["probabilities"]) # one probability per option, summing to 1 | |
| ``` | |
| HTTP server with a browser demo at `http://127.0.0.1:8000`: | |
| ```bash | |
| devision-serve --checkpoint lukbit/devision --port 8000 | |
| curl -s http://127.0.0.1:8000/v1/systemone -H 'Content-Type: application/json' \ | |
| -d '{"image": "https://example.com/photo.jpg", "state": "", "questions": {"has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"}}}' | |
| ``` | |
| Over HTTP, `image` is an http(s) URL, a data URI or base64; the server never reads its own files. Responses follow Jev; `confidence` for `choice` is `(n * p_max - 1) / (n - 1)`. Invalid requests return 422 (`devision.InvalidRequest` in Python). | |
| **Good for:** whether something is present and how many (ask counts as a choice), what kind of thing or scene, colours and materials, up/down and left/right in photos (as a choice). **Not for:** comparing two places on a diagram, chart or map; relative-position yes/no questions; small text; other languages or several images. Use the probabilities as thresholds and send uncertain cases to a person or a stronger model, after checking the thresholds on your own data. | |
| ## Evaluation | |
| Accuracy through `decide` with the fitted temperatures. *Mismatched*: the same questions with every picture swapped for an unrelated one. Test sets use held-out pictures; they are project subsets, not official leaderboard scores. | |
| | Set | Questions | Accuracy | Mismatched | ECE | | |
| |---|---:|---:|---:|---:| | |
| | COCO object presence | 1,120 | 0.938 | 0.493 | 0.020 | | |
| | VQAv2 multiple choice | 1,420 | 0.898 | 0.399 | 0.026 | | |
| | COCO size | 1,100 | 0.860 | 0.504 | 0.036 | | |
| | POPE, project filtered set | 8,676 | 0.851 | 0.527 | 0.058 | | |
| | GQA val subset | 992 | 0.784 | 0.532 | 0.042 | | |
| | COCO position | 1,274 | 0.872 | 0.493 | 0.025 | | |
| | VQAv2 yes/no subset | 1,000 | 0.710 | 0.518 | 0.022 | | |
| | COCO relative position | 1,950 | 0.711 | 0.484 | 0.047 | | |
| | VSR, project held-out split | 904 | 0.679 | 0.481 | 0.048 | | |
| | Visual7W, project held-out split | 1,000 | 0.721 | 0.422 | 0.030 | | |
| | Fresh counting test (unseen pictures) | 600 | 0.740 | 0.507 | 0.065 | | |
| Against Laya Vision 201M on the same questions (its published per-question predictions; difference in points, 95% interval from paired resampling by picture): | |
| | Set | Questions | Laya Vision | deVision | Difference | | |
| |---|---:|---:|---:|---| | |
| | VQAv2 yes/no | 4,887 | 0.717 | 0.725 | +0.8 [-0.9, +2.3] | | |
| | A-OKVQA | 1,138 | 0.598 | 0.626 | +2.7 [-0.6, +6.0] | | |
| | ScienceQA with images | 2,097 | 0.824 | 0.766 | -5.8 [-7.9, -3.6] | | |
| | ScienceQA natural science, needs the picture | 323 | 0.700 | 0.455 | -24.5 [-31.6, -17.6] | | |
| On the full POPE random / popular / adversarial sets (3,000 questions each) deVision scores **0.891 / 0.868 / 0.791**, against Laya Vision's published **0.836 / 0.819 / 0.777** (aggregate scores only, not paired). | |
| CPU latency: P50 148 ms, P95 161 ms (Darwin arm64, 4 threads, FP32, one question per request, warm-up excluded). | |
| ## Model | |
| | | | | |
| |---|---| | |
| | Vision | Frozen SigLIP2-B/16, 256 × 256 letterboxed input, 64 visual tokens after a 2 × 2 merge and an MLP projector | | |
| | Decision | Laya-initialised ModernBERT-large and Laya's decision head; each option is scored at its `[MASK]`, then a temperature-scaled softmax | | |
| | Size | About 518M parameters, FP32, 2.07 GB (LoRA merged) | | |
| Training: the projector is first aligned on COCO captions, then the decision is trained with Laya's RLCD objective on yes/no and multiple-choice questions (projector, decision head and a LoRA on ModernBERT), round after round. This release continued for one pass over 353,830 questions, 315,342 of them from thirteen public datasets (Objects365, TallyQA, CLEVR, CLEVR-Math, Super-CLEVR, FigureQA, MapQA, IconQA, SNLI-VE, Vision-Flan, VisOnlyQA, SpatialSense, PixMo-Count) and the rest replaying earlier data (VQAv2, GQA, A-OKVQA, ScienceQA, VSR, Visual7W, COCO-derived questions). Temperatures are fitted per question type and option count. Details: the [training log](https://github.com/byebyebruce/devision/blob/master/docs/training-log.md). | |
| ## Limitations | |
| - **Comparing two places named in the question** (two magnet poles, two chart series, two map regions) stays near chance; on the ScienceQA questions that need the picture it is about 24 points behind Laya Vision. | |
| - **Relative-position yes/no questions** are weak: right on both a picture and its mirror for only 19% of pairs (65% for relative-position choice questions). | |
| - **Presence leans towards "yes"**: it says an absent object is there more often than the previous release (POPE adversarial 0.791). | |
| - English only, one image, `noul` and `choice` only; 256 × 256 input, so no small text. | |
| - Calibration was fitted on this project's data and may not hold on yours. | |
| ## Licence | |
| Weights and code: Apache-2.0, like the three base models ([Laya](https://huggingface.co/convaiinnovations/laya), [SigLIP2](https://huggingface.co/google/siglip2-base-patch16-256), [ModernBERT](https://huggingface.co/answerdotai/ModernBERT-large)). The training datasets (listed in this card's metadata) have their own terms, some non-commercial (e.g. ScienceQA, CC BY-NC-SA 4.0); check them against your use. | |
| ## Links | |
| [Code and training records](https://github.com/byebyebruce/devision) (this release: tag `model-v0.2`) · [evaluation details](evaluation/results.md) · [provenance](provenance.json) · [Laya Vision](https://github.com/r33drichards/laya-vision) | |