# deVision v0.2 evaluation Generated by `scripts/release/hf.py` from the evaluation of this release. | Set | Questions | Accuracy | Mismatched | ECE | |---|---:|---:|---:|---:| | COCO object presence | 1,120 | 0.938 | 0.493 | 0.020 | | VQAv2 multiple choice | 1,420 | 0.898 | 0.399 | 0.026 | | COCO size | 1,100 | 0.860 | 0.504 | 0.036 | | POPE, project filtered set | 8,676 | 0.851 | 0.527 | 0.058 | | GQA val subset | 992 | 0.784 | 0.532 | 0.042 | | COCO position | 1,274 | 0.872 | 0.493 | 0.025 | | VQAv2 yes/no subset | 1,000 | 0.710 | 0.518 | 0.022 | | COCO relative position | 1,950 | 0.711 | 0.484 | 0.047 | | VSR, project held-out split | 904 | 0.679 | 0.481 | 0.048 | | Visual7W, project held-out split | 1,000 | 0.721 | 0.422 | 0.030 | | Fresh counting test (unseen pictures) | 600 | 0.740 | 0.507 | 0.065 | ## Laya Vision 201M, same questions | Set | Questions | Laya Vision | deVision | Difference | |---|---:|---:|---:|---| | VQAv2 yes/no | 4,887 | 0.717 | 0.725 | +0.8 [-0.9, +2.3] | | A-OKVQA | 1,138 | 0.598 | 0.626 | +2.7 [-0.6, +6.0] | | ScienceQA with images | 2,097 | 0.824 | 0.766 | -5.8 [-7.9, -3.6] | | ScienceQA natural science, needs the picture | 323 | 0.700 | 0.455 | -24.5 [-31.6, -17.6] | On the full POPE random / popular / adversarial sets (3,000 questions each) deVision scores **0.891 / 0.868 / 0.791**, against Laya Vision's published **0.836 / 0.819 / 0.777** (aggregate scores only, not paired). ## Temperatures | Bucket | Temperature | |---|---:| | Choice, 2 options | 1.8836 | | Choice, 3–5 options | 1.9790 | | Noul (yes / no) | 1.6382 | | Choice, any other count | 1.9568 |