jev-vision-27b

A LoRA on Qwen/Qwen3.8-27B for typed decisions over text and images: give it a state (text, images, or both), a question and a list of options, and it returns a probability for every option in one forward pass.

The other strong entrants on the Decision Index vision board were trained on text only and answer image questions zero-shot. This one was trained on vision decisions directly (spatial and depth questions, charts, documents and receipts, web screenshots, caption word-order pairs, memes, science diagrams), alongside a text decision corpus.

How it answers

Options are labelled A, B, C, ... in the prompt. The answer is read from the label logits at the answer position and softmaxed over the listed options only. One forward pass per question, no sampling, no thinking.

  • Up to 62 options per question (A-Z, a-z, 0-9, all single tokens). Larger questions are refused as unsupported, never truncated.
  • Context up to 65,536 tokens on the served path.
  • Images: up to 8 per request, max_pixels 4,000,000.

Run it

pip install vllm fastapi uvicorn peft transformers git+https://github.com/apolinario/decision-index
cd code
python -m train.merge --ckpt .. --out ../merged      # adapter -> plain bf16 model (one-time)
python -m serve.server --model ../merged --perms 1 --port 8000

Then POST /v1/systemone with {"model", "state", "questions"}, the Decision Index wire format. state may be a string, a JSON object, or a list mixing text and {"image": "https://... | data:..."} parts. The code is in code/.

The same model also runs in-process as a kit engine: --engine serve.engine:JVEngine --option adapter=<path>.

Speed

Path median mean p80
Text decisions, 1 client, A100 80GB, vLLM 120 ms 590 ms 614 ms
Image decisions, RTX PRO 6000, vLLM 254 ms 399 ms 635 ms

The image row was measured with two option orders averaged; the shipped setting is one order. The full public 0.3 text suite (140,620 requests) took about 4 hours on one H100 with 16 concurrent clients.

Training

  • Base: Qwen/Qwen3.8-27B, LoRA r=32, alpha=64, on the language model's attention, linear-attention and MLP projections. The vision tower is frozen.
  • Loss: KL to soft or label-smoothed targets at the answer position.
  • Text: 18k rows of SargeDev/jev-distill-corpus-v3.
  • Vision: about 13k rows built from train splits only (COCO boxes, ChartQA, DocumentVQA, CORD, Multimodal-Mind2Web train, Hateful Memes train with balanced classes, ScienceQA, COCO caption word swaps, SAT). Every training image was pHash-checked against the evaluation images, and near-duplicates were dropped.
  • 2,000 steps, then 400 more at a lower learning rate with the balanced Hateful Memes rows and without SAT.

Limits

Questions with more than 62 options are refused, which affects CLINC150, BANKING77 and POP909 in the public suite. Benchmark numbers on this card come from our own replica of the public vision benchmarks, not the official board.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for e12ex2/jev-vision-27b

Base model

Qwen/Qwen3.8-27B
Adapter
(188)
this model