Post
37
Imajev-4b is #1 of 91 on JevBench and #3 of 56 on DecisionBench š
Some context first. I'm a process improvement / business consultant and have worked with Fortune 500 companies on their processes around refunds, returns and customer support. In every process map, the decision nodes were handled by a person, because putting ambiguity into code is very hard.
When Jev came out, I could clearly see it fitting those decision nodes. But Jev only reads text, and many of these decisions start with a photo. So I set out to build the same idea for text and images in a single open model, and that became imajev.
Training went badly at first. My first big fine-tune on about 500k short decisions made the 9B worse at reasoning (64.9 down to 42.3 on JevBench hard). I spent the next couple of weeks generating hard questions with open-weight models and keeping only the ones where two models agreed on the answer. That brought it back.
Results this week, both run by the benchmarks' own maintainers:
š„ JevBench v1.4.2.2 (scored 27 Sep): #1 of 91, 67.37 vs Jev 1.13.0 at 63.29
š DecisionBench (eng, v1): #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1 Flash. The two above it are the benchmark team's own models.
ā Zero invalid answers: 23,900 of 23,900 on DecisionBench and 308 of 308 on JevBench's sealed set
Its strength is that its confidence can be trusted, and it says "can't tell" instead of guessing.
What it is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions. It returns a probability for every allowed answer plus "unknown", in one forward pass, so it can't produce a malformed answer.
š§ Weights: mohit67890/imajev-4b
š» Code: https://github.com/mohit67890/imajev
š JevBench: https://benchmarkheaven.com/jev-models/v1.4.2.2
š DecisionBench: Hanno-Labs/decision-bench-leaderboard
Thanks to the Qwen Team
Qwen for the base model
Some context first. I'm a process improvement / business consultant and have worked with Fortune 500 companies on their processes around refunds, returns and customer support. In every process map, the decision nodes were handled by a person, because putting ambiguity into code is very hard.
When Jev came out, I could clearly see it fitting those decision nodes. But Jev only reads text, and many of these decisions start with a photo. So I set out to build the same idea for text and images in a single open model, and that became imajev.
Training went badly at first. My first big fine-tune on about 500k short decisions made the 9B worse at reasoning (64.9 down to 42.3 on JevBench hard). I spent the next couple of weeks generating hard questions with open-weight models and keeping only the ones where two models agreed on the answer. That brought it back.
Results this week, both run by the benchmarks' own maintainers:
š„ JevBench v1.4.2.2 (scored 27 Sep): #1 of 91, 67.37 vs Jev 1.13.0 at 63.29
š DecisionBench (eng, v1): #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1 Flash. The two above it are the benchmark team's own models.
ā Zero invalid answers: 23,900 of 23,900 on DecisionBench and 308 of 308 on JevBench's sealed set
Its strength is that its confidence can be trusted, and it says "can't tell" instead of guessing.
What it is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions. It returns a probability for every allowed answer plus "unknown", in one forward pass, so it can't produce a malformed answer.
š§ Weights: mohit67890/imajev-4b
š» Code: https://github.com/mohit67890/imajev
š JevBench: https://benchmarkheaven.com/jev-models/v1.4.2.2
š DecisionBench: Hanno-Labs/decision-bench-leaderboard
Thanks to the Qwen Team