Feature Request: JEV-4B-VL — 12GB VRAM Multimodal Decision Model

#2
by Exo87 - opened

JEV-27B-VL is impressive but needs ~80GB VRAM, making it inaccessible to most users. A ~4B variant could run/train on a single 12GB GPU while leaving 2–4GB free for a game or other apps.

Proposal:
Release JEV-4B-VL, distilled from JEV-27B-VL, preserving the calibrated decision head and System 1/System 2 modes. Suggested backbone: Qwen3-VL-4B or Gemma 3 4B. Use 4-bit QLoRA + Unsloth for low-VRAM training/inference.

Target:

· 4B params
· 6–10GB VRAM with 4-bit/LoRA
· Trainable on RTX 3060/4070 12GB
· Apache-2.0, Unsloth-compatible scripts, quantized weights

Why feasible:
4B VL models already fit in ~8–10GB with QLoRA; 2B VL QLoRA runs in ~6.7GB. A 4B JEV variant is realistic and would democratize low-latency multimodal decisions on consumer hardware.

AutoTrust AI Lab org

Hi Marwan (@Exo87 ) ,

Thanks for taking the time to write such a detailed, well-reasoned proposal. This kind of feedback is genuinely helpful to us.

We hear you on hardware accessibility. A quantised version of JEV-27B-VL is already in our plan, which will significantly reduce the memory needed to run it.

The idea of a distilled JEV-4B-VL for 12GB-class GPUs is a great suggestion. We've noted it, along with your points on backbone options and the QLoRA/Unsloth workflow, and will take it into account as we plan future releases.

We'll share updates as plans firm up. Thanks again, and please keep the feedback coming and you can follow us on X @AutoTrustAI for more updates.

— The AutoTrust AI team

AutoTrust AI Lab org

@Exo87 If the need is a multimodal decision model you can run on consumer hardware now, rather than a JEV distill specifically: I built Seb-9B, an Apache-2.0 model that takes a text, JSON or image state plus one question with fixed options and returns a probability for every option in one forward pass: https://huggingface.co/ironbcc/seb-9b

It's 9B rather than 27B. The GGUF build (https://huggingface.co/ironbcc/seb-9b-GGUF) is 9.8 GB at Q8_0 plus a 0.9 GB vision projector, so 8-bit is too tight for your 12 GB target with headroom. A 4-bit quant would fit, but I've only validated the 8-bit builds and haven't measured accuracy below that.

Differences from what you described: it isn't distilled from JEV-27B-VL, it has no System 2 mode, choice questions take at most 20 options, and I haven't run it head-to-head against JEV-27B-VL. The card compares it with TypeSafe Jev 1.13 on identical rows and reports image results on a held-out set.

а в доту будет шпилить или это сложно для нее?

Sign up or log in to comment