OneJev-4B-GGUF

OneJev-4B is the 4-billion-parameter model of OmniJev's OneJev family, a multimodal "System One" decision model built as a full fine-tune of Qwen3.5-4B on 99,193 questions drawn from real agent runs, videos, and images. It accepts a screenshot, photo, video, or plain text along with a few typed questions, such as Noul yes/no or Choice over labeled options, and returns a calibrated probability for every option in a single forward pass, with no text generation. It is served with the qev package, which speaks TypeSafe's System One API and adds a media field for images and video. On one H200 with a 1280x720 screenshot, it answers a single question in 64 ms, or 10 questions in one request in 104 ms (10.4 ms per question). Its results are reported against Jev 1.13, Jev-Omni 12B, and Qwen3.8-27B in thinking mode, on a held-out OneJev test set. The card presents these only as a graphic, so the individual scores are not readable from the text, and it notes that Jev 1.13's figures are its published ones and that it reads text only. OneJev-4B sits between OneJev-0.8B (2.2 GB) and OneJev-9B (18.8 GB), alongside OneJev-27B and an FP8 build of it, and it is released under Apache 2.0 (10.4 GB of weights).

Model Files

File Name Quant Type File Size File Link Description
OneJev-4B.BF16.gguf BF16 9.7 GB Link Full BF16 weights. Highest quality, largest file size.
OneJev-4B.Q3_K_L.gguf Q3_K_L 2.69 GB Link Lower quality but usable, good for low RAM availability.
OneJev-4B.Q3_K_M.gguf Q3_K_M 2.54 GB Link Low quality.
OneJev-4B.Q4_K_M.gguf Q4_K_M 3.07 GB Link Good quality, default size for most use cases, recommended.
OneJev-4B.Q4_K_S.gguf Q4_K_S 2.92 GB Link Slightly lower quality with more space savings, recommended.
OneJev-4B.Q5_K_M.gguf Q5_K_M 3.51 GB Link High quality, recommended.
OneJev-4B.Q5_K_S.gguf Q5_K_S 3.43 GB Link High quality, recommended.
OneJev-4B.Q6_K.gguf Q6_K 3.99 GB Link Very high quality, near perfect, recommended.
OneJev-4B.Q8_0.gguf Q8_0 5.16 GB Link Extremely high quality, generally unneeded but max available quant.
OneJev-4B.mmproj-bf16.gguf mmproj-bf16 676 MB Link Multimodal projection file in BF16 format. Used for vision/language models.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
-
GGUF
Model size
5B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/OneJev-4B-GGUF

Finetuned
Qwen/Qwen3.5-4B
Quantized
(2)
this model

Collections including prithivMLmods/OneJev-4B-GGUF