OpenJev-Flash-9B-GGUF

OpenJev Flash 9B is the small, fast member of the OpenJev family: an open-weights decision model built on Qwen3.5-9B and fine-tuned with supervised and teacher-guided fine-tuning, though the training recipe and data are not released. It answers typed questions about text (choice with up to 52 options per pass, yes/no noul, or ordered score) with a calibrated probability for every option. The labels are written in the request, so a new task or domain needs a JSON change, not retraining. It reads the scores of the option letters at the first output position, so there is no free-form text to parse. Each decision takes about 40 ms on one H100 (FP8), 1.6x faster than OpenJev 27B, at roughly 4.6 US cents per 1,000 decisions. On JevBench's 231 public items it gets 188 correct (81.4%), on par with Cloudflare's Clef-Flash (82.3%) and ahead of Kev-9B and Nimble 9B (both 79.2%), and it does best on ambiguous cases and on judging responses. No JevBench item was used in training or tuning. On the card's own 10,000-question mix from 34 public sources it scores 79.4%, behind OpenJev 27B (84.1%) and the hosted Jev API (85.4%). It is served with vLLM plus a small /v1/systemone helper shim that mirrors the hosted Jev API and the 27B's interface, with FP8, MLX 4-bit and 8-bit, and GGUF builds available. The weights are CC BY-NC 4.0, so non-commercial use only unless you license it commercially, and the helper and serve files are Apache 2.0. The project is independent of TypeSafe. openjev/OpenJev-Flash-9B on Hugging Face — OpenJev-Flash-9B.

Model Files

File Name Quant Type File Size File Link Description
OpenJev-Flash-9B.BF16.gguf BF16 17.9 GB Link Full BF16 weights. Highest quality, largest file size.
OpenJev-Flash-9B.Q3_K_L.gguf Q3_K_L 4.93 GB Link Lower quality but usable, good for low RAM availability.
OpenJev-Flash-9B.Q3_K_M.gguf Q3_K_M 4.62 GB Link Low quality.
OpenJev-Flash-9B.Q4_K_M.gguf Q4_K_M 5.63 GB Link Good quality, default size for most use cases, recommended.
OpenJev-Flash-9B.Q4_K_S.gguf Q4_K_S 5.35 GB Link Slightly lower quality with more space savings, recommended.
OpenJev-Flash-9B.Q5_K_M.gguf Q5_K_M 6.47 GB Link High quality, recommended.
OpenJev-Flash-9B.Q5_K_S.gguf Q5_K_S 6.31 GB Link High quality, recommended.
OpenJev-Flash-9B.Q6_K.gguf Q6_K 7.36 GB Link Very high quality, near perfect, recommended.
OpenJev-Flash-9B.mmproj-bf16.gguf mmproj-bf16 922 MB Link Multimodal projection file in BF16 format. Used for vision/language models.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Releases / v0.6.0 — https://github.com/ggml-org/llama.cpp/releases/tag/v0.6.0

Downloads last month
-
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/OpenJev-Flash-9B-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(5)
this model

Collections including prithivMLmods/OpenJev-Flash-9B-GGUF