qjev+ (JEV decision model)

qjev+ is a 27B JEV decision model: give it a conversation, a document or an agent trajectory plus a set of options, and it returns the right choice, fast and reliably. It is built for the decisions that sit inside real products: routing, intent and category classification, content moderation, and approving what an agent is about to do.

“The best decision model is not the one that says the most — it is the one that makes the right call when it matters.”

Highlights

  • Accurate single-choice decisions across Vietnamese and English, with the full general ability of Qwen3.8-27B.
  • Strong safety out of the box: high accuracy on Vietnamese and English content-safety benchmarks, and the best agent-safety score in its class on ATBench, with a very low false-block rate on benign inputs (NotInject).
  • Agent-ready: judges whether the next tool call in a multi-step agent run should go ahead.
  • JEV-ready: built for JEV decision serving, returning typed decisions (enum, boolean, multi-field JSON) in milliseconds instead of generating text. On a single H200, a decision takes about 47 ms p50.
  • Drop-in: standard Qwen3.8 architecture and chat template; runs on transformers, vLLM and SGLang with structured output.

Results

Public benchmarks. Single-label decisions, temperature 0, one option chosen from the allowed set, default threshold, no per-task tuning.

Benchmark qjev+ Qwen3.8-27B Quyet-1.0-Large pplx-decider-v1-27b Gemma-4-26B-A4B-it
ViCulturaBench, accuracy (4,000, labels re-verified) 91.5 91.5 95.2 90.0 90.7
Aegis 2.0, accuracy (150 sampled from test) 78.7 78.7 81.3 75.3 78.7
ATBench, accuracy (1,000 agent trajectories) 73.6 70.2 58.2 60.8 65.0
ATBench, unsafe recall 62.8 46.5 44.5 22.1 48.5
ATBench, false-block rate (lower is better) 15.7 6.4 28.2 1.0 18.7
NotInject, false-block rate (339 benign, lower is better) 2.1 2.4 6.8 0.9 5.3
Prompt protection, accuracy (385) 97.9 97.9 97.9 97.7 99.0
Latency p50, one H200, JEV (ms) 47 48 33 47 21

All models run through the same JEV serving path (SGLang), same prompts and label sets.

Usage

Ask for JSON with an enum field via structured output, thinking disabled:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="-")
r = client.chat.completions.create(
    model="qjev-plus-27b",
    temperature=0,
    messages=[{"role": "system", "content": "Classify the user's message."},
              {"role": "user", "content": "Hoá đơn của tôi chưa được xuất sau 3 ngày."}],
    response_format={"type": "json_schema", "json_schema": {"name": "decision", "strict": True, "schema": {
        "type": "object",
        "properties": {"intent": {"type": "string", "enum": ["question", "request", "complaint", "other"]}},
        "required": ["intent"], "additionalProperties": False}}},
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
print(r.choices[0].message.content)

License

Apache 2.0, same as the base model.

Downloads last month
2
Safetensors
Model size
28B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for beyoru/qjev-plus-27b

Base model

Qwen/Qwen3.8-27B
Finetuned
(481)
this model
Quantizations
1 model