Instructions to use Kailune-AI/pjev-2b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Kailune-AI/pjev-2b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Kailune-AI/pjev-2b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Kailune-AI/pjev-2b") model = AutoModelForCausalLM.from_pretrained("Kailune-AI/pjev-2b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Kailune-AI/pjev-2b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Kailune-AI/pjev-2b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kailune-AI/pjev-2b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Kailune-AI/pjev-2b
- SGLang
How to use Kailune-AI/pjev-2b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Kailune-AI/pjev-2b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kailune-AI/pjev-2b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Kailune-AI/pjev-2b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kailune-AI/pjev-2b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Kailune-AI/pjev-2b with Docker Model Runner:
docker model run hf.co/Kailune-AI/pjev-2b
pjev-2b
A small typed-decision model. Give it evidence, a question and a fixed set of allowed answers; it returns a calibrated probability for every option in one forward pass. It never generates text, so there is nothing to parse and nothing to hallucinate.
The readout is the base model's own output layer. Options are rendered A., B., C., the prompt
ends where the answer letter goes, and the logits at that single position are restricted to the
letters in play and softmaxed. No decision head is added and no new parameters are introduced —
which is why the pretrained model's own sense of its uncertainty survives fine-tuning instead of
being relearned from scratch, and why the merged model is stock LlamaForCausalLM that any runtime
loads without adaptation.
This repository holds the merged standalone model: base weights with the adapter folded in, tokenizer, and calibration file. No separate adapter, no separate base download.
Accuracy per parameter
JevBench's 231 public items, scored with the benchmark's own modules, one H100, in-process, batch 1. Competitor figures are recomputed from the benchmark's published per-item records on the identical items, except imajev-4b, which publishes aggregates only.
| System | Params | Easy | Standard | Hard |
|---|---|---|---|---|
| imajev-4b | 4B | 100.0 | 98.6 | 72.1 |
| djev | — | 100.0 | 98.6 | 67.6 |
| pjev-2b | 2.5B | 100.0 | 81.9 | 61.3 |
| SemIf | 4B | 100.0 | 98.6 | 61.3 |
| jqv | 32B | 100.0 | 95.8 | 61.3 |
| reflex-4B | 4B | 100.0 | 94.4 | 60.4 |
| kev 8B | 8B | 100.0 | 93.1 | 45.0 |
| Laya | 0.4B | 95.8 | 69.4 | 35.1 |
pjev-2b is not the most accurate model here, and two open systems above it are. What it is: the smallest model that holds the 4B line. At 61.3 on the hard tier it matches a 4B and a 32B, beats every 8B-and-under open model we could pair against, and does it at a quarter to a thirteenth of their parameter count.
Paired on the 111 hard items, McNemar's exact test with a 20,000-sample paired bootstrap:
| Opponent | Diff | 95% CI | p | Verdict |
|---|---|---|---|---|
| Laya | +26.1 | +12.6, +38.7 | 0.0003 | win |
| the 8B-and-under open field | +16.2 to +31.5 | — | ≤0.0055 | win |
| the 4B class | ±0.9 | — | ≥0.28 | tie |
These are public-half numbers and are not an official JevBench score. The official board also scores 308 sealed items that only its operator can run, and every ranked system drops substantially from public to sealed accuracy. Treat this as a setup check, not a ranking.
Several questions, one piece of evidence
Real traffic rarely asks one question about a document. Route it, flag its urgency, and check it against a policy, and that is three decisions over the same evidence. The state is encoded once and the per-question suffixes batch: six questions cost 48 ms in total, 8.0 ms each, against 121 ms one at a time. One question is still cheapest on the direct path at 20.5 ms; the crossover is at two.
Quick start
Everything needed to serve it is in this repository — the weights, and the three MIT-licensed files that define the decision prompt and the letter readout, so the served model answers exactly as it was trained.
pip install torch transformers
hf download Kailune-AI/pjev-2b --local-dir pjev-2b
cd pjev-2b
python serve.py --model . --calib ./calib.json --port 8000
That gives you a TypeSafe-compatible endpoint at POST /v1/systemone on the transformers backend,
which is all you need to evaluate it. For throughput, serve.py can front a vLLM engine instead:
pip install "vllm>=0.30.0", start vLLM on another port, and pass --vllm-url.
A real request and the real response it returns — copy-pasted from a live server, not illustrative:
import requests, json
req = {
"state": "Ticket #48120\nFrom: dana.k@example.com\n\n"
"I was charged twice for order #77219 on 3 March. Both charges show as settled "
"on my card. I only placed the order once. Please refund the duplicate.",
"questions": {
"queue": {"type": "choice", "instructions": "Route this ticket to the team that should own it.",
"criteria": {"billing": "Payments, charges, refunds and invoices.",
"shipping": "Delivery, tracking and returns in transit.",
"account": "Login, profile and account settings.",
"other": "Anything that fits none of the above."}},
"urgent": {"type": "noul", "instructions": "This ticket needs a response within 24 hours.",
"criteria": {"true": "Money is at stake or the customer is blocked.",
"false": "It can wait for the normal queue."}},
"frustration": {"type": "score", "instructions": "How frustrated does the customer sound?",
"criteria": ["calm", "mildly annoyed", "clearly upset", "angry"]},
},
}
print(json.dumps(requests.post("http://127.0.0.1:8000/v1/systemone", json=req).json(), indent=2))
{
"answers": {
"queue": {
"type": "choice", "choice": "billing", "confidence": 0.938,
"probabilities": {"billing": 0.938, "shipping": 0.028, "account": 0.012, "other": 0.022}
},
"urgent": {"type": "noul", "noul": 0.881, "value": true, "confidence": 0.881},
"frustration": {
"type": "score", "score": 2, "expected": 2.004, "confidence": 0.665,
"probabilities": {"0": 0.043, "1": 0.102, "2": 0.665, "3": 0.191}
}
},
"usage": {"input_tokens": 543, "output_tokens": 0},
"latency_s": 0.062
}
Three typed decisions over one ticket, 62 ms, zero output tokens. Route to billing, threshold the
urgent probability against your own cost of being wrong, and note that frustration returns a
distribution over the ordered levels rather than a single guess.
Give your options descriptions. The criteria strings are part of the prompt and the model
reads them. The same ticket with bare null criteria routes to other at 0.48 instead of billing
at 0.94 — the labels alone do not tell it what billing means.
What to expect
Accuracy. Strong on easy and standard items, mid-field on hard ones. If your decisions look like routing, classification, policy checks against a short document, or yes/no judgements with clear criteria, this is the right size. If they need multi-hop reasoning over long policies, the hard-tier number above is the honest guide.
Confidence. Calibration is the axis where this model does not lead. Hard-tier expected calibration error is 0.120; several competing systems are better there, and the ones that are fit a temperature on a held-out set. We ship at temperature 1.0 — see below — which is a defensible default and also a measurable cost on the hardest items.
Speed. 20.5 ms per decision on one H100, one question at a time; 8.0 ms per question when several share one state.
Options. Up to 26 options are read directly as letters. Beyond that a tournament round is used, which is an approximation rather than a single joint softmax.
Abstention. There is none. If the evidence supports no option, the model still distributes probability across the options you supplied. Threshold on confidence and route the low end to a person.
How it works
- The state, the question and the option list are rendered into a prompt that ends immediately before the answer letter.
- One forward pass. At that final position the logits of the tokens
A,B,C, … are gathered — each is a single token in this tokenizer, asserted at load — and everything else is discarded. - Softmax over the letters in play gives the distribution. For
noulquestions the two letters are the yes and no branches; forscorequestions they are the ordered levels. - Nothing is decoded. A request with several questions about one state runs the shared prefix once and batches the per-question suffixes.
Why temperature 1.0. Our held-out calibration folds turned out easier than genuinely hard items, so every temperature fitted on them came out below 1 and made the model more confident exactly where it should have been less. Rather than ship a temperature that flatters the easy cases, we ship uncalibrated and say so. If you have a few hundred labelled examples of your own traffic, fitting one temperature on them is the first improvement to try, and on our own hard-tier numbers it is worth real points.
The merged weights here were checked against the unmerged base-plus-adapter on 40 JevBench easy and standard items: identical answers on 40 of 40, maximum probability difference 0.0000.
Limitations
- Mid-field on hard reasoning. Two open systems in the table above beat it on the hard tier. It is a 2B; this is the trade.
- Not the best calibrated. Hard-tier ECE 0.120, shipped uncalibrated. Fit your own temperature.
- No abstention output. Nothing detects "the evidence does not answer this".
- 26 options before a tournament round, which is an approximation.
- English in practice. Evaluated in English only.
- Public-benchmark numbers only. We have no measurement on any sealed evaluation, and every ranked system on the public board drops substantially when one is run.
- Not a safety, medical, legal or hiring certificate. It returns probabilities over options you supply.
Licence and credits
Model weights: Apache-2.0, matching the openbmb/MiniCPM5-2B base.
The three serving files (serve.py, decision_core.py, letter_adapter.py) are redistributed
unmodified from the Eikos project under the MIT Licence,
Copyright (c) 2026 Caio Vicentino — see LICENSE-eikos and NOTICE.
Fine-tuned on the public
caiovicentino1/eikos-decisions
dataset (CC BY 4.0, with an ODC-BY-1.0 subset), used as released. That corpus and the recipe it
documents are its author's work; please carry the attribution if you build on this.
- Downloads last month
- 356
Model tree for Kailune-AI/pjev-2b
Base model
openbmb/MiniCPM5-2B
