Instructions to use BricksDisplay/jevling-e2b-v0.1-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use BricksDisplay/jevling-e2b-v0.1-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16 # Run inference directly in the terminal: ./llama-cli -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
Use Docker
docker model run hf.co/BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
- LM Studio
- Jan
- Ollama
How to use BricksDisplay/jevling-e2b-v0.1-GGUF with Ollama:
ollama run hf.co/BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
- Unsloth Desktop
- Docker Model Runner
How to use BricksDisplay/jevling-e2b-v0.1-GGUF with Docker Model Runner:
docker model run hf.co/BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
- Lemonade
How to use BricksDisplay/jevling-e2b-v0.1-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull BricksDisplay/jevling-e2b-v0.1-GGUF:BF16
Run and chat with the model
lemonade run user.jevling-e2b-v0.1-GGUF-BF16
List all available models
lemonade list
- Atomic Chat
Jevling-E2B-v0.1 โ GGUF
Quantised build of BricksDisplay/jevling-e2b-v0.1 for on-device use (16 GB RAM). This is a System One decision model: one forward pass, no generated text, several typed questions per call, each answered as a calibrated probability. It does not work with stock llama.cpp chat/completion endpoints โ they can load the weights but cannot ask a typed question or read the answer slot.
Use the maintained implementation
tools/system-one in mybigday/system-one-llama.cpp, branch feat/system-one:
git clone -b feat/system-one https://github.com/mybigday/system-one-llama.cpp
cd system-one-llama.cpp && cmake -B build && cmake --build build --target llama-system-one llama-system-one-batch -j
One call, three typed questions:
build/bin/llama-system-one -m jevling-e2b-v0.1-q8_0-embf16.gguf \
--state "Customer: I was charged twice for the same subscription this month. Please refund the duplicate charge." \
--choice "queue:Which team should handle this ticket?:billing=payments and refunds,technical=a product fault,account=login or profile settings" \
--noul "refund:Is the customer asking for a refund?" \
--score "urgency:How urgent is this?:routine,soon,urgent,critical" --json
Answers come back as probabilities per option (--json is the same shape the server returns: noul P(true), choice with a distribution, score as the expected level). Pass option descriptions โ the model was trained with them. Use --system-one-request request.json for anything with colons or many questions; llama-system-one-batch for throughput. The prompt template and readout live inside the GGUF (tokenizer.chat_template.system_one; the readout is derived from the model, not declared); do not supply a chat template of your own.
Measured: โ1.7 s per 5-question request on 16 CPU threads, โ2.9 GB RSS; run with --swa-full. Keep flash-attention off on CPU and threads = physical cores.
Or the /v1/systemone endpoint
This file also answers upstream llama.cpp's decision-model endpoint, which needs a build that tokenizes the prompt piece by piece โ the boundaries this model was trained on. That is one commit on top of upstream master:
git clone -b system-one/decision-segments https://github.com/mybigday/system-one-llama.cpp
cd system-one-llama.cpp && cmake -B build -DCMAKE_BUILD_TYPE=Release && cmake --build build --target llama-server -j
build/bin/llama-server -m jevling-e2b-v0.1-q8_0-embf16.gguf -c 2048 -t $(nproc) --n-outputs-max-per-seq 16
curl http://localhost:8080/v1/systemone -H "Content-Type: application/json" -d '{
"state": "Customer: I was charged twice and nobody replied.",
"questions": {
"queue": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "payments and refunds", "technical": "a product fault"}},
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["routine", "soon", "urgent", "critical"]}
}
}'
Stock llama.cpp must not be used for this. It will load the file and answer, and its answers will be wrong without saying so: a single pass over the whole prompt merges across one of the training boundaries and comes out one token shorter, which moves a probability by up to 0.28 and changes a few decisions in a thousand. Nothing is raised, because the separator the template writes is simply undefined there and renders as empty. The branch above is a requirement, not a suggestion.
Measured against the fp32 reference these weights were validated on, 255 states / 1275 answers: this file's decision agreement is noul 1.0000 ยท choice 1.0000 ยท score 1.0000, accuracy unchanged on all three, 2 of 1275 flips (both ties).
Evaluation
All numbers are accuracy on datasets the models were not trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120โ150 unless noted. Both models of the series are shown; this card's model in bold.
| benchmark | task | Jevling-0.8B-v0.1 | Jevling-E2B-v0.1 |
|---|---|---|---|
| MASSIVE (en-US) | scenario classification, 18-way | .675 | .733 |
| BBC News | topic, 5-way | .933 | .958 |
| TREC | question type, 6-way | .858 | .850 |
| PAWS | paraphrase yes/no | .508 | .625 |
| CommitmentBank | NLI, 3-way | .893 | .875 |
| StrategyQA | yes/no reasoning | .483 | .567 |
| PubMedQA | yes/no/maybe | .758 | .667 |
| SciQ | 4-way science QA | .942 | .975 |
| Social IQa | 3-way | .575 | .725 |
| TruthfulQA (MC) | multiple choice | .450 | .633 |
| XStoryCloze (en) | 2-way | .933 | .958 |
| QuALITY | long-document 4-way QA | .417 | .500 |
| RewardBench | pairwise preference | .600 | .817 |
| Hermes function-calling | tool choice | .996 | .988 |
| Financial PhraseBank | sentiment, 3-way | .608 | .658 |
| JevBench easy / original / hard (231 items) | typed decisions | 1.000 / .833 / .441 | 1.000 / .903 / .441 |
| zh-TW kiosk set (ours, synthetic-derived, 255 states) | intent acc / completeness AUROC / is-order / noise / size | .969 / .971 / .996 / 1.000 / 1.000 | .973 / .989 / .995 / 1.000 / 1.000 |
JevBench hard (.44) is where every open model we know of sits (TypeSafe's Jev API scores .73 on it); on the zh-TW kiosk set the same harness scores the two models at .971 / .973 and the Jev API at .931 โ the kiosk set is our own synthetic-derived data, so read that comparison as "fit for the distribution it was built for", not as a general claim.
Limitations
- Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4โ.5); compute arithmetic in code and put the result in the state.
- Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it.
- When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention.
- Not a chat model: it does not generate text.
Licence and release status
v0.1 is a research / non-commercial release. Some of the public datasets in the training mix carry non-commercial or research-only terms, so these weights are released under CC-BY-NC-4.0 on top of the Gemma Terms of Use (gemma-4 base).
- Downloads last month
- 212
16-bit
Model tree for BricksDisplay/jevling-e2b-v0.1-GGUF
Base model
google/gemma-4-E2B