Instructions to use lumierenoir/klev-e4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use lumierenoir/klev-e4b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Desktop
klev-e4b
klev (KV + Jev) is a 4B decision model: given a typed question β a single choice,
a noul (true/false) judgement, a score, or a multi-choice item β it returns calibrated
probabilities over the options plus an explicit rejection channel. It is built on
unsloth/gemma-4-e4b-it-unsloth-bnb-4bit, fine-tuned with QLoRA, and read out through a
small pointer head instead of text generation. It runs 4-bit on a single 16 GB GPU.
It was built using Unsloth: the 4-bit QLoRA
fine-tune, the delimiter embedding deltas, the pointer head and the teacher-KL cache all run
through Unsloth's patched Gemma 4 stack (FastModel). See "Built with Unsloth" below, and
the unsloth usage snippet under Usage.
The name comes from KV + Jev: a Kev-lineage pointer readout serving Jev's typed decision protocol. Working name during development was M5.
Built on Kev
klev-e4b is based on Kev and reuses part of its code. The record format, the pointer readout, the metrics and the decision-v7 data are Kev's (Apache-2.0, Jared Palmer), vendored rather than reimplemented so that klev stays byte-compatible with the Kev family β which is what makes the Kev-4B column in Results a like-for-like baseline:
| from Kev | reused here |
|---|---|
kev/model.py (PointerHead) |
model/head.py β the q/k readout below, plus klev's own garbage candidate |
kev/api.py, kev/data.py |
data/format.py β TypeSafe record format, render(), record builder |
kev/metrics.py |
data/metrics.py β the scorer used for every number in Results (verbatim) |
kev/suite.py + jaredpalmer/kev-suites |
data/suites.py β frozen, sha256-checked suites |
kev.calibrate |
scripts/calibrate.py β the fitted pointer temperature in head.pt |
| decision-v7 | the 15,576 training rows (the same partitions Kev-0.8B/4B/9B trained on) |
Kev is Apache-2.0 and every vendored file carries a header naming its origin. No Kev model weights are used: the base is Gemma 4 E4B IT, the head and LoRA are trained fresh. What is klev's own is the Gemma base, the Unsloth QLoRA training stack, the KL anchor, the delimiter embedding deltas, the rejection channel and the stitch below.
Built with Unsloth
klev was built using Unsloth (2026.9.7), which is what makes the recipe fit one consumer GPU β 4-bit NF4 QLoRA of a 4B base, pointer head and cached-teacher KL on a single RTX 4080 16 GB.
| stage | Unsloth usage |
|---|---|
| base load | unsloth.FastModel.from_pretrained(..., load_in_4bit=True) |
| fine-tune | FastModel.get_peft_model for the rank-16 LoRA + use_gradient_checkpointing="unsloth", driven by a transformers.Trainer subclass (DistillTrainer) |
| trainable tokens | delimiter rows via peft trainable_token_indices (Unsloth's Gemma 4 embedding path) |
| teacher KL | cached base logits, read through Unsloth's patched hidden-state/logits access |
| inference / evals | FastModel for every decision, chat, drift and stitch eval |
Reproducing the run needs unsloth imported before transformers / peft / trl,
because it patches them at import time:
pip install unsloth # 2026.9.7
python -c "import unsloth, transformers, peft, trl" # unsloth first
Files
| file | what it is |
|---|---|
adapter/ |
rank-16 LoRA (alpha 32) + tokenizer with the 5 delimiter special tokens (<unused0>..<unused4>) |
head.pt |
pointer head weights and the trainable delimiter-embedding deltas |
train_config.json |
training hyperparameters (decision-v7, 2 epochs, KL anchor 0.3) |
stitch/costbr-lora/ |
external third-party LoRA (CoSt-BR, pt-BR stance) trained on the same base |
stitch/probe.pt |
27-example nearest-class-mean probe in that LoRA's prompt space + T, beta, ext_weight |
stitch/report.json |
stitch bench report (CoSt-BR, 718 rows) |
This is not a chat model: decisions are read from the pointer head over the
option/decision tokens. Reuse model/delimiters.py, model/head.py and
eval/eval_decisions.py from the code repo to load and score it.
Architecture
- Base: Gemma 4 E4B IT, 4-bit NF4 QLoRA, loaded and trained with Unsloth (
FastModel). - Five Gemma reserved tokens (
<unused0>..<unused4>) mark state / option / decision positions; their embedding rows are trained as deltas (no vocabulary resize). - Pointer head (vendored from Kev, see "Built on Kev"):
q/kdot product over option boundary tokens, softmax over the K options plus a learned garbage candidate for rejection. - Trainable: 43.7M params (LoRA + deltas + head).
Training
- decision-v7: 15,576 decision rows, 2 epochs, 8.5 h on one RTX 4080 16 GB (~6.5 s/step).
- Objective: pointer CE over K+1 + KL anchor 0.3 to the frozen base on decision states.
- The KL anchor keeps the representation base-aligned (drift: mean KL 1.864, 88.4 % next-token argmax agreement on held-out Alpaca prompts), which is what makes the model composable with external LoRAs (see below).
Results
Same records and scorer for all models. Best per row in bold.
| suite | klev-e4b | Kev-4B | Winnow-E4B |
|---|---|---|---|
| MMLU | 0.657 | 0.725 | 0.701 |
| ARC-Challenge | 0.856 | 0.916 | 0.896 |
| HellaSwag | 0.752 | 0.800 | 0.795 |
| Stanceosaurus ar / en / ru / es | 0.421 / 0.431 / 0.437 / 0.596 | 0.428 / 0.485 / 0.580 / 0.655 | 0.449 / 0.443 / 0.583 / 0.653 |
| CoSt-BR (majority 0.404) | 0.400 | 0.323 | 0.415 |
| in-distribution multilingual | 0.785 | 0.794 | 0.823 |
| OOD multilingual (XNLI/Belebele/XStory/PAWS-X) | 0.751 | 0.756 | 0.850* |
* Winnow answered 3,520/4,000 OOD rows. Per task: XNLI is flat (0.74β0.76), Belebele favours Winnow (0.877 vs 0.699), XStoryCloze favours klev (0.924 vs 0.874 Kev-4B / 0.922 Winnow).
Own decision suite: decision-v7 dev 0.860 (Brier 0.212), transfer-v4 dev 0.751. Chat-mode retention stays level with the base (MMLU 0.322 β 0.389).
Stitch: importing an external LoRA with β€30 examples
Because of the KL anchor, a LoRA trained outside this fine-tune composes with klev without
wrecking its decisions (decision-v7 dev 0.860 β 0.855 with stitch/costbr-lora active).
Its knowledge is linearly readable from the hidden state of its own prompt format, so
fusing that readout into the pointer logits with two scalars imports it:
logits_final = logits_ptr(h_sys) + beta * log p_alp
p_alp = softmax(-d^2 / T) # NCM in the external LoRA's prompt space
| CoSt-BR (718 test rows) | ext weight 0.5 | ext weight 1.0 |
|---|---|---|
| klev alone (pointer) | 0.405 | 0.389 |
| 27-example probe, LoRA space | 0.457 | 0.475 |
| stitched klev | 0.521 | 0.511 |
| external LoRA alone (generation) | 0.522 | 0.522 |
Latency (RTX 4080, 4-bit, single stream):
| median | p90 | |
|---|---|---|
| vanilla decision | 205 ms | 231 ms |
| stitched | 456 ms | 495 ms |
| overhead | +251 ms (2.23Γ) | +264 ms |
The overhead is one extra forward over the LoRA's prompt; the fusion itself is negligible.
Probe hyperparameters: T = 344.5, beta = 4.0, ext_weight = 0.5 (a robust cell; the
train-selected cell gives 0.521 β see stitch/report.json).
Usage
Built and served with Unsloth β import unsloth must come before transformers / peft
(it patches them at import time), which is why it is first below.
import unsloth # noqa: F401 β must precede transformers / peft / trl
import torch
from huggingface_hub import snapshot_download
from peft import PeftModel
from unsloth import FastModel
ckpt = snapshot_download("lumierenoir/klev-e4b", allow_patterns=["adapter/*", "head.pt"])
model, processor = FastModel.from_pretrained(
"unsloth/gemma-4-e4b-it-unsloth-bnb-4bit", max_seq_length=2048, load_in_4bit=True)
tokenizer = getattr(processor, "tokenizer", processor)
tokenizer.add_special_tokens({"additional_special_tokens":
["<unused0>", "<unused1>", "<unused2>", "<unused3>", "<unused4>"]})
model = PeftModel.from_pretrained(model, f"{ckpt}/adapter")
saved = torch.load(f"{ckpt}/head.pt", map_location="cpu", weights_only=False)
# apply saved["deltas"] to the delimiter embeddings and saved["head"] to PointerHead,
# then read decisions exactly as in eval/eval_decisions.py (code repo).
License and data
- Base model: Apache-2.0 (
google/gemma-4-E4B-it,unsloth/gemma-4-e4b-it-unsloth-bnb-4bit). - This release (adapter, head, probe): Apache-2.0.
- Kev (
github.com/jaredpalmer/kev, Apache-2.0, Jared Palmer) β klev is built on it and reuses its record format,PointerHead, metrics, suite loading, calibration procedure and decision-v7 data. See "Built on Kev" above. - Unsloth (
github.com/unslothai/unsloth, Apache-2.0) β the training and inference stack this model was built with. See "Built with Unsloth" above. - Training data: decision-v7 records built from permissively licensed sets (Apache-2.0 / MIT / CC-BY-4.0), attribution manifest in the code repo.
stitch/costbr-lorawas trained intraining-conversational-stance/on the CoSt-BR dataset (pt-BR Reddit conversations); check that dataset's terms before redistribution.
Code
Architecture, training, evals and the stitch experiment:
github.com/felipepenhorate/klev (SPEC.md, docs/m5βm9, eval/stitch_demo.ipynb).
- Downloads last month
- -