klev-e4b

klev (KV + Jev) is a 4B decision model: given a typed question β€” a single choice, a noul (true/false) judgement, a score, or a multi-choice item β€” it returns calibrated probabilities over the options plus an explicit rejection channel. It is built on unsloth/gemma-4-e4b-it-unsloth-bnb-4bit, fine-tuned with QLoRA, and read out through a small pointer head instead of text generation. It runs 4-bit on a single 16 GB GPU.

It was built using Unsloth: the 4-bit QLoRA fine-tune, the delimiter embedding deltas, the pointer head and the teacher-KL cache all run through Unsloth's patched Gemma 4 stack (FastModel). See "Built with Unsloth" below, and the unsloth usage snippet under Usage.

The name comes from KV + Jev: a Kev-lineage pointer readout serving Jev's typed decision protocol. Working name during development was M5.

Built on Kev

klev-e4b is based on Kev and reuses part of its code. The record format, the pointer readout, the metrics and the decision-v7 data are Kev's (Apache-2.0, Jared Palmer), vendored rather than reimplemented so that klev stays byte-compatible with the Kev family β€” which is what makes the Kev-4B column in Results a like-for-like baseline:

from Kev reused here
kev/model.py (PointerHead) model/head.py β€” the q/k readout below, plus klev's own garbage candidate
kev/api.py, kev/data.py data/format.py β€” TypeSafe record format, render(), record builder
kev/metrics.py data/metrics.py β€” the scorer used for every number in Results (verbatim)
kev/suite.py + jaredpalmer/kev-suites data/suites.py β€” frozen, sha256-checked suites
kev.calibrate scripts/calibrate.py β€” the fitted pointer temperature in head.pt
decision-v7 the 15,576 training rows (the same partitions Kev-0.8B/4B/9B trained on)

Kev is Apache-2.0 and every vendored file carries a header naming its origin. No Kev model weights are used: the base is Gemma 4 E4B IT, the head and LoRA are trained fresh. What is klev's own is the Gemma base, the Unsloth QLoRA training stack, the KL anchor, the delimiter embedding deltas, the rejection channel and the stitch below.

Built with Unsloth

klev was built using Unsloth (2026.9.7), which is what makes the recipe fit one consumer GPU β€” 4-bit NF4 QLoRA of a 4B base, pointer head and cached-teacher KL on a single RTX 4080 16 GB.

stage Unsloth usage
base load unsloth.FastModel.from_pretrained(..., load_in_4bit=True)
fine-tune FastModel.get_peft_model for the rank-16 LoRA + use_gradient_checkpointing="unsloth", driven by a transformers.Trainer subclass (DistillTrainer)
trainable tokens delimiter rows via peft trainable_token_indices (Unsloth's Gemma 4 embedding path)
teacher KL cached base logits, read through Unsloth's patched hidden-state/logits access
inference / evals FastModel for every decision, chat, drift and stitch eval

Reproducing the run needs unsloth imported before transformers / peft / trl, because it patches them at import time:

pip install unsloth            # 2026.9.7
python -c "import unsloth, transformers, peft, trl"   # unsloth first

Files

file what it is
adapter/ rank-16 LoRA (alpha 32) + tokenizer with the 5 delimiter special tokens (<unused0>..<unused4>)
head.pt pointer head weights and the trainable delimiter-embedding deltas
train_config.json training hyperparameters (decision-v7, 2 epochs, KL anchor 0.3)
stitch/costbr-lora/ external third-party LoRA (CoSt-BR, pt-BR stance) trained on the same base
stitch/probe.pt 27-example nearest-class-mean probe in that LoRA's prompt space + T, beta, ext_weight
stitch/report.json stitch bench report (CoSt-BR, 718 rows)

This is not a chat model: decisions are read from the pointer head over the option/decision tokens. Reuse model/delimiters.py, model/head.py and eval/eval_decisions.py from the code repo to load and score it.

Architecture

  • Base: Gemma 4 E4B IT, 4-bit NF4 QLoRA, loaded and trained with Unsloth (FastModel).
  • Five Gemma reserved tokens (<unused0>..<unused4>) mark state / option / decision positions; their embedding rows are trained as deltas (no vocabulary resize).
  • Pointer head (vendored from Kev, see "Built on Kev"): q/k dot product over option boundary tokens, softmax over the K options plus a learned garbage candidate for rejection.
  • Trainable: 43.7M params (LoRA + deltas + head).

Training

  • decision-v7: 15,576 decision rows, 2 epochs, 8.5 h on one RTX 4080 16 GB (~6.5 s/step).
  • Objective: pointer CE over K+1 + KL anchor 0.3 to the frozen base on decision states.
  • The KL anchor keeps the representation base-aligned (drift: mean KL 1.864, 88.4 % next-token argmax agreement on held-out Alpaca prompts), which is what makes the model composable with external LoRAs (see below).

Results

Same records and scorer for all models. Best per row in bold.

suite klev-e4b Kev-4B Winnow-E4B
MMLU 0.657 0.725 0.701
ARC-Challenge 0.856 0.916 0.896
HellaSwag 0.752 0.800 0.795
Stanceosaurus ar / en / ru / es 0.421 / 0.431 / 0.437 / 0.596 0.428 / 0.485 / 0.580 / 0.655 0.449 / 0.443 / 0.583 / 0.653
CoSt-BR (majority 0.404) 0.400 0.323 0.415
in-distribution multilingual 0.785 0.794 0.823
OOD multilingual (XNLI/Belebele/XStory/PAWS-X) 0.751 0.756 0.850*

* Winnow answered 3,520/4,000 OOD rows. Per task: XNLI is flat (0.74–0.76), Belebele favours Winnow (0.877 vs 0.699), XStoryCloze favours klev (0.924 vs 0.874 Kev-4B / 0.922 Winnow).

Own decision suite: decision-v7 dev 0.860 (Brier 0.212), transfer-v4 dev 0.751. Chat-mode retention stays level with the base (MMLU 0.322 β†’ 0.389).

Stitch: importing an external LoRA with ≀30 examples

Because of the KL anchor, a LoRA trained outside this fine-tune composes with klev without wrecking its decisions (decision-v7 dev 0.860 β†’ 0.855 with stitch/costbr-lora active). Its knowledge is linearly readable from the hidden state of its own prompt format, so fusing that readout into the pointer logits with two scalars imports it:

logits_final = logits_ptr(h_sys) + beta * log p_alp
p_alp = softmax(-d^2 / T)          # NCM in the external LoRA's prompt space
CoSt-BR (718 test rows) ext weight 0.5 ext weight 1.0
klev alone (pointer) 0.405 0.389
27-example probe, LoRA space 0.457 0.475
stitched klev 0.521 0.511
external LoRA alone (generation) 0.522 0.522

Latency (RTX 4080, 4-bit, single stream):

median p90
vanilla decision 205 ms 231 ms
stitched 456 ms 495 ms
overhead +251 ms (2.23Γ—) +264 ms

The overhead is one extra forward over the LoRA's prompt; the fusion itself is negligible. Probe hyperparameters: T = 344.5, beta = 4.0, ext_weight = 0.5 (a robust cell; the train-selected cell gives 0.521 β€” see stitch/report.json).

Usage

Built and served with Unsloth β€” import unsloth must come before transformers / peft (it patches them at import time), which is why it is first below.

import unsloth  # noqa: F401  β€” must precede transformers / peft / trl
import torch
from huggingface_hub import snapshot_download
from peft import PeftModel
from unsloth import FastModel

ckpt = snapshot_download("lumierenoir/klev-e4b", allow_patterns=["adapter/*", "head.pt"])
model, processor = FastModel.from_pretrained(
    "unsloth/gemma-4-e4b-it-unsloth-bnb-4bit", max_seq_length=2048, load_in_4bit=True)
tokenizer = getattr(processor, "tokenizer", processor)
tokenizer.add_special_tokens({"additional_special_tokens":
    ["<unused0>", "<unused1>", "<unused2>", "<unused3>", "<unused4>"]})
model = PeftModel.from_pretrained(model, f"{ckpt}/adapter")
saved = torch.load(f"{ckpt}/head.pt", map_location="cpu", weights_only=False)
# apply saved["deltas"] to the delimiter embeddings and saved["head"] to PointerHead,
# then read decisions exactly as in eval/eval_decisions.py (code repo).

License and data

  • Base model: Apache-2.0 (google/gemma-4-E4B-it, unsloth/gemma-4-e4b-it-unsloth-bnb-4bit).
  • This release (adapter, head, probe): Apache-2.0.
  • Kev (github.com/jaredpalmer/kev, Apache-2.0, Jared Palmer) β€” klev is built on it and reuses its record format, PointerHead, metrics, suite loading, calibration procedure and decision-v7 data. See "Built on Kev" above.
  • Unsloth (github.com/unslothai/unsloth, Apache-2.0) β€” the training and inference stack this model was built with. See "Built with Unsloth" above.
  • Training data: decision-v7 records built from permissively licensed sets (Apache-2.0 / MIT / CC-BY-4.0), attribution manifest in the code repo.
  • stitch/costbr-lora was trained in training-conversational-stance/ on the CoSt-BR dataset (pt-BR Reddit conversations); check that dataset's terms before redistribution.

Code

Architecture, training, evals and the stitch experiment: github.com/felipepenhorate/klev (SPEC.md, docs/m5–m9, eval/stitch_demo.ipynb).

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support