Instructions to use Jwuthrich/selfjev-4b-vision-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Jwuthrich/selfjev-4b-vision-8bit with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="Jwuthrich/selfjev-4b-vision-8bit")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Jwuthrich/selfjev-4b-vision-8bit") model = AutoModelForCausalLM.from_pretrained("Jwuthrich/selfjev-4b-vision-8bit", device_map="auto") - Notebooks
- Google Colab
- Kaggle
SelfJev-4B Vision, 8-bit (LLM.int8)
Structured decisions from text. A 4B backbone, 8-bit, for 12 GB GPUs (the 4-bit build is the one for 8 GB).
SelfJev-4B Vision with its LoRA merged into Qwen3.5-4B and the weights quantized to 8-bit (LLM.int8) with bitsandbytes (LLM.int8 8-bit). It answers yes/no, one-choice, scored and multi-select questions over the options you supply and returns probabilities, through the SelfJev engine and API. The checkpoint is 4.8 GB (the full merged model is about 9.3 GB).
Quickstart
Needs an NVIDIA GPU (CUDA) and selfjev 0.4.0 or later.
pip install "selfjev[serve,gpu,quant]"
selfjev serve --quantized-model Jwuthrich/selfjev-4b-vision-8bit
--quantized-model takes this repository (or a local folder) in place of --adapter. The first start also downloads the
Qwen3.5-4B tokenizer; the vision tower is only fetched from the base model if you send an image. The HTTP API is the same as
the full model's: API guide.
Quality
Scored on all 3,657 questions of SelfJev Decision Bench through the native tree engine on an A10G whose memory was capped at 7.2 GiB (the usable memory of an 8 GB card).
| Build | Decision Bench (3,657) | Text decisions (1,991) | AI response review (946) | Compact challenge (720) |
|---|---|---|---|---|
| bf16, full model | 95.9% | 96.1% | 92.5% | 100% |
| 8-bit (LLM.int8) (this repository) | 95.3% | 95.2% | 92.0% | 100% |
| 4-bit (sibling build) | 95.0% | 94.7% | 92.0% | 100% |
The suites are authored with AI-checked labels and have informed research decisions, so read them as a development benchmark. The full-model rows are the vision release's own reports; the compact bf16 figure was re-run on the same A10G. Reports: reports/quant.
Memory and speed
Peak GPU memory for one question and the median time of one request (A10G, --max-batch-tokens 4096, selfjev 0.4.2;
memory is what the tensors allocate):
| Text length | 8-bit (LLM.int8) | bf16 (full model) |
|---|---|---|
| 2,048 tokens | 4.8 GiB, 558 ms | 8.3 GiB, 434 ms |
| 8,192 tokens | 5.8 GiB, 1,767 ms | 9.2 GiB, 1,545 ms |
| 16,384 tokens | out of memory at 7.2 GiB | 10.4 GiB, 3,159 ms |
This build does not fit an 8 GB card for long texts (it ran out of memory at 16K tokens and on the longest text-decision questions under the 7.2 GiB cap); use the 4-bit build there. It suits 12 GB cards. Its text-decision score (95.2%) was measured without the memory cap.
Limits
- Scored on text only. Images use the base model's vision tower in bf16 (not quantized, loaded on the first image); the quantized path has not been scored on images.
- Texts above 16K tokens were not validated (the model was trained on texts up to 16K).
- Quantization costs accuracy: 0.6 points on the pooled benchmark and 0.9 on text decisions; it is also slower than 4-bit (bitsandbytes int8 matmuls). The numbers above are the measurement; there is no calibration step to recover it.
- Same license and intended use as the full model; built from Qwen3.5-4B (Apache 2.0).
How this was made
selfjev quantize --adapter weights/selfjev_4b_vision --bits 8bit: the adapter is merged into the pinned Qwen3.5-4B
(851bf6e) in bf16, then loaded with bitsandbytes LLM.int8 8-bit and saved. Source:
Jwuthri/SelfJev.
- Downloads last month
- 26