basal-1.5-max

GitHub Website Collection

basal-1.5-max (11B) is the largest and most accurate model of the family, for long documents, judging and hard English decisions. Part of the basal-1.5 family of typed-decision models for Polish and English.

What it is. Inspired by System 1 (fast, intuitive) decision models such as Jev: instead of writing an answer, the model reads a state (a message, a document, retrieved passages, an agent trace, a web page as JSON) and answers a typed question about it (choice, yes/no noul, score, and in the basal engine multi and act) by returning a calibrated probability for each allowed answer, in one forward pass, without generating text. The answer never falls outside the options you give, and the probability says how sure the model is.

What it is for: a dynamic classifier. The classes are described in the request, in plain language, so one model serves many tasks without retraining: routing with changing categories, rule and policy checks, questions about long documents, RAG decisions (relevance, sufficiency, conflicts, claim support, "cannot be determined"), urgency and risk scores, judging answers, agent and tool decisions, and triage with a confidence threshold.

model params use
basal-1.5 4.5B main model
basal-1.5-max 11B most accurate in the family
basal-1.5-mini 1.5B fastest, distilled from max

Each model also comes as -MLX-8bit and -MLX-fp4 (Apple Silicon) and -GGUF (Ollama, llama.cpp). NVIDIA ModelOpt checkpoints for vLLM: -FP8 and -NVFP4.

Results

Werdykt (our hidden Polish–English benchmark of typed decisions: 10 categories × 2 languages, one protocol for every model, open models on one H100; macro accuracy over the 20 cells): 0.773 (Polish 0.777, English 0.769), median latency 67.1 ms per decision with SGLang, $0.108 per 1,000 decisions. Leaderboard and public samples: https://basal.si5.pl/.

benchmark basal-1.5-max (11B) basal-1.0-4.5B
Werdykt v1, macro 0.773 0.660
PL sealed test (3,000 new decisions, scored once) 0.933 0.874
PL decisions (held-out test of basal-1.0, 7,081 items) 0.940 0.886
PL general (knowledge, exams, reading; 3,079 items) 0.794 0.737
EN external panel (10 public English decision tasks, 3,842 items) 0.737 0.671
Public EN decision benchmark (231 items; in-domain for 1.5) 0.848 0.740
Coverage at 1% error (share decided automatically with the shipped threshold) 62.2% 55.2%

All served with the basal engine on one H100 (--mode fast, bf16), both option orders averaged. Definitions: README.

Against basal-1.5 (4.5B), fp32 readouts of both release checkpoints: held-out test 0.893 vs 0.875 (+3.5 over basal-1.0-4.5B's 0.858), Polish general knowledge test 0.799 vs 0.745, decided automatically at the 1% / 5% error targets 62.2% / 84.5% vs 51.0% / 76.6%; held-out dev sets: action selection 0.903, ordinal scores 0.867, hard English decisions 0.972. On the English external panel basal-1.5-max scores 0.737 against 0.671 for basal-1.0-4.5B (+6.9, paired 95% interval [+5.3, +8.5]), higher on Circa as well.

Quick start

1. Install into a fresh uv environment (no git needed):

uv venv --python 3.12 ~/basal-env && source ~/basal-env/bin/activate
uv pip install torch==2.11.0 --index-url https://download.pytorch.org/whl/cu128
uv pip install "basal[fp8] @ https://github.com/rkinas/basal/archive/refs/tags/v1.5.0.tar.gz"

Install torch first, from the CUDA 12.8 index: the newest torch on PyPI may need a newer GPU driver, and the torchvision preinstalled on cloud GPU images breaks transformers. DGX Spark and B300: use the cu130 build (hardware notes).

2. Start the server and wait until it prints basal: model ... ready:

basal-serve --model Remek/basal-1.5-max --mode fast --port 8000

The first start in mode fast compiles the model for a few minutes; --mode fast-nocompile starts in seconds and --mode fp8 is faster on workstation, consumer and desktop GPUs. vLLM and SGLang: --mode vllm / --mode sglang in separate environments (vLLM: basal[vllm]; SGLang: install sglang[srt]==0.5.21 first, then basal with --no-deps, as in the README). Apple Silicon: basal-1.5-max-MLX-8bit; Ollama: basal-1.5-max-GGUF. FP8 / NVFP4 checkpoints for vLLM: basal-1.5-max-FP8 · basal-1.5-max-NVFP4.

3. Ask questions (in a second terminal). Several questions about one state go into one request and are answered over one shared state:

curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
  "state": "Klient: od wczoraj nie mogę zalogować się do bankowości internetowej, system pokazuje błąd hasła.",
  "questions": {
    "dept":   {"type": "choice", "instructions": "Do którego działu skierować zgłoszenie?",
               "criteria": {"cards": "Reklamacje kart", "online": "Wsparcie bankowości elektronicznej", "loans": "Kredyty"}},
    "urgent": {"type": "noul", "instructions": "Czy klient nie może korzystać z usługi?"}}}'

The response has a calibrated probability for every option, the chosen key and the confidence:

{"model": "basal-1.5-max",
 "answers": {"dept": {"type": "choice", "choice": "online", "probabilities": {"cards": …, "online": …, "loans": …}, "confidence": …},
             "urgent": {"type": "noul", "noul": …, "probabilities": {"true": …, "false": …}, "confidence": …}},
 "usage": {"input_tokens": …, "output_tokens": 0, "questions": 2, "branches": 2, "latency_ms": …}}

From Python (the basal package includes a client):

from basal.client import Basal
b = Basal("http://127.0.0.1:8000")
state = "Klient: od wczoraj nie mogę zalogować się do bankowości internetowej."
a = b.choice(state, "Do którego działu skierować zgłoszenie?",
             {"cards": "Reklamacje kart", "online": "Wsparcie bankowości elektronicznej", "loans": "Kredyty"})
print(a["choice"], a["confidence"])                                      # chosen key and its probability
print(b.yes_no(state, "Czy klient zgłasza problem techniczny?")["noul"])  # P(yes)
print(b.score(state, "Jak pilne jest zgłoszenie?", ["niska", "średnia", "wysoka"])["score"])  # expected level 0-2

Many items at once: basal-run --input items.jsonl --output answers.jsonl. Full API (multi, act, "evidence": true, "facts": "auto", option keys): API and features.

Engines

The status column refers to basal-1.5 (4.5B); basal-1.5-max uses the same engines and API, and its own measurements are listed where available.

engine basal-serve mode hardware what it is for quick start status (basal-1.5, 4.5B)
basal engine (PyTorch) fast (also fast-nocompile, fp8, nvfp4); eager --device mps|cpu NVIDIA GPUs (sm80+); Apple Silicon (MPS) and CPU with eager the primary server: packed requests (SOAM), evidence spans, fp8 on Hopper and Blackwell basal-serve --model Remek/basal-1.5-max --mode fast verified: served check on one H100, 13.7 ms p50 per question over HTTP, 64.4 decisions/s with 32 clients
vLLM vllm NVIDIA GPUs an alternative CUDA server basal-serve --model Remek/basal-1.5-max --mode vllm (extra basal[vllm]) verified offline (one H100, basal-bench, 1,000 development decisions): agreement 0.994 with the fp32 engine, accuracy 0.928 vs 0.930, 147 decisions/s
SGLang sglang NVIDIA GPUs an alternative CUDA server (RadixAttention prefix sharing) basal-serve --model Remek/basal-1.5-max --mode sglang (own environment: sglang[srt]==0.5.21, then basal with --no-deps) verified offline (one H100, basal-bench, 1,000 development decisions): agreement 0.995 with the fp32 engine, accuracy 0.925 vs 0.930, 104.6 decisions/s; on Werdykt the same answer as the basal engine on 98.3% of items (97.0% for states over 8k tokens)
MLX mlx Apple Silicon Macs: -MLX-8bit (recommended) or -MLX-fp4 (half the memory) basal-serve --model Remek/basal-1.5-max-MLX-8bit --mode mlx (extra basal[mlx]) verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.992 (-MLX-8bit) / 0.970 (-MLX-fp4), 587 / 581 ms per question
Ollama ollama CPU, Apple Silicon, consumer GPUs local and desktop use with Ollama (-GGUF) 1. hf download Remek/basal-1.5-max-GGUF --local-dir basal-1.5-max-GGUF 2. in that folder: ollama create basal-1.5-max:q8_0 -f Modelfile.Q8_0 3. basal-serve --model ./basal-1.5-max-GGUF --mode ollama --ollama-model basal-1.5-max:q8_0 verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.994 (Q8_0) / 0.980 (Q4_K_M), 518 / 542 ms per question
llama.cpp llamacpp CPU, Apple Silicon, consumer GPUs llama-server with the -GGUF files; the engine's token ids are sent as is 1. llama-server -m basal-1.5-max-GGUF/basal-1.5-max-Q8_0.gguf -c 4096 --port 8080 2. basal-serve --model ./basal-1.5-max-GGUF --mode llamacpp verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.994 (Q8_0) / 0.980 (Q4_K_M), 427 / 446 ms per question

basal-serve is the decision layer in every row. vLLM, SGLang, MLX, Ollama and llama.cpp run the weights; the typed answers (a calibrated probability for each of your options, both option orders, the state shared by several questions, multi, act and facts) come from basal-serve --mode <engine> in front of them, with the same HTTP API in every mode. A plain ollama run or llama-cli gives free text only. Evidence spans need the PyTorch engine (fast, eager); vLLM and SGLang start with a 4,096-token context (longer states are refused); raise it with --max-len (e.g. 32768 for long documents). Install and first request for each engine: quick start.

Speed

One decision = one question in both option orders (the default), median latency. About 22 GB of GPU memory for the bf16 weights; on 24 GB cards use the 4.5B model (--mode fp8 was not measured for max).

Measured on one H100 (other GPUs were not measured for max):

engine measurement result
basal engine, fast (bf16) HTTP load test, one question per request, 32 concurrent clients 27.3 ms p50 / 46.6 ms p95 for one question, 33.4 decisions/s; bf16 with SOAM agrees with the fp32 reference on 0.993 of 1,500 decisions
vLLM – not measured for max (verified on basal-1.5: agreement 0.994 with the fp32 engine)
SGLang Werdykt, 5,000 decisions, against the basal engine's fast mode the same answer on 99.4% of decisions (98.4% for states over 8k tokens); not measured offline for max

The HTTP and offline numbers are not comparable with each other. Apple Silicon (MLX, Apple M4 Pro, 24 GB): 1.36 s per decision.

Several questions per request share the state: in an in-process engine comparison on one H100 (4.5B, same node and requests) 5 questions take 37.8 ms instead of 111.7 ms as separate prompts, 12 questions 80.3 ms instead of 222.0 ms (SOAM). Every GPU we measured: HARDWARE.md.

Files

  • model.safetensors, tokenizer, chat_template.jinja: the model (Llama architecture, 50 layers, 32k Polish vocabulary)
  • CALIBRATION.json: per-type temperatures and confidence thresholds for 1% / 5% error (applied by the basal server, also in the vLLM, SGLang, MLX and Ollama modes)
  • evidence_head.pt: the pointer head behind "evidence": true (experimental)
  • basal.json: prompt format and readout protocol

How it works

The model receives a fixed chat prompt with the state, the question and lettered options; the assistant turn is prefilled with {"answer": " and the decision is the softmax over the next-token logits of the option letters only (one forward pass, no text generation). The server asks every question with the options in original and reversed order and averages the two distributions (this reduces sensitivity to option order), then applies the calibrated temperature of the question type from CALIBRATION.json, fitted on exactly this averaged prediction.

Use the confidence. CALIBRATION.json stores confidence thresholds chosen on calibration data, before testing, for a target error of 1% or 5% among accepted decisions; on the test data they accept 62.2% / 84.5% of decisions at 0.58% / 4.1% observed error. Accept decisions above the threshold automatically and route the rest to a person. The thresholds are validated on descriptions-only prompts ("option_keys": "hide") and yes/no questions; for other formats, refit them on your own labelled requests.

How to ask. One condition per yes/no question (or a multi with one label per condition); choice with described sides for named binary outcomes; ask in the language of the state; put the facts and rules in the state; offer "cannot be determined" when the state may not contain the answer; compute exact numbers in code (or send "facts": "auto" for Polish dates and amounts).

Training

Fine-tuned from speakleash/Bielik-PL-11B-v3.0-Instruct for typed decisions in Polish and English: rules and deadlines, real documents, long documents, retrieval decisions, contracts, routing, ordinal scores, judging answers, action selection and "cannot be determined" answers. The recipe includes a reinforcement-learning (RL) stage.

Limitations

  • Validate on your own documents before relying on it: hidden and held-out test sets do not cover every domain.
  • Polish world knowledge of a small model is limited: provide the relevant facts in the state. Legal rules change; the model does not know rules introduced after its training.
  • Exact arithmetic and thresholds are a weak spot of one-pass decisions: compute them in code or use "facts": "auto".
  • Averaging the original and reversed option order reduces, but does not remove, sensitivity to option order for three or more options. Up to 10 options per question.
  • act and evidence spans are experimental (features).
  • Decisions with serious consequences for people should be reviewed by a person.

Citation

@software{kinas2026basal15,
  title   = {basal-1.5: typed-decision models and inference engine for Polish and English},
  author  = {Kinas, Remigiusz},
  year    = {2026},
  version = {1.5.0},
  url     = {https://github.com/rkinas/basal}
}

The method of the previous release is described in the basal-1.0 technical report (PDF, doi:10.5281/zenodo.23022986).

License and attribution

Apache-2.0. Fine-tuned from speakleash/Bielik-PL-11B-v3.0-Instruct (Apache-2.0).

Downloads last month
101
Safetensors
Model size
11B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Remek/basal-1.5-max

Finetuned
(1)
this model
Quantizations
5 models

Space using Remek/basal-1.5-max 1

Collection including Remek/basal-1.5-max