decisor-4b

decisor-4b — Technical Preview v0.1.0, Apache-2.0. A flock of southern lapwings (quero-quero) on a grass field, one with wings raised in defense.

Under Labs · Code and examples · Release announcement

decisor-4b — quero-quero — Technical Preview (v0.1.0).

We are developing a family of language models with a focus on Brazilian Portuguese and Brazil's legal, tax and fiscal contexts. decisor-4b is the first technical preview in this effort, built on the Qwen3.5-4B checkpoint.

Read this first. This is a technical preview (v0.1.0) for experimentation and integration, not a production-ready deployment. This release distributes the FP8 model for SGLang only, served with the patched SGLang image from the decisor repository. Option probabilities aren't calibrated estimates of correctness. Legal and tax outputs require review by a qualified professional against current authoritative sources.

The model was trained for typed decisions in English and Brazilian Portuguese, with additional emphasis on Brazilian legal and tax contexts. Training data and the training recipe aren't released. Training was designed to preserve instruction following while specializing the model for typed decisions. We evaluated free-form generation on the BF16 reference checkpoint with a small automated battery; we haven't established broader instruction-following quality or tested generation on FP8.

The model supports two workloads:

  • Typed decisions: picks one of the supplied options for a state and question, and returns a probability distribution over all of them, using option scores from a single prefill pass.
  • Instruction following: responds to instructions and generates text. The GitHub SDK and CLI in this release focus on typed decisions.
Focus area Scope
General-purpose decisions Classification and selection across varied tasks, including English-language tasks
Brazilian Portuguese Additional training emphasis on Brazilian Portuguese language and context
Brazilian legal and tax contexts Additional domain-focused training for Brazilian legal, tax and fiscal content

The focus areas overlap. They describe training emphasis on a single model, not separate models, heads or routing modes. Support for other languages comes from Qwen3.5-4B, and performance varies by language, domain and task.

Its nickname, quero-quero, comes from the southern lapwing, the state bird of Rio Grande do Sul, Brazil.

Code and integration

decisor on GitHub provides a minimal Python SDK, the patched SGLang image, and a CLI demo with examples in English and Brazilian Portuguese.

Follow the GitHub quickstart to download the engine image, start SGLang, and try the demo.

How it works

state + question + options → prefill → option scores → decision + probabilities

The SDK formats the request using the model's decision prompt format and reads option scores from the next-token logits produced during prefill. The current engine request emits one token, but the SDK uses the scores — not the generated text — to select an option.

Successful calls return one of the supplied option IDs and probabilities normalized over the supplied options. If your task allows an "insufficient information" outcome, include it among the options and state when it should apply.

Diagram: the decisor 4B model serves two workloads — typed decisions (prompt + options, prefill readout, class distribution) and instruction following (instructions, generated response). This release distributes the FP8 model for SGLang only.

Decision prompt format

The SDK renders each request into the model's decision prompt format. You can also use it directly, without the SDK. Rendered example:

You make decisions with exactly one letter. Read the state, apply the question, and pick from the listed options.
Semantics: a field absent from the state is unknown — it is neither true nor false. Use the insufficient-information option, when listed, only if a missing fact would change the decision; otherwise decide with the facts given.

STATE (JSON):
{"policy": "Refunds are allowed for unused items within 30 days of delivery.", "request": {"days_since_delivery": 12, "item_unused": true}}
Q: Is the refund allowed under this policy?
Options: (A) Refund allowed (B) Refund not allowed (C) Not enough information
A:

Options are assigned letters in order — (A), (B), (C) — and the prompt ends with the answer marker \nA: (not option A).

The SDK sends the rendered prompt to SGLang's /generate endpoint as a raw completion. It doesn't apply the chat template or produce a thinking trace on this path.

For each option letter, the SDK checks three candidate forms: A, " A" (with a leading space) and (A, retaining only valid single-token continuations in the prompt context. It combines the logprobs of eligible forms for each option with logsumexp, then normalizes across the supplied options.

Using this prompt directly requires reproducing that token selection and scoring procedure. Sampling a letter doesn't reproduce the SDK's aggregated decision or probability distribution. The implementation is available in the decisor SDK.

Model details

Field Value
Model ID decisor-4b
Version 0.1.0 — technical preview
Publisher Under Labs
Backbone Fine-tuned from Qwen/Qwen3.5-4B, the original post-trained checkpoint
Format compressed-tensors, FP8 dynamic (per-channel weights, per-token activations) — 128 tensors in FP8 (96 MLP + 32 full-attention); 427 tensors kept in BF16 (including Gated DeltaNet (GDN) layers, norms and embeddings). "FP8" doesn't mean 8-bit everywhere.
Weight size 7.13 GB (7,125,997,144 bytes)
Evaluation context We used an 8,192-token context for the FP8 evaluations below.

Distributed format: FP8 only. This repository contains the model weights, configuration, tokenizer and chat template. The v0.1.0 weights won't change; any change to the weights ships as a new version. Pin revision="v0.1.0" to stay on this release.

Intended use

decisor-4b is intended for experimentation with classification and selection tasks where the application supplies the decision criteria and candidate outcomes.

Don't use this preview as the sole basis for consequential decisions about people, including credit, employment, benefits or sanctions. Legal and tax applications require review by a qualified professional and verification against current authoritative sources.

Usage

The supported deployment path for this release is the patched SGLang v0.5.20 image provided by the decisor repository.

The stock SGLang v0.5.20 loader doesn't correctly load this fused checkpoint; the patched SGLang image includes the required loader fix and the logprobs guard, which keeps the scheduler from crashing when a token_ids_logprob request shares a batch with one that doesn't ask for logprobs. Use it rather than the stock SGLang snippet from Hugging Face's Use this model button. The published image is underlabsai/decisor-sglang:0.5.20, digest sha256:60aee212ef25b0d303213b1c144b32f01403191b923d93d09a610770f8a6ad60.

GPU Status
NVIDIA RTX 5090 (32 GB) functional checks (demo and example requests)
NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB) all reported measurements

We haven't tested vLLM or direct Transformers inference for this release.

On first startup, the engine downloads the model weights from this Hugging Face repository and caches them for later runs.

Python

With the SDK installed from the repository's ./sdks/python directory, the engine running, and SGLANG_API_KEY available in your shell (standalone Python doesn't load .env automatically):

import os
from decisor import decide

result = decide(
    "http://127.0.0.1:8768",
    api_key=os.environ["SGLANG_API_KEY"],
    state={
        "policy": "Refunds are allowed for unused items within 30 days of delivery.",
        "request": {"days_since_delivery": 12, "item_unused": True},
    },
    question="Is the refund allowed under this policy?",
    options=[
        {"id": "approved", "text": "Refund allowed"},
        {"id": "rejected", "text": "Refund not allowed"},
        {"id": "insufficient", "text": "Not enough information"},
    ],
)
print(result["decision"])

Example result object (probabilities rounded for display):

{
  "decision": "approved",
  "options": [
    {"id": "approved", "prob": 1.0},
    {"id": "rejected", "prob": 0.0},
    {"id": "insufficient", "prob": 0.0}
  ]
}

Limits: the SDK accepts 2 to 18 options per request and rejects more; ties resolve to the first option. This is an SDK limit. We've evaluated FP8 with up to 10 options. Failures raise an exception (connection, HTTP status or timeout).

Evaluation

We ran all evaluations ourselves, under the conditions stated with each result. The BF16 reference checkpoint is our fine-tuned decisor model before FP8 quantization; we keep it internally and don't distribute it. Each result names the checkpoint it was measured on. We aren't releasing the exact adapted items and internal evaluation tools for tc193, the development set, the 1,621-question evaluation or the generation battery with this preview. Links to the source datasets provide context, but aren't enough to reproduce those evaluations exactly.

We start with decision quality on the FP8 model distributed here, followed by BF16 reference checkpoint results and FP8 throughput.

tc193

tc193 is our internal 193-item adaptation of LegalBench BR, licensed under CC BY-SA 4.0, used here for evaluation under our decision protocol (10 options per item, in a fixed order, scored by option letter). The adapted items aren't redistributed.

Seven items include the gold category name in the case text, which may give away the answer. We kept all seven in the reported score; one of them, lbbr-632, is still under review for exclusion. This is a problem with the test items, not evidence of train/test contamination, but it limits how much tc193 can tell you.

Model tc193 score Runtime and conditions
decisor-4b FP8 (this release) 142/193 (73.58%) patched SGLang image (same patches as the image we distribute), NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB), concurrency 1
decisor-4b BF16 reference checkpoint 143/193 (74.09%) same image, GPU, engine flags and concurrency as the FP8 row
Qwen3.5-4B (commit 851bf6e8) 74/193 (38.34%) in-process Transformers harness, BF16, NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB), no thinking
Most-frequent-category baseline 53/193 (27.46%) always selects "Direito Tributário"; calculated from the answer key

All three models answered the same 193 items, each with the same 10 options in the same order, using the same prompts and letter-based scoring. The runtime differs between the Qwen and decisor rows.

We trained decisor on this decision prompt format. The comparison shows how each model performs with that format. It doesn't show how Qwen would do with a prompt suited to it, and it doesn't separate the effect of training from the runtime difference.

Item by item, the BF16 reference checkpoint got 76 items right that Qwen missed and missed 7 that Qwen got right (net +69). For FP8, the counts are 75 and 7 (net +68). Between BF16 and FP8, three decisions changed: two from right to wrong and one from wrong to right.

"Direito Tributário" is the most frequent answer (53 of 193) and always appears as option J. Always choosing option A would score 33/193 (17.10%).

BF16–FP8 decision agreement

We compared the stored predictions from the BF16 and FP8 runs.

Evaluation set Same answer Different answer
tc193 190/193 3/193
Internal development set (320 questions) 316/320 4/320

Within each set, both checkpoints ran on the same patched SGLang image, engine flags and concurrency. These counts compare the selected options without using the answer keys. Agreement shows how closely FP8 preserved the BF16 decisions; both checkpoints can agree on an incorrect answer.

BF16 reference checkpoint results

We ran the evaluations in this table on the BF16 reference checkpoint, except the Qwen3.5-4B row. "Adapted" means we used a subset or our own decision protocol; those rows aren't official full-benchmark scores.

Benchmark / dataset Correct Accuracy Evaluation setup
JevBench-231 177/231 76.62% JevBench harness v1.4.1, frozen v1.2 protocol
Gevva0-650 649/650 99.85% Public suite, revision 845e11cd
BFCL (adapted) 127/130 97.69% Adapted
iSarcasmEval 126/135 93.33% Adapted
FinEntity 80/92 86.96% Adapted
SATA-Bench 130/154 84.42% Adapted
NLI4CT 39/50 78.00% Adapted
BPoMP 262/340 77.06% Adapted
WinoGrande 37/50 74.00% Adapted
GSM8K (adapted) 72/100 72.00% Adapted
HellaSwag 35/50 70.00% Adapted
CLadder 35/50 70.00% Adapted
MuSR 32/50 64.00% Adapted
cfcolor 32/50 64.00% Adapted
Amazon ESCI 32/50 64.00% Adapted
ANLI 29/50 58.00% Adapted
Humicroedit 29/50 58.00% Adapted
VAST 26/50 52.00% Adapted
CRUXEval (adapted) 25/50 50.00% Adapted
GPQA Diamond 23/50 46.00% Adapted
ARC-Challenge 14/15 93.33% Adapted
ARC-Easy 13/15 86.67% Adapted
MMLU 12/15 80.00% Adapted
SGD/SGD-X 11/15 73.33% Adapted
SimpleBench 0/10 0.00% Adapted
Total — 23 adapted datasets 1,221/1,621 75.32% Adapted; up to 18 options per question
Qwen3.5-4B on the same 1,621 questions 926/1,621 57.13% Same questions, harness and protocol as the total above

JevBench conditions: our decision prompt format and option scoring, thinking disabled, serial, one pass, no retries; reproduced twice with identical results. Self-measured — not a leaderboard entry.

Gevva0: ECE 0.0027, Brier 0.0036, OOD 120/120. The suite has no official leaderboard, and 649/650 is near its ceiling. The ECE and Brier scores describe calibration on this suite for the BF16 reference checkpoint; they don't establish calibration for the FP8 model or other tasks and datasets.

The 23 adapted datasets ran through the same in-process Transformers harness for both models, with our decision prompt format and scoring, without thinking. We trained decisor on this format and used the same format to evaluate Qwen3.5-4B. We haven't tested Qwen with a prompt optimized for it in these comparisons. The total gives each question equal weight, so larger subsets contribute more. We haven't evaluated the distributed FP8 model on this combined set.

We evaluated JevBench-231 and Gevva0-650 only on the BF16 reference checkpoint. The evaluation route we used couldn't load the FP8 model; we haven't rerun them through SGLang.

BFCL adapted measures selection under our decision protocol, not successful execution of real tools. The ARC, MMLU and SGD/SGD-X subsets contain only 15 questions each and SimpleBench contains 10 — don't generalize these scores to the full benchmarks.

Free-form generation

We evaluated the BF16 reference checkpoint on an internal free-form generation battery: 72 prompts scored automatically on a 108-point scale, with objective and keyword-based checks and partial credit.

Model Points out of 108 Runtime
decisor-4b BF16 reference checkpoint 73.0 Separate SGLang setup; memory and context settings weren't recorded
Qwen3.5-4B 67.5 SGLang; 32,768-token context

Both runs used non-thinking mode with the same frozen generation parameters (temperature 0). decisor scored 5.5 points higher on this battery, but the runtime settings weren't fully matched. The scores aren't a count of correct responses, and separate writing-rubric and model-judge reviews aren't included.

We haven't run this evaluation on the distributed FP8 model.

Decision throughput — FP8 (SGLang)

We measured throughput with a dedicated evaluation harness under controlled prefix-cache conditions. The repository disables the radix cache by default, so these numbers don't describe the default setup. We used the FP8 model on an NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB) with a patched SGLang image carrying the same patches as the image we distribute. This isn't an end-to-end benchmark of the current Python SDK or CLI.

The workload was the 193 tc193 prompts at concurrency 1, with cache conditions verified through per-request token accounting. The cold result is one exploratory repetition with an initially empty prefix cache. For the warm result, we report the median of three repetitions with a prewarmed prefix cache, without restarting the engine between them.

Condition (FP8) Valid decisions/s
cold, concurrency 1 (exploratory, 1 rep) 33.7
warm, concurrency 1 (median of 3) 49.5

Valid-decision throughput excludes requests whose option scores tied after rounding to four decimal places.

Cold and warm describe the cache state during measurement.

Bar chart: valid-decision throughput at concurrency 1 — cold 33.7 decisions/s (one exploratory repetition), warm 49.5 decisions/s (median of three repetitions with a prewarmed prefix cache).

Measurement protocol

Across the four measurement windows, 772 requests completed without failure or timeout. Of these, 768 counted toward valid-decision throughput; four were excluded under the tie rule above.

The engine ran with --mem-fraction-static 0.35 --max-running-requests 16 --context-length 8192 for these windows.

Peak VRAM usage across these windows was 37,135–37,249 MiB. This is configuration-specific memory usage, not a minimum VRAM requirement.

Limitations

What we tested

  • FP8 inference with the patched SGLang image on an NVIDIA RTX 5090 (32 GB) and an RTX PRO 6000 Blackwell Server Edition (96 GB).
  • SDK and CLI integration with the running FP8 model, including four example requests in English and Brazilian Portuguese that produced their expected decisions (functional checks, not a quality benchmark).
  • FP8 decision quality on the 193-item tc193 adaptation, subject to the evaluation caveats above.
  • BF16–FP8 decision agreement on tc193 and a 320-question development set.
  • FP8 throughput on the RTX PRO 6000 under the specific harness, workload and cache conditions reported above.

Evaluation boundaries

  • We haven't tested smaller GPUs. Weight-file size alone doesn't show runtime compatibility or sufficient VRAM.
  • We measured JevBench-231, Gevva0-650, the 23-dataset table and the generation battery on the BF16 reference checkpoint.
  • We haven't repeated our earlier end-to-end testing of the BF16 setup in full for FP8. The current integration checks don't replace it.
  • Our FP8 evaluations used 10 options per item in tc193 and four in an internal development set of 320 questions; we report agreement, not accuracy, for that set. We haven't evaluated FP8 with more options.
  • We used an 8,192-token context for the FP8 evaluations reported here. These evaluations don't establish quality at longer contexts.
  • The tools we used to create the FP8 weights included a dependency combination outside their declared support ranges. This concerns the conversion environment, which is separate from the patched SGLang image used to serve the model.
  • We haven't run a dedicated tax evaluation. The tc193 results cover legal tasks only and don't show how the model performs on tax content.
  • We compared Qwen3.5-4B with decisor on typed decisions and free-form generation. The decision evaluations use our decision prompt format; we trained decisor on this format and used the same format to evaluate Qwen3.5-4B. We haven't tested Qwen with a prompt optimized for it in these comparisons. The 1,621-question comparison used the same evaluation harness for both models; the tc193 and generation comparisons used different runtime setups, described alongside their results.

Behavioral limitations

Option probabilities aren't calibrated. The model can be confidently wrong, and its decisions can change with wording, language and how the input is laid out.

In small exploratory tests, changing the field order or the policy wording changed some decisions. In one case, instructions hidden in the input data flipped a correct decision to a wrong one. These tests weren't systematic, so they don't tell you how often this happens in practice.

Treat instructions inside untrusted input as a prompt-injection risk. Keeping your policy separate from user content makes prompts clearer, but the model won't always respect that boundary.

Write options that are clearly distinct, say when the insufficient-information option applies, and test the model on your own representative and adversarial cases. Keeping the field order fixed doesn't make the model insensitive to it.

In the C=1 serving benchmark, FP8 scored 142/192 in the initially empty-cache window and 144/192 in each of the three prewarmed windows. One tied-score request was excluded from each FP8 window. The BF16 reference checkpoint scored 143/193 in all four windows, with no ties excluded. These counts use the serving benchmark's tie-exclusion rule. They are separate from the 142/193 FP8 quality result above and aren't enough to establish a general accuracy advantage over BF16.

Generated responses can include fabricated facts, legal citations or tax guidance. Check them against current authoritative sources.

License and attribution

Component License / attribution
Model weights Apache-2.0. See LICENSE and NOTICE.
Python SDK and integration code Apache-2.0. See the decisor repository.
Backbone Fine-tuned from Qwen3.5-4B, licensed under the Apache License, Version 2.0.

The patched SGLang image includes a fused-checkpoint loader fix and the logprobs guard from SGLang PR #35052, with an in-build regression test from PR #35852. Upstream license files are preserved.

Copyright 2026 Under Serviços de Internet Ltda.

Developed by Under Labs. Provided as-is, without a commitment to user support.

Citation

@misc{decisor4b2026,
  title  = {decisor-4b},
  author = {{Under Labs}},
  year   = {2026},
  url    = {https://huggingface.co/underlabs/decisor-4b},
  note   = {Technical preview, version 0.1.0}
}

For the SDK and serving integration:

@software{decisor2026,
  title  = {decisor: Python SDK and SGLang runtime for decisor-4b},
  author = {{Under Labs}},
  year   = {2026},
  url    = {https://github.com/underlabs-ai/decisor},
  note   = {Model: https://huggingface.co/underlabs/decisor-4b}
}

Contact

For non-sensitive bug reports and reproducible model-behavior issues, use GitHub Issues. Do not include API keys, personal data or confidential inputs.

To report a security vulnerability, use GitHub private vulnerability reporting. Do not open a public issue.

Downloads last month
207
Safetensors
Model size
5B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for underlabs/decisor-4b

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(897)
this model