DECISIONx

A decision model. Ask a typed question about a piece of text; get calibrated probabilities back. It never generates text.

DECISIONx answers the small decisions inside software — which queue gets this ticket, is this email spam, how urgent is this incident — in a single forward pass, with a probability you can put in an if statement.

Question type You give it You get back
choice any list of options the best option, and a probability for every option
score a scale (e.g. 1–5, with optional labels) the expected score, its spread, and the full distribution
null a yes/no statement the probability that it is true

Options can be anything, in any order. They do not need to have been seen in training.

Quick start

pip install torch transformers peft safetensors huggingface_hub
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("thinkingdbx/DECISIONx-1.5B")
sys.path.insert(0, path)                      # the small inference package ships with the weights

from decisionx.model import DecisionModel
m = DecisionModel.load(path, device="cuda")   # or "cpu"; the base model downloads on first use

m.decide("choice", "Which team should handle this ticket?",
         state="I was charged twice for my order last week and I want the extra payment back.",
         options=["billing", "shipping", "technical support", "account security"])

What it answers

Real outputs from this checkpoint (CPU, no calibration):

state question answer
choice "I was charged twice for my order last week…" Which team should handle this ticket? (billing / shipping / technical support / account security) billing — 0.92 (technical support 0.08)
choice "Can you move my 3pm meeting with Priya to tomorrow morning?" What kind of request is this? (calendar / email / travel / weather / music) calendar — 0.9996
null "The package arrived damaged and I would like my money back please." The customer is asking for a refund. 0.94 true
null "Subject: Q3 planning — Hi team, attaching the draft roadmap…" This email is spam. 0.0006
null "Subject: You've been selected! Congratulations, you have won a free cruise…" This email is spam. 0.149 ⚠️
score "Production database is down and no customer can log in." How urgent is this? (1 can wait … 5 drop everything) 3.9 ± 1.0 (mode 4)
score "Could you update the footer copyright year…when you get a chance?" How urgent is this? 2.3 ± 1.0 (mode 2)

The spam row is the model's main weakness, shown on purpose: it rates the cruise email 230× more spam-like than the planning email, so its ranking is right, but on a task it has never seen it can put the threshold in the wrong place. That is what task calibration is for.

Calibrating a new task

Give it 20–60 labelled examples of your decision; it fits a temperature and a per-option bias for that task, and the answers you get afterwards are both better placed and honestly confident.

from decisionx.schema import Decision
from decisionx.calibration import TaskCalibration

examples = [Decision("null", email, "This email is spam.", label=is_spam)
            for email, is_spam in my_labelled_emails[:32]]
cal = m.calibrate(examples)              # strength chosen by cross-validation on your examples
cal.save("spam.calibration.json")

m.decide("null", "This email is spam.", state=new_email,
         calibration=TaskCalibration.load("spam.calibration.json"))

Biases are keyed by option text, so option order does not matter. On a task that is already well calibrated, cross-validation will usually choose to change nothing.

Results

Tasks it has never seen

Nine tasks held out from training entirely — new label sets, new domains, new phrasings. Accuracy, beside the score of always answering the most common label:

task type DECISIONx + calibration (32 examples) most-common label
DBpedia — 14 article categories choice 0.930 0.931 0.102
Rotten Tomatoes — critic recommends? null 0.922 0.918 0.505
Counterfactual statements null 0.900 0.896 0.888
Toxicity null 0.867 0.879 0.902
Financial news — bearish/bullish/neutral choice 0.832 0.846 0.687
TREC — question type (6) choice 0.818 0.818 0.276
MASSIVE — assistant intents (60) choice 0.735 0.750 0.072
Enron — spam? null 0.615 0.905 0.507
SST-5 — review stars (1–5) score 0.507 0.515 0.297
mean 0.792 0.829 0.470

Expected calibration error on these tasks: 0.089 out of the box, 0.053 after calibrating with 32 examples, 0.046 with 64.

labelled examples per task 0 8 16 32 64
mean accuracy (9 unseen tasks) 0.792 0.817 0.824 0.829 0.835

Each calibration figure is the mean of 20 random draws of the examples, scored on the rest of the task.

Tasks it was trained on

On the test slices of its training tasks it averages 0.84 accuracy with expected calibration error around 0.03 — for example banking intents (77 options) 0.82, CLINC intents (151) 0.92, NLI 0.84–0.92, FEVER claim checking 0.90, paraphrase detection 0.95.

Full per-task numbers are in results/eval_report.json; the calibration study is in results/calibration_heldout.json.

How it works

The state, the question and the numbered options are laid out in the backbone's chat format. The backbone reads them once. At the last token of each option, a small MLP turns the hidden state into one score; a softmax over the options gives the probabilities. There is no language-model head and no decoding, so a decision costs one forward pass however many options there are.

A score question is answered by scoring the points of the scale; a yes/no question by scoring "yes" and "no". One head serves all three types. After training, a temperature per question type is fitted on validation data so the probabilities mean what they say.

Backbone Qwen2.5-1.5B-Instruct, adapted with LoRA (rank 16, all attention and MLP projections)
Trainable parameters 18.5M (LoRA) + 0.6M (scoring head)
Training data 269k decisions from 24 public datasets: intents, ticket routing, topics, emotion and sentiment, natural-language inference, claim checking, paraphrase, multiple-choice reasoning, review and rating scores, semantic similarity, readability. Several phrasings per task; label-set tasks also posed as yes/no statements at varied base rates
Training one epoch on a single NVIDIA L4, about 5.6 hours
State window 768 tokens; longer text is truncated

Limitations

  • Thresholds on new yes/no tasks. Ranking transfers better than base rates. Calibrate any yes/no decision that matters, with 32 or more labelled examples.
  • Scores are the weakest type. On a new 1–5 scale it is about half right exactly and about half a point off on average.
  • Small model. At 1.5B parameters, decisions that need multi-step reasoning or specialist knowledge will be worse than a large LLM's.
  • English only, and measured on public benchmark data. It has not been evaluated on any real production decision stream.
  • Long inputs are truncated at 768 tokens.
  • Options read in order. Each option sees those before it. Training shuffled option order to remove position bias, but lists of 100+ options were never trained on in full.

License and intended use

Research and evaluation only. The backbone is Apache-2.0, but some of the training data carries non-commercial terms (Yelp, QQP, ANLI, SICK, CommonLit, Amazon reviews) and some has unclear terms. These weights are therefore not licensed for commercial use. A version trained only on permissively licensed data is planned.

Files

adapter/ LoRA weights (safetensors) and config
head.safetensors the scoring head
tokenizer/ tokenizer and chat template
decisionx.json model config and per-type temperatures
decisionx/ the inference package (DecisionModel, Decision, TaskCalibration)
results/ evaluation report and calibration study
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thinkingdbx/DECISIONx-1.5B

Adapter
(1507)
this model

Datasets used to train thinkingdbx/DECISIONx-1.5B