Instructions to use thinkingdbx/DECISIONx-1.5B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thinkingdbx/DECISIONx-1.5B with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
DECISIONx
A decision model. Ask a typed question about a piece of text; get calibrated probabilities back. It never generates text.
DECISIONx answers the small decisions inside software — which queue gets this
ticket, is this email spam, how urgent is this incident — in a single forward
pass, with a probability you can put in an if statement.
| Question type | You give it | You get back |
|---|---|---|
choice |
any list of options | the best option, and a probability for every option |
score |
a scale (e.g. 1–5, with optional labels) | the expected score, its spread, and the full distribution |
null |
a yes/no statement | the probability that it is true |
Options can be anything, in any order. They do not need to have been seen in training.
Quick start
pip install torch transformers peft safetensors huggingface_hub
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("thinkingdbx/DECISIONx-1.5B")
sys.path.insert(0, path) # the small inference package ships with the weights
from decisionx.model import DecisionModel
m = DecisionModel.load(path, device="cuda") # or "cpu"; the base model downloads on first use
m.decide("choice", "Which team should handle this ticket?",
state="I was charged twice for my order last week and I want the extra payment back.",
options=["billing", "shipping", "technical support", "account security"])
What it answers
Real outputs from this checkpoint (CPU, no calibration):
| state | question | answer | |
|---|---|---|---|
| choice | "I was charged twice for my order last week…" | Which team should handle this ticket? (billing / shipping / technical support / account security) | billing — 0.92 (technical support 0.08) |
| choice | "Can you move my 3pm meeting with Priya to tomorrow morning?" | What kind of request is this? (calendar / email / travel / weather / music) | calendar — 0.9996 |
| null | "The package arrived damaged and I would like my money back please." | The customer is asking for a refund. | 0.94 true |
| null | "Subject: Q3 planning — Hi team, attaching the draft roadmap…" | This email is spam. | 0.0006 |
| null | "Subject: You've been selected! Congratulations, you have won a free cruise…" | This email is spam. | 0.149 ⚠️ |
| score | "Production database is down and no customer can log in." | How urgent is this? (1 can wait … 5 drop everything) | 3.9 ± 1.0 (mode 4) |
| score | "Could you update the footer copyright year…when you get a chance?" | How urgent is this? | 2.3 ± 1.0 (mode 2) |
The spam row is the model's main weakness, shown on purpose: it rates the cruise email 230× more spam-like than the planning email, so its ranking is right, but on a task it has never seen it can put the threshold in the wrong place. That is what task calibration is for.
Calibrating a new task
Give it 20–60 labelled examples of your decision; it fits a temperature and a per-option bias for that task, and the answers you get afterwards are both better placed and honestly confident.
from decisionx.schema import Decision
from decisionx.calibration import TaskCalibration
examples = [Decision("null", email, "This email is spam.", label=is_spam)
for email, is_spam in my_labelled_emails[:32]]
cal = m.calibrate(examples) # strength chosen by cross-validation on your examples
cal.save("spam.calibration.json")
m.decide("null", "This email is spam.", state=new_email,
calibration=TaskCalibration.load("spam.calibration.json"))
Biases are keyed by option text, so option order does not matter. On a task that is already well calibrated, cross-validation will usually choose to change nothing.
Results
Tasks it has never seen
Nine tasks held out from training entirely — new label sets, new domains, new phrasings. Accuracy, beside the score of always answering the most common label:
| task | type | DECISIONx | + calibration (32 examples) | most-common label |
|---|---|---|---|---|
| DBpedia — 14 article categories | choice | 0.930 | 0.931 | 0.102 |
| Rotten Tomatoes — critic recommends? | null | 0.922 | 0.918 | 0.505 |
| Counterfactual statements | null | 0.900 | 0.896 | 0.888 |
| Toxicity | null | 0.867 | 0.879 | 0.902 |
| Financial news — bearish/bullish/neutral | choice | 0.832 | 0.846 | 0.687 |
| TREC — question type (6) | choice | 0.818 | 0.818 | 0.276 |
| MASSIVE — assistant intents (60) | choice | 0.735 | 0.750 | 0.072 |
| Enron — spam? | null | 0.615 | 0.905 | 0.507 |
| SST-5 — review stars (1–5) | score | 0.507 | 0.515 | 0.297 |
| mean | 0.792 | 0.829 | 0.470 |
Expected calibration error on these tasks: 0.089 out of the box, 0.053 after calibrating with 32 examples, 0.046 with 64.
| labelled examples per task | 0 | 8 | 16 | 32 | 64 |
|---|---|---|---|---|---|
| mean accuracy (9 unseen tasks) | 0.792 | 0.817 | 0.824 | 0.829 | 0.835 |
Each calibration figure is the mean of 20 random draws of the examples, scored on the rest of the task.
Tasks it was trained on
On the test slices of its training tasks it averages 0.84 accuracy with expected calibration error around 0.03 — for example banking intents (77 options) 0.82, CLINC intents (151) 0.92, NLI 0.84–0.92, FEVER claim checking 0.90, paraphrase detection 0.95.
Full per-task numbers are in results/eval_report.json; the calibration study
is in results/calibration_heldout.json.
How it works
The state, the question and the numbered options are laid out in the backbone's chat format. The backbone reads them once. At the last token of each option, a small MLP turns the hidden state into one score; a softmax over the options gives the probabilities. There is no language-model head and no decoding, so a decision costs one forward pass however many options there are.
A score question is answered by scoring the points of the scale; a yes/no question by scoring "yes" and "no". One head serves all three types. After training, a temperature per question type is fitted on validation data so the probabilities mean what they say.
| Backbone | Qwen2.5-1.5B-Instruct, adapted with LoRA (rank 16, all attention and MLP projections) |
| Trainable parameters | 18.5M (LoRA) + 0.6M (scoring head) |
| Training data | 269k decisions from 24 public datasets: intents, ticket routing, topics, emotion and sentiment, natural-language inference, claim checking, paraphrase, multiple-choice reasoning, review and rating scores, semantic similarity, readability. Several phrasings per task; label-set tasks also posed as yes/no statements at varied base rates |
| Training | one epoch on a single NVIDIA L4, about 5.6 hours |
| State window | 768 tokens; longer text is truncated |
Limitations
- Thresholds on new yes/no tasks. Ranking transfers better than base rates. Calibrate any yes/no decision that matters, with 32 or more labelled examples.
- Scores are the weakest type. On a new 1–5 scale it is about half right exactly and about half a point off on average.
- Small model. At 1.5B parameters, decisions that need multi-step reasoning or specialist knowledge will be worse than a large LLM's.
- English only, and measured on public benchmark data. It has not been evaluated on any real production decision stream.
- Long inputs are truncated at 768 tokens.
- Options read in order. Each option sees those before it. Training shuffled option order to remove position bias, but lists of 100+ options were never trained on in full.
License and intended use
Research and evaluation only. The backbone is Apache-2.0, but some of the training data carries non-commercial terms (Yelp, QQP, ANLI, SICK, CommonLit, Amazon reviews) and some has unclear terms. These weights are therefore not licensed for commercial use. A version trained only on permissively licensed data is planned.
Files
adapter/ |
LoRA weights (safetensors) and config |
head.safetensors |
the scoring head |
tokenizer/ |
tokenizer and chat template |
decisionx.json |
model config and per-type temperatures |
decisionx/ |
the inference package (DecisionModel, Decision, TaskCalibration) |
results/ |
evaluation report and calibration study |
- Downloads last month
- -