kambo-v1-decide

kambo-v1-decide is Kambo-v1 fine-tuned to make typed decisions: you give it a state (a message, a document, a log line) and a set of questions about it, and it returns a probability for every allowed answer. It doesn't generate text.

Each question has a type:

Type You give You get back
Choice a question and a list of options (up to 255 per question) probability per option, and the most likely option
Noul a yes/no question P(True)
Score a question and ordered levels probability per level, and the expected level

Every answer is read off one forward pass. The state is encoded once and shared by all questions in the call. Each option gets a single-token label, so all options are scored from one logit vector.

Usage

import sys
from huggingface_hub import snapshot_download

path = snapshot_download("VikramPal/kambo-v1-decide")
sys.path.insert(0, path)
from kambo_decide import Decider, Choice, Noul, Score

d = Decider(path)
out = d.decide(
    "Customer: I was charged twice for my card this month and nobody answers. "
    "Fix it or I'm closing my account.",
    [
        Choice("Which team should handle this?", ["billing", "technical support", "sales"]),
        Noul("Does the message mention a payment card?"),
        Score("How angry is the customer?", ["calm", "annoyed", "furious"]),
    ],
)
for r in out:
    print(r["question"], "->", r["answer"], r["probs"])

Each result also carries confidence (top probability) and label_mass: the share of the raw next-token distribution that landed on a valid label. A low label_mass means the input didn't fit the question well, and it's a usable out-of-distribution signal on its own.

The Decider scores each question in the given option order and in reversed order, then averages the two. This cancels the position bias that otherwise favours option "A". To skip that and halve the cost, pass permute=False.

Tested with transformers==4.53.2 and torch 2.x. On a single A100, four questions about one state take about 150 ms in one call. A 77-option question takes about the same.

Adding a "none" option

The model was trained with a "none of these" option that is sometimes the correct answer. Add one to any Choice when the input may not fit any of your options:

labels = ["card arrival", "lost or stolen card", "exchange rate", "none of these"]
d.decide("Customer message: who won the 1998 world cup?",
         [Choice("Which banking request is the customer making?", labels)])

Optional calibration

decide(..., calib={"choice": {"T": t, "b": offsets}}) applies a temperature and per-option offsets fitted on your own labelled data. kambo_decide.fit_temperature and kambo_decide.fit_bias fit them from the output of Decider.raw_logprobs. The model is usable without this: the raw probabilities in the tables below are uncalibrated.

Results

All numbers come from 400 test items per task, run on one A100 in bf16. "Calibrated" means a temperature plus per-option offsets fitted on a separate 200-item dev split of the same task. The baseline is the unmodified Kambo-v1, scored with the same prompts and code.

Task How it relates to training Majority class Kambo-v1 kambo-v1-decide kambo-v1-decide, calibrated
IMDB sentiment (2-way) task never seen in training 52.0% 68.5% 86.8% 86.0% (ECE 0.021)
BoolQ (yes/no over a passage) train split used; test is the held-out validation split 59.0% 59.0% 67.8% 70.5%
Banking77 (77 intents) train split used; test is the held-out test split 2.7% 2.5% 74.5% 75.3%

Out-of-domain inputs: we asked the 77-way Banking77 question about 200 trivia questions (SQuAD validation) and about 200 real banking messages.

Kambo-v1 kambo-v1-decide
Off-topic inputs rejected via a "none of these" option 16.0% 99.5%
Banking messages answered normally with the same option present 73.5% 86.5%
Mean confidence on banking messages / off-topic inputs (no "none" option) 0.04 / 0.04 0.79 / 0.11
Off-topic inputs with confidence ≥ 0.9 0% 0.5%
AUROC, banking vs off-topic, using confidence 0.39 0.97
AUROC, banking vs off-topic, using label_mass 0.96 0.98

Without a "none" option, the model doesn't force a confident answer onto off-topic input. Its confidence drops to 0.11 on average, so a confidence threshold separates in-domain from off-topic inputs.

Training

This is a full fine-tune of every weight except the MoE routers, which are frozen so the expert balance learned in pretraining is kept. The architecture is unchanged.

  • Objective: cross-entropy over the option labels only, plus 0.1 × cross-entropy of the correct label over the full vocabulary. The first term is the expected log-score of the answer distribution, so it rewards calibrated probabilities rather than just a correct top pick. The second keeps probability mass on valid labels.
  • Schedule: 2,500 steps, batch size 32 (80,000 examples), AdamW (β = 0.9, 0.95), learning rate 1e-5 with 100 warmup steps and cosine decay, gradient clip 1.0, bf16 autocast with fp32 master weights. About 2.1 hours on one A100 40 GB.
  • Data: up to 8,000 examples from each source, train splits only. The examples are recast as Choice, Noul or Score questions. Options are randomly subsampled and shuffled. With probability 0.15 the true answer is removed and "none of these" becomes correct, and "none of these" is often added as a distractor.
Source Used as
MNLI, ANLI (r1–r3) entailment as 3-way choice or yes/no
BoolQ yes/no over a passage
Banking77, CLINC150 (incl. out-of-scope → "none") intent choice
AG News, DBpedia-14, TREC topic / question-type choice
PubMedQA (artificial) yes/no over an abstract
Civil Comments toxicity as yes/no and as a 3-level score
SQuAD (train) questions off-topic inputs for the "none" answer

Sentiment data was deliberately left out, so IMDB measures transfer to an unseen task. The training datasets keep their own licenses. Notably, ANLI is CC BY-NC 4.0.

Limitations

  • English only. The questions and states in training were all English.
  • It decides from what is stated. It is weaker on implicit cues: for "Fix it or I'm closing my account", P(threatening to leave) is 0.46.
  • Each task was measured on 400 items with one training seed, so differences of a couple of points are within noise.
  • Calibrating on a 200-item dev set doesn't always help: on BoolQ it raised accuracy but worsened ECE (0.054 raw → 0.086). Calibrate on more data, or use the raw probabilities.
  • Don't use it as the only signal for high-stakes decisions.

License

Apache-2.0, the same as the base model.

Downloads last month
32
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VikramPal/kambo-v1-decide

Finetuned
(2)
this model