Instructions to use VikramPal/kambo-v1-decide with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use VikramPal/kambo-v1-decide with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="VikramPal/kambo-v1-decide", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("VikramPal/kambo-v1-decide", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
kambo-v1-decide
kambo-v1-decide is Kambo-v1 fine-tuned to make typed decisions: you give it a state (a message, a document, a log line) and a set of questions about it, and it returns a probability for every allowed answer. It doesn't generate text.
Each question has a type:
| Type | You give | You get back |
|---|---|---|
Choice |
a question and a list of options (up to 255 per question) | probability per option, and the most likely option |
Noul |
a yes/no question | P(True) |
Score |
a question and ordered levels | probability per level, and the expected level |
Every answer is read off one forward pass. The state is encoded once and shared by all questions in the call. Each option gets a single-token label, so all options are scored from one logit vector.
Usage
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("VikramPal/kambo-v1-decide")
sys.path.insert(0, path)
from kambo_decide import Decider, Choice, Noul, Score
d = Decider(path)
out = d.decide(
"Customer: I was charged twice for my card this month and nobody answers. "
"Fix it or I'm closing my account.",
[
Choice("Which team should handle this?", ["billing", "technical support", "sales"]),
Noul("Does the message mention a payment card?"),
Score("How angry is the customer?", ["calm", "annoyed", "furious"]),
],
)
for r in out:
print(r["question"], "->", r["answer"], r["probs"])
Each result also carries confidence (top probability) and label_mass: the share of the raw next-token distribution that landed on a valid label. A low label_mass means the input didn't fit the question well, and it's a usable out-of-distribution signal on its own.
The Decider scores each question in the given option order and in reversed order, then averages the two. This cancels the position bias that otherwise favours option "A". To skip that and halve the cost, pass permute=False.
Tested with transformers==4.53.2 and torch 2.x. On a single A100, four questions about one state take about 150 ms in one call. A 77-option question takes about the same.
Adding a "none" option
The model was trained with a "none of these" option that is sometimes the correct answer. Add one to any Choice when the input may not fit any of your options:
labels = ["card arrival", "lost or stolen card", "exchange rate", "none of these"]
d.decide("Customer message: who won the 1998 world cup?",
[Choice("Which banking request is the customer making?", labels)])
Optional calibration
decide(..., calib={"choice": {"T": t, "b": offsets}}) applies a temperature and per-option offsets fitted on your own labelled data. kambo_decide.fit_temperature and kambo_decide.fit_bias fit them from the output of Decider.raw_logprobs. The model is usable without this: the raw probabilities in the tables below are uncalibrated.
Results
All numbers come from 400 test items per task, run on one A100 in bf16. "Calibrated" means a temperature plus per-option offsets fitted on a separate 200-item dev split of the same task. The baseline is the unmodified Kambo-v1, scored with the same prompts and code.
| Task | How it relates to training | Majority class | Kambo-v1 | kambo-v1-decide | kambo-v1-decide, calibrated |
|---|---|---|---|---|---|
| IMDB sentiment (2-way) | task never seen in training | 52.0% | 68.5% | 86.8% | 86.0% (ECE 0.021) |
| BoolQ (yes/no over a passage) | train split used; test is the held-out validation split | 59.0% | 59.0% | 67.8% | 70.5% |
| Banking77 (77 intents) | train split used; test is the held-out test split | 2.7% | 2.5% | 74.5% | 75.3% |
Out-of-domain inputs: we asked the 77-way Banking77 question about 200 trivia questions (SQuAD validation) and about 200 real banking messages.
| Kambo-v1 | kambo-v1-decide | |
|---|---|---|
| Off-topic inputs rejected via a "none of these" option | 16.0% | 99.5% |
| Banking messages answered normally with the same option present | 73.5% | 86.5% |
| Mean confidence on banking messages / off-topic inputs (no "none" option) | 0.04 / 0.04 | 0.79 / 0.11 |
| Off-topic inputs with confidence ≥ 0.9 | 0% | 0.5% |
| AUROC, banking vs off-topic, using confidence | 0.39 | 0.97 |
AUROC, banking vs off-topic, using label_mass |
0.96 | 0.98 |
Without a "none" option, the model doesn't force a confident answer onto off-topic input. Its confidence drops to 0.11 on average, so a confidence threshold separates in-domain from off-topic inputs.
Training
This is a full fine-tune of every weight except the MoE routers, which are frozen so the expert balance learned in pretraining is kept. The architecture is unchanged.
- Objective: cross-entropy over the option labels only, plus 0.1 × cross-entropy of the correct label over the full vocabulary. The first term is the expected log-score of the answer distribution, so it rewards calibrated probabilities rather than just a correct top pick. The second keeps probability mass on valid labels.
- Schedule: 2,500 steps, batch size 32 (80,000 examples), AdamW (β = 0.9, 0.95), learning rate 1e-5 with 100 warmup steps and cosine decay, gradient clip 1.0, bf16 autocast with fp32 master weights. About 2.1 hours on one A100 40 GB.
- Data: up to 8,000 examples from each source, train splits only. The examples are recast as
Choice,NoulorScorequestions. Options are randomly subsampled and shuffled. With probability 0.15 the true answer is removed and "none of these" becomes correct, and "none of these" is often added as a distractor.
| Source | Used as |
|---|---|
| MNLI, ANLI (r1–r3) | entailment as 3-way choice or yes/no |
| BoolQ | yes/no over a passage |
| Banking77, CLINC150 (incl. out-of-scope → "none") | intent choice |
| AG News, DBpedia-14, TREC | topic / question-type choice |
| PubMedQA (artificial) | yes/no over an abstract |
| Civil Comments | toxicity as yes/no and as a 3-level score |
| SQuAD (train) questions | off-topic inputs for the "none" answer |
Sentiment data was deliberately left out, so IMDB measures transfer to an unseen task. The training datasets keep their own licenses. Notably, ANLI is CC BY-NC 4.0.
Limitations
- English only. The questions and states in training were all English.
- It decides from what is stated. It is weaker on implicit cues: for "Fix it or I'm closing my account", P(threatening to leave) is 0.46.
- Each task was measured on 400 items with one training seed, so differences of a couple of points are within noise.
- Calibrating on a 200-item dev set doesn't always help: on BoolQ it raised accuracy but worsened ECE (0.054 raw → 0.086). Calibrate on more data, or use the raw probabilities.
- Don't use it as the only signal for high-stakes decisions.
License
Apache-2.0, the same as the base model.
- Downloads last month
- 32
Model tree for VikramPal/kambo-v1-decide
Base model
VikramPal/kambo-v1