Onebox 1.1 (9B)

Onebox makes typed decisions about a state. You send a state (text or JSON) and one or more questions, each with a fixed set of allowed options. You get back a probability for every option of every question. The model never writes free text, so there is nothing to parse and nothing to repair.

The name comes from Newcomb's problem, where the decision is whether to take one box or two. Onebox takes its decision in one pass.

This is a research release from the newcombs lab. We use it to study how well a local model makes typed decisions and how often it falls for misleading words in the input. This is our measure of whether a model has learned to read a text closely and understand it, rather than to just pick the option whose word fits best.

What changed since Onebox 1

Onebox 1 fell for words that name an option, most of all when a page or form named after a topic was broken. For 1.1 we added contrast pairs to the training data. Each pair has two short texts with the same word. In one the word is a lure, in the other it really names the right option.

We tried several adapter sizes and kept the one that worked best. The base model, the inference code and the request format are the same.

What is in this release

  • onebox.safetensors: weights we trained for Qwen3.5-9B, low rank adapters and a small scoring head. Only the text part of the base model is used.
  • onebox.py: the inference code. One file, no training code.
  • onebox_config.json: base model, sizes and the calibration temperature.

Question types

type what you send what you get
choice named options, optionally with a description each a probability per option
noul a yes or no question the probability of yes
score ordered levels a probability per level and the expected level

The order in which you list the options does not change the answer. Every option is read on its own against the same state, and the options are compared only at the end.

How to run it

pip install torch "transformers>=5.10" safetensors huggingface_hub
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("newcombs/onebox-1.1-9b")
sys.path.insert(0, path)
from onebox import load

model = load(path)  # fetches Qwen/Qwen3.5-9B on first use, about 19 GB
print(model.decide(
    {"ticket": "Server down since 3am, customers cannot log in"},
    {
        "team": {
            "type": "choice",
            "instructions": "Which team handles this?",
            "criteria": {"billing": "payments, invoices", "technical": "outages, bugs, login", "sales": "new contracts"},
        },
        "urgent": {"type": "noul", "instructions": "Does this need action today?"},
        "severity": {"type": "score", "instructions": "How severe is this?", "criteria": ["minor", "major", "critical"]},
    },
))

The request and the answer have the same shape as TypeSafe's System One API (state, questions with type, instructions and criteria). load picks CUDA, Apple MPS or the CPU on its own. In bf16 the model needs about 18 GB of memory.

Output (rounded to four places):

{'answers': {'team': {'type': 'choice',
                      'choice': 'technical',
                      'confidence': 0.9641,
                      'probabilities': {'billing': 0.0185, 'technical': 0.9641, 'sales': 0.0174}},
             'urgent': {'type': 'noul', 'noul': 0.9705},
             'severity': {'type': 'score',
                          'score': 1.5098,
                          'confidence': 0.5153,
                          'legend': {'0': 'minor', '1': 'major', '2': 'critical'},
                          'probabilities': {'0': 0.0055, '1': 0.4791, '2': 0.5153}}}}

Results

All numbers are measured by us with one harness. A decision counts as correct when the most likely option matches the label.

typed-decisions

Public benchmark LocalLLaMA/typed-decisions, official test split: 400 cases with 2000 decisions from four workflows (agent traces, customer service, invoices, security incidents).

model trained on these four workflows? accuracy
always the most common answer n/a 46.1 %
Laya 421M no 36.2 %
openjev 4B v5 no, according to its manifest 63.8 %
Clef-Flash (9B) not as far as its card says 70.3 %
Julia-1 unknown 72.6 %
TypeSafe Jev (closed, published number) unknown 72.7 %
Laya 421M, fine-tuned yes 76.6 %
Onebox 1 (9B) yes 79.3 %
Onebox 1.1 (9B) yes 80.4 %

Onebox was trained on the training split of this benchmark, so this is the number for workflows it has seen. On a workflow it has never seen, expect clearly less.

By question type on the test split, Onebox 1.1 reaches 78.0 % on choice (Onebox 1: 75.3 %), 85.7 % on yes or no questions (86.2 %) and 78.1 % on score (77.1 %). Its expected calibration error before temperature scaling is 0.133 (Onebox 1: 0.150). On the validation split it reaches 82.2 % (80.5 %).

The step from Onebox 1 is about one point and close to what chance alone can move on 2000 decisions. We report it mainly because it did not get worse.

Misleading words

A lure is a word in the input that points toward the wrong option. A customer writes "I can't open the billing page." The page does not load, so this belongs to the technical team, but the word "billing" pulls toward billing.

We built two sets of 200 pairs for this, written by two different language models and checked by us. In each pair, one case has such a lure and the other uses the same word where it really points to the right option. The table shows how often both cases of a pair are right, and in brackets how often the model picked the lure. With descriptions, every option also came with a short description.

model set A set A, with descriptions set B set B, with descriptions
Onebox 1 (9B) 85.0 % (9.5 %) 90.0 % (7.5 %) 60.5 % (25.0 %) 75.5 % (19.5 %)
Onebox 1.1 (9B) 91.0 % (6.0 %) 92.5 % (7.0 %) 68.5 % (21.0 %) 82.0 % (15.0 %)
Clef-Flash (9B) 90.5 % (6.0 %) 94.5 % (5.5 %) 69.5 % (25.5 %) 83.0 % (14.5 %)

These sets are not neutral ground for us. Our contrast pairs cover some of the kinds of lures these sets test (never the test pairs themselves), and Clef-Flash was trained on none of them. So here are both parts separately (both sets, names only):

model pairs right, kinds our training covers pairs right, other kinds lure picked, other kinds
Onebox 1 (9B) 68.3 % 77.6 % 12.5 %
Onebox 1.1 (9B) 77.4 % 82.3 % 12.0 %
Clef-Flash (9B) 75.5 % 84.9 % 13.0 %

On the kinds it was trained on, Onebox 1.1 is now ahead. The clearest case is a broken page named after a topic: Onebox 1.1 picks the lure in 28 % of these cases, Onebox 1 in 42 % and Clef-Flash in 36 %. On the other kinds Clef-Flash is still ahead, and Onebox 1.1 falls for the lure as often as Onebox 1. The model learned the kinds of lures it saw, not to read more carefully in general.

On an older, smaller lure set it falls for the lure about half as often as Onebox 1.

The billing page is not solved either. Onebox 1.1 still picks billing for "I can't open the billing page" (91.7 %). When the text names an error ("shows a 500 error", "stuck on a spinner", "blank white screen") it picks technical in all four of our test sentences.

One more finding, from our own analysis: when we replace the lure word with a neutral token, the model picks the right option in about two thirds of the cases where it had fallen for the lure. The rest of the text was enough. The single word overrules what the model understood.

Other tasks

150 items per task, fixed sample, none of them seen in training.

task options Onebox 1 Onebox 1.1
AG News 4 90.0 % 87.3 %
emotion 6 54.7 % 54.7 %
XNLI (English) 3 83.3 % 86.0 %
Banking77 77 55.3 % 57.3 %
MASSIVE 59 70.0 % 75.3 %

Speed

With onebox.py in bf16 on a MacBook Pro with M5 Max (Apple MPS), a question takes about 0.7 seconds (400 questions from the typed-decisions test). On the Mac, PyTorch only has slow reference kernels for the linear attention layers of Qwen3.5. Our own MLX server for Apple Silicon is faster, but it is not part of this release. We have not measured onebox.py on an NVIDIA GPU.

Questions with many options stay cheap because the shared part of the input is read once.

How this model was chosen

We trained several variants of 1.1 and fixed a selection rule before measuring them. That rule picked a different variant, by a narrow margin. We release this one because it is ahead on the typed-decisions test and validation split, falls for fewer of the older lure cases and is better calibrated. Because we chose among several variants on the test sets, all numbers of the chosen model are slightly optimistic.

Training data

  • Public sets with their own labels: SNLI, ANLI, MultiNLI, BoolQ (training splits only).
  • Decision cases we built ourselves, including the contrast pairs against lures. None of them overlap with the lure test sets.
  • The training split of typed-decisions, never the test split.

Training ran in one stage on a single GPU. The numbers on this card were measured with MLX on a Mac. We checked onebox.py against them: on the first 400 decisions of the typed-decisions test it picks the same option in 399.

Limits

  • Words that name an option still pull Onebox toward it, most of all for kinds of lures that were not in its training data and for a broken page described without an error word.
  • On workflows it has not seen, expect less than the 80.4 % above.
  • Mostly English. We have not measured other languages.
  • It was trained on inputs of up to 768 tokens (about 550 English words). onebox.py reads up to 32,000 tokens and cuts longer states from the beginning. How well it decides on longer inputs is not measured yet.
  • It reads text only, no images.
  • It is a research model. Do not use it for decisions about people without a human checking the result.

License

The weights and onebox.py are released under the Apache License 2.0. The base model Qwen3.5-9B is licensed under Apache 2.0 by the Qwen team.

Citation

@misc{newcombs2026onebox11,
  title  = {Onebox 1.1: typed decisions in one pass},
  author = {{newcombs}},
  year   = {2026},
  url    = {https://huggingface.co/newcombs/onebox-1.1-9b}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for newcombs/onebox-1.1-9b

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(980)
this model