Instructions to use convaiinnovations/laya with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use convaiinnovations/laya with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="convaiinnovations/laya")# pip install -U transformers accelerate # Load model directly from transformers import LayaTypedDecisions model = LayaTypedDecisions.from_pretrained("convaiinnovations/laya", device_map="auto") - Laya
How to use convaiinnovations/laya with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
choice:11+ temperature (0.1006) saturates confidence: answers return at 1.00, wrong ones included
Thanks for releasing this with the weights and the eval numbers, the Jev-compatible system_one shape made it very easy to try.
Reporting something that looks unintended in the shipped calibration, because it silently removes the confidence signal for any choice with more than ten options.
What happens
temperature_by_options in rl_agent_config.json fits one temperature per option-count bucket:
{'choice:2': 1.906, 'choice:3-5': 1.760, 'choice:6-10': 1.0000, 'choice:11+': 0.1006,
'score:3-5': 1.251, 'noul:2': 1.983}
choice:11+ is 0.1006. Since the code does z = logits / temperature, that multiplies the logits by about ten and saturates the softmax, so essentially every answer comes back at confidence 1.00.
Minimal reproduction
Same model, same eight support tickets, same instructions. The only thing that changes is how many options the question offers, which moves it across the bucket boundary at ten.
| options | bucket (temperature) | accuracy | mean confidence | min confidence |
|---|---|---|---|---|
| 10 | choice:6-10 (1.0000) | 7/8 | 0.715 | 0.151 |
| 11 | choice:11+ (0.1006) | 6/8 | 0.968 | 0.744 |
| 20 | choice:11+ (0.1006) | 5/8 | 1.000 | 1.000 |
Adding a single option takes mean confidence from 0.715 to 0.968. At twenty options every answer reports 1.000, including the three that are wrong.
It is the temperature, not the option count
Re-scoring the same k=20 forward passes at temperature 1.0 instead of 0.1006 gives mean confidence around 0.80 rather than 1.000, with the same accuracy. So this is not the entropy normalisation in confidence_from_probs reacting to a larger k, and it is not the model: temperature only rescales confidence, and the chosen option never changes.
On a separate 20-option routing task of my own, rescoring at 1.0 also restored a usable gap between confidence-when-right and confidence-when-wrong, where the shipped value left almost none.
Why it seems worth fixing
The main reason to reach for a calibrated decision model is to gate on the probability and ask a human when the model is unsure. Above ten options the returned confidence carries no information, so a gate built on it passes everything through, and the errors arrive looking maximally certain. That is the one failure mode calibration is supposed to prevent.
Worth double-checking how that bucket was fit, or capping it. Happy to share the reproduction script if useful.
Repro script used:
from rl_agent_api import RLAgent
agent = RLAgent('.', device='cpu')
tickets = [("my card was charged twice for one order", "billing"),
("the package never arrived", "shipping")] # ... 8 total
for options in (depts[:10], depts[:11], depts):
for text, want in tickets:
out = agent.system_one(text, {"q": {"type": "choice",
"instructions": "Which department should handle this ticket?",
"criteria": {d: None for d in options}}})
print(len(options), out["answers"]["q"]["confidence"])
This is still the right report, and I want to add what has changed since β partly because it
affects the recommendation at the end of it.
Your measurement was on the raw value, and it was correct: the loader now clamps temperatures to[0.5, 5], and that landed two days after you posted (2026-09-21). So 0.1006 no longer reaches
the softmax; it becomes 0.5 and the load prints a warning naming the affected bucket. Measured
on the English checkpoint at k=20, one state:
choice:11+ |
top-1 | entropy | options >1% |
|---|---|---|---|
| 0.1006 (as shipped) | 1.0000 | 0.0000 | 1 / 20 |
| 0.5 (as served, after the clamp) | 0.9514 | 0.2909 | 3 / 20 |
| 1.0 (neutral) | 0.5664 | 1.8509 | 16 / 20 |
So your point mass is confirmed exactly, and the clamp is a real improvement β but it is not a
calibration, and that is the part worth knowing if you gate on it. At 20 options the clamped
distribution still puts ~95% on one label, and right and wrong answers overlap almost completely:
mean answer-confidence 0.915 when right against 0.906 when wrong.
I measured what that does to a threshold. Twelve cases whose content makes exactly one option
correct, at two option counts:
| gate | 3 options | 20 options |
|---|---|---|
>= 0.85 |
3/12 selected, all correct β precision 1.000 | 9/12 selected, 5 correct β 0.556 |
At 20 options a gate at 0.85 selects 75% of the cases at 55.6% precision, below the model's own
58.3% accuracy on the same set β so it is not merely uncalibrated at high k, it is inverted
there. Filed as NandhaKishorM/laya#394, and
the README no longer presents a confidence threshold as permission to act without measuring it on
your own data.
One correction to the framing, in the model's favour: choice:11+ is one temperature for
every count from 11 up, and it is measurably worse at the low end of its own bucket than at the
top. At k=13 the clamped value still leaves top-1 at 0.9880 with a single option above 1%. So the
bucket is not uniformly saturated. If the intent is that confidence be usable at high option
counts, that bucket needs refitting rather than clamping, and that is a checkpoint change rather
than a code one.
Your k=10 β k=11 step across the bucket boundary is still the cleanest demonstration of this I
have seen, and it is the reason I looked at the 13-vs-20 split at all.