AI & ML interests

AI Machine Learning and Nvidia

jbrashear 
posted an update 1 day ago
view post
Post
88
Jeb update for the week. Most of what changed came out of a review.

Dipankar went back to the per-question eval records we publish, re-tempered them himself, and found the 4B and 9B were still shipping the score temperature fit to the training target. The 27B had already moved off it.

So on 2026-09-29 we refit the 4B and 9B without retraining. Score questions now use the hard fit, T 0.83 on both. Choice and noul stay where they were, and accuracy doesn't move. Across the 20 evaluation sets, question-weighted ECE goes from 0.0863 to 0.0800 on the 4B and from 0.0796 to 0.0709 on the 9B.

It isn't free. Jevals HelpSteer2 gets worse, 0.048 to 0.074 on the 4B and 0.039 to 0.082 on the 9B.

Then we ran the check he pushed us toward. HelpSteer2 has 418 validation rows that none of our reported eval sets use and that never went near the training pool. Those rows want a score temperature of 1.26 on the 4B and 1.29 on the 9B, not 0.83.

So the problem is how our calibration split drew HelpSteer2. If your traffic looks like HelpSteer2, keep the older train fit (1.20 on the 4B, 1.22 on the 9B, both in temperatures.json) or refit on your own labels. Both model cards say that now. For v3, the HelpSteer2 part of calibration comes from held-out validation data.

The community quants have also gone further than I expected. bartowski's 9B GGUF is at 3,542 downloads, and mradermacher's 4B, 9B and 27B GGUFs are at 600, 860 and 779. That's more than our own repos put together. Thank you both. And thank you Dipankar, that was exactly the kind of review I hoped for when we put the records up.

A question for anyone running Jeb on real traffic: are you using the shipped temperatures as they are, or refitting on your own labels? If you refit, I'd like to know what you ended up with.

The review thread: https://huggingface.co/posts/jbrashear/856060098865438
jbrashear 
posted an update 3 days ago
view post
Post
117
I've been training Jebadiah, small models that don't chat. You give one a typed question about some JSON state and it gives you back a probability for every option, from one forward pass.

There are three sizes, 4B, 9B and 27B, all Apache 2.0, with GGUF and MLX builds. Each repo ships a temperatures.json fitted on held-out data, so the probabilities are calibrated per question type.

New this week: pip install jebadiah-decide, then jeb serve. It runs them through Ollama, LM Studio, llama-server, vLLM or MLX and puts /v1/systemone and /v1/decide on localhost. The setup for each runtime is here: https://github.com/getainode/jebadiah/blob/main/docs/run-locally.md

If you try it on a routing or triage step of your own, I'd like to hear how it does.

frontier-infra/jebadiah-open-system-one-decision-models-6ab80765ddd3fa0b3eba5213
  • 5 replies
·