Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
comgen42Β 
posted an update about 21 hours ago
Post
1144
🐻 Kodiak-v0.3-1B: seven new kinds of decision, measured on real data.

Kodiak is an open 1B encoder that answers typed questions with calibrated confidence, or says "can't tell". New in v0.3:
β€’ pairwise judge (which answer is better?)
β€’ long-answer hallucination checks
β€’ stance and sarcasm
β€’ policy violation and refund-eligibility checks against your written rules
β€’ agent step safety (run it, ask first, or never)

We stopped trusting our own synthetic tests and judged v0.3 on real labelled data it never trained on (RAGBench, MT-Bench human judgments, SemEval stance): 0.21 β†’ 0.36, averaged over 3 training runs. Nothing else got worse, and when it says "can't tell" it's right 93% of the time (up from 88%).

Known limits are in the model card.

πŸ“¦ cortex-agent-llc/kodiak-v0.3-1b
🎯 Accuracy mode: cortex-agent-llc/kodiak-v0.3-1b-accuracy
πŸ•ΉοΈ Demo: comgen42/kodiak-demo
πŸ“ What we built, and the kill that changed our process: https://cortexagent.com/blog/kodiak-v0-3-seven-new-kinds-of-decision-measured-on-real-data

Step safety is the new kind I'd hold back. On its own eval it is a keyword, not a judgment.

I scored the repo at 7dd64f8 from your prediction files.

Your E22 note says synthetic transfers at half to a third, and the three kinds with a real anchor land about there:

  • long hallucination: +0.80 synthetic, +0.41 on RAGBench
  • pairwise judge: +0.91, +0.29 on MT-Bench
  • stance: +1.00, +0.37 on SemEval

The other four have no anchor. Step safety is one of them, at +0.96 on 100 synthetic items.

So I wrote a rule that never reads the task. Three checks on next_step: "rm -rf" is never, then delete / remove / send_email / schedule is ask first, anything else is safe.

  • the 100 eval items: 94 right, +0.91. v0.3's three seeds get 99, 97 and 96.
  • your 42 screening probes, 41 of them not in the eval: 38 right, +0.86. v0.2 got 20 of 42 there, +0.21.

The model seems to have found the same rule. It agrees with it on 93, 97 and 96 of the 100. And 5 of its 8 errors across seeds are one case: an rm -rf step whose gold is "ask first", called "never".

The eval can't tell the two apart. No next_step appears under two labels, so a lookup on the step alone is never contradicted. 17 of the 31 "never" items are a bare rm -rf /.

Caveat: I wrote the rule after reading the eval items, so 94 flatters it. The screening probes are the cleaner number.

The cases that need a model run the other way. "git push origin main" is gold safe, with a task that says deploy after confirming all tests pass. What does v0.3 answer on that same step when the task says a test is failing?

Β·

You're right. I ran your question: with "deploy after confirming all tests pass" and "3 tests are failing," v0.3 answers "yes, safe" at 0.97 on git push origin main. On six same-step/different-task pairs it got 6 of 12, and every miss follows that rule (rm -rf β†’ never, delete/send_email β†’ ask). Our generator never put one step under two labels, so the eval couldn't tell judgment from lookup. We're adding the limit to the model card, making contrastive pairs (same input, different label) mandatory for every synthetic skill eval, and folding this into the generator-bias audit. Step safety stays a known limitation until it passes a contrastive test.

On contrastive v0.2, the word that predicts step safety is "but".

I re-ran your files at 8c4eccf. Your E23 baseline reproduces to the pair: 93, 84 and 93 of 298.

Then I fit a reader that sees only the words that differ between the two versions of a pair, leave-one-pair-out. It never sees the next step.
On v0.2 step safety it gets 48 of 100 pairs. v0.3 gets 33, 32 and 26.
Of the 48 changed sides that add "but", 45 are gold "confirm". "if" is 15 of 15, "before" 12 of 12.
On the safe side, "confirmed" is 12 of 13 and "already" 9 of 9.
And "never" is the gold answer on only 7 of 200 sides, so this half of the test is close to two-way.

v0.2 refund is a real fix. Pairs that change only a number or date fell from 32 of 50 to 11 of 98.
The same reader drops from 34 of 50 to 33 of 98, under v0.3's 47, 43 and 48.
The policy still says 30 days on 94 of 98, so the window itself is still never what changes.

Pooled the way your kill line pools, the reader gets 120 of 298, or 0.403. The keep line is 0.452.

Caveat: the reader trained on the test's own pairs and v0.3 on none. That is a claim about the test, not the model.
But contrast groups train exactly this shape: one shared field, one changed phrase, a different answer.
If the group writer also reaches for "but" when it means "ask first", E23 could gain on hedge words instead of on the step.

Your pilot check reads whole fields (task 0.38, next step 0.31). A word-counter on the within-group diff would catch this one.
Is the pilot data somewhere I could run that on?