Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
pollix 
posted an update about 7 hours ago
Post
717
First stuntd model is on the Hub :)

pollix/stuntd-support-triage is three small heads on the Laya encoder that triage a support ticket in one request: category, urgency and needs_human. About 50 MB each, all three answers come back at a p50 of 71ms through the daemon.

On 1,000 tickets they never saw, each head answers on its own when it's sure: category 99.9%, needs_human 92%, urgency 76%. A ticket only skips the big model when all three are sure, that's 72.7% of them, and all three are right on 97.1% of those.

It's the support demo from the repo, so the tickets are generated and the teacher is a rule. The point is to show what a head looks like and how fast it is, then you train the same thing on your own traffic with your own LLM as the teacher.

hf download pollix/stuntd-support-triage --local-dir support-heads


Model: pollix/stuntd-support-triage
Everything in one place: pollix/stuntd-6abe0a33303828e10c72ab41
Code: https://github.com/bladedevoff/stuntd

The 1,000 eval tickets are unseen states, not unseen tickets.

I ran examples/support/generate.py at 51186ae1 with client.py's own draw (seed 10001, training states skipped).

Every label is a function of three things: the template, the impact sentence and the tail sentence. Channel, plan, product, when, amount and number are slot fill no rule reads.

997 of the 1,000 eval tickets have a template + impact + tail that is already in the 3,000 training rows. Median 17 training copies each. Across train and eval, no such combination ever carries two different labels, on any of the three questions.

So a lookup table on those three fields answers 997/1000, all correct.

That makes 97.1% a measure of something else: how often a field no rule reads moves a head. Your category meta.json points the same way. Two of its five confident errors are the "What is included in the {amount} EUR tier?" template, labelled billing, answered other.

One teacher detail before anyone copies the rule: 21 of the 368 needs_human=true eval tickets are stuck payouts. They trip the "money back" cue through "our sellers are waiting for their money back". Nobody asked for a refund.

Two cheap tests would say what happens on someone's own traffic:

  1. hold out one template per category, train on the rest
  2. flip only channel and plan on each eval ticket, count changed answers

Which one do you expect the 72.7% to survive?

·

You're right, and thanks for actually running it. I checked on my side too: 997/1000 at seed 10001, 258 template + impact + tail combos in total, none of them with two labels. So 72.7% / 97.1% is mostly the heads remembering combos, not generalizing.

Ran both tests.

One template per category held out (5 of 20), trained on the rest, 1,000 tickets from the held-out templates only:

  • category: sure on 100%, right on 66.1%
  • urgency: sure on 44.8%, right on 96.2% of those
  • needs_human: sure on 96%, right on 87.1% of those
  • all three sure on 42.6%, all three right on 51.9% of those

Channel and plan flipped on the published heads: category changes on 2.7%, needs_human on 3.5%, urgency on 7.1%.

So 72.7% doesn't survive the first one. And category stays fully sure while it's wrong a third of the time, because the threshold is picked on a random holdout of the same templates and says nothing about a template it never saw. urgency's threshold held up, category's didn't.

On real traffic that's what check sampling is for: a share of local answers still goes to the teacher and the site drops back to shadow when agreement falls. But the demo eval should have shown this. I'll put both results on the model card, call the eval unseen states instead of unseen tickets, and fix the money back cue in the rule. Same 21 of 368 payouts here btw, good catch.

Sure on 100%, right on 66.1% is the number worth keeping. The threshold learned the templates, not the task.

Cheap fix before the card: pick each head's threshold on a holdout split by template, not at random. The same leave-one-template-out you just ran, used for calibration and not only for the eval. Then "sure" has to mean sure on a template it never saw.

urgency held up and category didn't, so the per-head thresholds would likely move by very different amounts. That gap is itself worth a line on the card.

On real traffic there is no template id, but distance to the nearest training ticket plays the same role. It flags a ticket before the head answers, where agreement drift only notices after a batch has gone wrong.

Would you route the far ones straight to the teacher, or keep them inside the sampled share?