AI & ML interests

None defined yet.

pollixย 
posted an update 2 days ago
view post
Post
58
stuntd 0.1.4 is out, small one.

Last time someone asked what the novelty gate does at the 0.99 quantile instead of 0.95. I checked: at 0.95 every head sends about 5% of totally normal traffic to the big model on purpose, and with three fields that adds up. So 0.99 is the default now.

It also ships examples/support/check.py. One command trains the demo heads twice (all templates, and with one template per category left out) and prints the whole sweep, screenshot is its output with the junk-input part cut. Every generalization number in the README comes from that script now, so you can rerun it instead of trusting me.

What it shows on the support demo:
- normal tickets answered locally: 70.6% at 0.99 vs 67.1% at 0.95 (72.3% with no gate), still 97% right
- tickets from templates the heads never saw: 0.3% answered locally at 0.99. Without the gate the heads answer 69% of those and get only 42% right
- the weather question, asdf and {} are stopped at every setting

Also auto_retrain finally works for Jev captures, it only counted OpenAI and Anthropic before.

pip install -U stuntd

https://github.com/bladedevoff/stuntd/releases/tag/v0.1.4
  • 1 reply
ยท
pollixย 
posted an update 6 days ago
view post
Post
7388
stuntd 0.1.3 is out, and most of it started with a comment under my last post :)

@dipankarsarkar ran the support demo himself and showed my "1,000 tickets the heads never saw" were new wording of known tickets, not new kinds. On templates left out of training the heads were sure on 40% of tickets and right on only 61% of those. They also answered "what's the weather in Paris" like it was a billing ticket.

So now every head keeps the encoder vectors of what it trained on, and anything far from all of them goes to the big model no matter how confident the head is.

Screenshot is the same three requests with the gate off and on. It's local Jev mode with no provider, so the gate hands them to zero-shot Laya. Behind OpenAI or Anthropic they'd go to your model.

On the support demo it costs about 6 points of local answers on normal tickets (76.6% to 70.6%, still 97% right), and tickets from unseen templates go from 40% answered locally to 0. Funny detail: a ticket I typed by hand counted as new too, these demo heads only ever saw 20 templates.

Also in this one: auto retrain counts distinct texts and waits between runs, report shows intervals, import skips duplicates, site rm/rename, /healthz and stuntd stop.

pip install -U stuntd


https://github.com/bladedevoff/stuntd/releases/tag/v0.1.3
  • 1 reply
ยท
pollixย 
posted an update 8 days ago
view post
Post
5293
First stuntd model is on the Hub :)

pollix/stuntd-support-triage is three small heads on the Laya encoder that triage a support ticket in one request: category, urgency and needs_human. About 50 MB each, all three answers come back at a p50 of 71ms through the daemon.

On 1,000 tickets they never saw, each head answers on its own when it's sure: category 99.9%, needs_human 92%, urgency 76%. A ticket only skips the big model when all three are sure, that's 72.7% of them, and all three are right on 97.1% of those.

It's the support demo from the repo, so the tickets are generated and the teacher is a rule. The point is to show what a head looks like and how fast it is, then you train the same thing on your own traffic with your own LLM as the teacher.

hf download pollix/stuntd-support-triage --local-dir support-heads


Model: pollix/stuntd-support-triage
Everything in one place: pollix/stuntd-6abe0a33303828e10c72ab41
Code: https://github.com/bladedevoff/stuntd
  • 3 replies
ยท
pollixย 
posted an update 9 days ago
view post
Post
6449
stuntd 0.1.2 is out ๐ŸŽ‰

stuntd sits in front of your LLM, learns its typed decisions and answers the confident ones locally with a small head on the Laya encoder by @convaiinnovations . About 20ms on GPU and 60ms on CPU, and anything it isn't sure about still goes to the big model.

New in 0.1.2:
- decisions with several fields, like category + urgency + needs_human in one call, answered locally only when every field is sure
- the Anthropic Messages API learns too, not only OpenAI
- auto_retrain: the daemon retrains a site in the background once enough new traffic comes in, so collect, train, shadow and live run on their own
- serve --lazy loads the checkpoint on the first request

Try it in the browser: pollix/stuntd
Code: https://github.com/bladedevoff/stuntd

pip install -U stuntd
  • 1 reply
ยท
pollixย 
posted an update 15 days ago
view post
Post
949
Built a Space for stuntd: pollix/stuntd

For each question it shows the same answer twice: zero-shot Laya, and a small head trained on the frozen Laya encoder. You get the option probabilities, the confidence threshold and the latency. When the head isn't sure, it hands the question back to the teacher instead of guessing.

On jevbench, same machine, n=500: agnews 86.0 โ†’ 92.6, banking77 38.2 โ†’ 69.6. Everything runs locally; heads train in 4-20 minutes on a laptop GPU.

Code (Apache-2.0): https://github.com/bladedevoff/stuntd
pollixย 
posted an update 5 months ago
view post
Post
161
Shipped StudioMI300 for the AMD x lablab hackathon. One English sentence becomes a 30-second cinematic reel, end-to-end on a single AMD Instinct MI300X.

Every model in the pipeline is Apache 2.0 or MIT.

๐ŸŽฌ Director Agent โ€” Qwen3.5-35B-A3B via vLLM with AITER MoE acceleration. Plans 6 shots, character bibles, music brief, per-shot voice-over.

๐ŸŽจ Character keyframes โ€” FLUX.2 klein 4B reference editing. No LoRA training step. Identity stays consistent across shots by construction.

๐ŸŽž๏ธ Animation โ€” Wan2.2-I2V-A14B with ParaAttention FBCache (lossless 2x) and selective torch.compile on transformer_2 (another 1.2x). End-to-end Wan2.2 inference went from 25.9 min to 10.4 min per 720p clip.

๐Ÿ” Vision Critic โ€” same Qwen3.5 checkpoint reloaded with a 10-label failure taxonomy (character drift, extras invade frame, camera ignored, walking backwards, hand artifact, wardrobe drift, neon glow leak, stylized AI look, random intimacy, object morphing). Bad clips auto-retry with targeted strategies. Up to 3 attempts.

๐ŸŽต Music โ€” ACE-Step v1 generates 30s instrumental from Director's brief.

๐Ÿ—ฃ๏ธ Narration โ€” Kokoro-82M, 9 languages. Director picks language to match setting. Tokyo to Japanese, Paris to French, Mumbai to Hindi.

The 192 GB HBM3 on MI300X is what lets four very different model architectures share one card sequentially. On a 24 GB consumer GPU this stack needs 4-5 separate machines wired together.

Space (live infra restoring after hackathon close, pls like this space):
lablab-ai-amd-developer-hackathon/studiomi300

Code (Apache 2.0):
https://github.com/bladedevoff/studiomi300

Special thanks to the FLUX, Wan2.2, ACE-Step and Kokoro teams for keeping serious generative AI open. The pipeline composes their work into something none of them alone can produce โ€” a complete cinematic artifact from a single prompt.

#AMDHackathon #ROCm #MI300X #OpenSource