Compressing visual tokens in vision-language models: 3x more requests per GPU on Qwen2-VL datas3nt • Jun 17 • 1
Intel XPU Kernel Skill: LLM-driven Triton kernel optimization for the Hugging Face Kernel Hub danf • Jun 17 • 11
Don’t Waste Tokens on Data Entry: Tag Customer Reviews Overnight with ZeroGPU Batch API its-maddy-a • Jun 16 • 1
Yui Home Assistant — teaching a 3B model to write Home Assistant automations that actually *work* build-small-hackathon • Jun 15
My First MO§ES™ SigRank — the diagnostic x-ray of the token economy build-small-hackathon • Jun 15 • 1
Keep the model on one axis: building *Penny Stock Persuasion* on a small Nemotron Han-Solo • Jun 15 • 1