view article Article Transformers now runs llama.cpp quants +1 marcsun13, ArthurZ, lysandre • 1 day ago • 42
view article Article Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem MultiverseComputingCAI • 2 days ago • 26
view article Article tokenizers v1: encode, decode and scaling, measured +2 ArthurZ, sbrandeis, mcpotato, lysandre • 2 days ago • 65
view article Article Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL +2 aminediroHF, qgallouedec, kashif, sergiopaniego • 13 days ago • 50
view article Article IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license ibm-research • 14 days ago • 60
ibm-granite/granite-timeseries-patchtst-fm-r2 Time Series Forecasting • 0.4B • Updated 14 days ago • 154k • 16
view article Article Training a coding model to paint watercolours with TRL and OpenEnv sergiopaniego • 20 days ago • 71
view article Article NeoMME: an efficient Multimodal-native and Multilingual Encoder Hcompany • 20 days ago • 110
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction Paper • 2608.26005 • Published 28 days ago • 163
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher Paper • 2608.26872 • Published 27 days ago • 65
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution Paper • 2608.25593 • Published 28 days ago • 69
view article Article Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers tomaarsen • 28 days ago • 144