How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus Paper • 2609.15504 • Published 7 days ago • 37
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training Paper • 2609.07108 • Published 14 days ago • 35
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics Paper • 2609.10712 • Published 12 days ago • 43
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published 25 days ago • 40
trl-internal-testing/tiny-MuseGlimmerForConditionalGeneration Image-Text-to-Text • 6.52M • Updated Aug 13 • 99.3k • 1
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents Paper • 2608.18423 • Published Aug 19 • 21
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published Aug 7 • 114
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development Paper • 2608.13417 • Published Aug 13 • 59
view article Article TutorMoments: Do AI tutors know when to help and when to hold back? allenai • Aug 7 • 31
Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements Paper • 2607.28661 • Published Jul 22 • 19
When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings Paper • 2608.03994 • Published Aug 4 • 9
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published Jul 31 • 113
Instella-MoE ✨ Collection Family of fully open 16B MoE LLM with 2.8B active params per token, trained on AMD Instinct™ MI300 & MI325 GPUs. https://arxiv.org/abs/2609.00791 • 6 items • Updated 2 days ago • 19