False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents Paper • 2609.39102 • Published 11 days ago • 572
TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models Paper • 2610.07767 • Published 5 days ago • 89
LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches Paper • 2610.06647 • Published 6 days ago • 97
autotrust/GLM5.3-Flash-E224-DGX-Spark Image-Text-to-Text • 128B • Updated about 3 hours ago • 15.1k • 560