view article Article Rewarding fact fidelity with an LLM judge in GRPO: lessons from 38,000 judged rewrites jialinyyzz • 7 days ago • 3
view article Article SmolLM3: smol, multilingual, long-context reasoner +21 eliebak, cmpatino, anton-l, edbeeching, m-ric, nouamanetazi, akseljoonas, guipenedo, hynky, clefourrier, SaylorTwift, kashif, qgallouedec, hlarcher, glutamatt, Xenova, reach-vb, ngxson, craffel, lewtun, loubnabnl, lvwerra, thomwolf • Jul 8, 2025 • 800
Measuring Mathematical Problem Solving With the MATH Dataset Paper • 2103.03874 • Published Mar 5, 2021 • 8