T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks Paper • 2609.11042 • Published 12 days ago • 63
yandex/AliceAI-Foundation-80B-A3B-Base Text Generation • 81B • Updated about 10 hours ago • 676 • 174
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics Paper • 2609.10712 • Published 13 days ago • 44
Nemotron Labs IMO 2026 Collection Checkpoints, training data and benchmark from 'An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics' (IMO 2026). • 7 items • Updated 10 days ago • 9
Hardwired-Neurons Language Processing Units as General-Purpose Cognitive Substrates Paper • 2508.16151 • Published Aug 22, 2025 • 2
view article Article Training a coding model to paint watercolours with TRL and OpenEnv sergiopaniego • 19 days ago • 71