T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks Paper • 2609.11042 • Published 14 days ago • 63
TaH Collection Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models • 7 items • Updated Aug 6 • 2
TaH Collection Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models • 7 items • Updated Aug 6 • 2
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published Jul 30 • 188