FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published Aug 25 • 152
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published Aug 24 • 211
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence Paper • 2608.11341 • Published Aug 11 • 71
SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs Paper • 2509.20758 • Published Sep 25, 2025 • 2
Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning Paper • 2602.01058 • Published Feb 1 • 45