AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Paper • 2608.05987 • Published 1 day ago • 57
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Paper • 2608.05139 • Published 3 days ago • 23
QQWorld: Quantile-Quantile Matching for World Model Regularization Paper • 2607.28415 • Published 9 days ago • 30
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published 9 days ago • 181