view article Article Your Inference Server is Secretly a Learner: Reef Infrastructure for Continual Self-Improving Agents quao627 • 5 days ago • 15
view article Article Your Inference Server is Secretly a Learner: Reef Infrastructure for Continual Self-Improving Agents quao627 • 5 days ago • 15
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published Aug 19 • 54 • 2
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published Aug 19 • 54
SPADE Collection The full SPADE release: paper, model checkpoints, grounding corpora, and synthetic environments. • 4 items • Updated Aug 20 • 1
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published Aug 19 • 54
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published Jul 26 • 95
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published Jul 26 • 95