ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 12 days ago • 214
WebWorld: The Browser as a World Model for Self-Improving Web Code Paper • 2608.30530 • Published 26 days ago • 10
WebWorld: The Browser as a World Model for Self-Improving Web Code Paper • 2608.30530 • Published 26 days ago • 10
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild Paper • 2608.23181 • Published Aug 24 • 34
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild Paper • 2608.23181 • Published Aug 24 • 34
FinanceComplexQA Collection Multilingual-Multimodal-NLP/FinanceComplexQA • 1 item • Updated Jul 23 • 1
MIRA Collection Group-specific quality scorers from MIRA for mid-training data selection. • 12 items • Updated May 27 • 2
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling Paper • 2606.18023 • Published Jun 16 • 161
MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection Paper • 2605.30288 • Published May 29 • 22
MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection Paper • 2605.30288 • Published May 29 • 22
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards Paper • 2605.21467 • Published May 20 • 86
ClawGym: A Scalable Framework for Building Effective Claw Agents Paper • 2604.26904 • Published Apr 29 • 55