sentence-transformers/all-MiniLM-L6-v2 Sentence Similarity • 22.7M • Updated Jun 1 • 253M • • 5.97k
Beyond Solver Verdicts: Generative Reward Models for Autoformalization Paper • 2609.11085 • Published 5 days ago • 33
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 7 days ago • 414
StudentSim: Training LLM-based Student Simulators Paper • 2609.01591 • Published 14 days ago • 488
Demystifying Agent Skills: Why They Work-Until They Don't Paper • 2608.14036 • Published Aug 14 • 170
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 29 days ago • 340
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published Aug 12 • 291
Running Featured 769 Agent Memory Leaderboard 🧠769 Unified memory evaluation · Results expected August 12.
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 312
Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMark Paper • 2607.11915 • Published Jul 5 • 4