JEV-as-a-Judge: Accept When Confident, Escalate When Unsure Paper • 2609.26550 • Published 2 days ago • 23
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation Paper • 2608.02287 • Published Aug 3 • 32
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces Paper • 2604.05172 • Published Apr 6 • 24
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks Paper • 2602.12670 • Published Feb 13 • 66