EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published Aug 20 • 175
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling Paper • 2608.26623 • Published 29 days ago • 20
StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments Paper • 2608.24804 • Published Aug 25 • 41
💫StarShell Collection Resources for paper "Terminal Agents Suffice for Enterprise Automation". Repo: https://github.com/servicenow/starshell • 7 items • Updated Aug 5 • 14
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published Jul 11 • 69
Beyond Transcription: Mechanistic Interpretability in ASR Paper • 2508.15882 • Published Aug 21, 2025 • 92
view article Article Can Voice Agents Handle Bilingual Customers? Benchmarking Frontier ASR on Code-Switched Speech ServiceNow-AI • Jun 9 • 46
view article Article EVA-Bench Data 2.0: 3 Domains, 121 Tools, 213 Scenarios ServiceNow-AI • Jun 4 • 42
view article Article EVA-Bench Data 2.0: 3 Domains, 121 Tools, 213 Scenarios ServiceNow-AI • Jun 4 • 42
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents Paper • 2605.13841 • Published May 13 • 78 • 4
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents Paper • 2605.13841 • Published May 13 • 78
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents Paper • 2605.13841 • Published May 13 • 78
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics Paper • 2605.12178 • Published May 12 • 66