Cross-platform-bench The benchmarks evaluate LM agent on SWE/Computer-use tasks across different operating systems. SWE-bench-Live/Windows Viewer • Updated Jun 28 • 61 • 155 SWE-bench-Live/OS-bench Viewer • Updated 23 days ago • 140 • 224
SWE-bench-Live The datasets for benchmarking and training of LLM coding agents. SWE-bench-Live/SWE-bench-Live Viewer • Updated Sep 18, 2025 • 3.69k • 10.9k • 7 SWE-bench-Live/MultiLang Viewer • Updated about 9 hours ago • 758 • 3.02k SWE-bench-Live/Windows Viewer • Updated Jun 28 • 61 • 155
Cross-platform-bench The benchmarks evaluate LM agent on SWE/Computer-use tasks across different operating systems. SWE-bench-Live/Windows Viewer • Updated Jun 28 • 61 • 155 SWE-bench-Live/OS-bench Viewer • Updated 23 days ago • 140 • 224
SWE-bench-Live The datasets for benchmarking and training of LLM coding agents. SWE-bench-Live/SWE-bench-Live Viewer • Updated Sep 18, 2025 • 3.69k • 10.9k • 7 SWE-bench-Live/MultiLang Viewer • Updated about 9 hours ago • 758 • 3.02k SWE-bench-Live/Windows Viewer • Updated Jun 28 • 61 • 155