Cross-platform-bench The benchmarks evaluate LM agent on SWE/Computer-use tasks across different operating systems. SWE-bench-Live/Windows Viewer • Updated 9 days ago • 66 • 1.67k SWE-bench-Live/OS-bench Viewer • Updated Jul 21 • 140 • 1.4k
SWE-bench-Live The datasets for benchmarking and training LLM coding agents. SWE-bench-Live/SWE-bench-Live Viewer • Updated 24 days ago • 3.69k • 85k • 9 SWE-bench-Live/MultiLang Viewer • Updated 9 days ago • 1.08k • 24.1k SWE-bench-Live/Windows Viewer • Updated 9 days ago • 66 • 1.67k
Cross-platform-bench The benchmarks evaluate LM agent on SWE/Computer-use tasks across different operating systems. SWE-bench-Live/Windows Viewer • Updated 9 days ago • 66 • 1.67k SWE-bench-Live/OS-bench Viewer • Updated Jul 21 • 140 • 1.4k
SWE-bench-Live The datasets for benchmarking and training LLM coding agents. SWE-bench-Live/SWE-bench-Live Viewer • Updated 24 days ago • 3.69k • 85k • 9 SWE-bench-Live/MultiLang Viewer • Updated 9 days ago • 1.08k • 24.1k SWE-bench-Live/Windows Viewer • Updated 9 days ago • 66 • 1.67k