benchmarks moonshotai/PerceptionBench Viewer • Updated 25 days ago • 3k • 6.88k • 49 SWE-bench/SWE-bench_Verified Benchmark • Updated 10 days ago • 500 • 102k • 148 datacurve/deep-swe Benchmark • Updated Jun 2 • 113 • 935 • 62 cais/hle Benchmark • Updated Jan 20 • 2.5k • 38.7k • 932
post-training Glint-Research/Fable-5-traces Traces • Updated Jun 29 • 4.67k • 70.9k • 723 Qyrou/reasoning-corpus-4K-5M-v1 Preview • Updated 4 days ago • 14.9k • 204 yannelli/laravel-11-qa Viewer • Updated Oct 4, 2024 • 12.6k • 75 • 6 Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset Viewer • Updated Jul 18 • 18.5M • 21.9k • 186
Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset Viewer • Updated Jul 18 • 18.5M • 21.9k • 186
benchmarks moonshotai/PerceptionBench Viewer • Updated 25 days ago • 3k • 6.88k • 49 SWE-bench/SWE-bench_Verified Benchmark • Updated 10 days ago • 500 • 102k • 148 datacurve/deep-swe Benchmark • Updated Jun 2 • 113 • 935 • 62 cais/hle Benchmark • Updated Jan 20 • 2.5k • 38.7k • 932
post-training Glint-Research/Fable-5-traces Traces • Updated Jun 29 • 4.67k • 70.9k • 723 Qyrou/reasoning-corpus-4K-5M-v1 Preview • Updated 4 days ago • 14.9k • 204 yannelli/laravel-11-qa Viewer • Updated Oct 4, 2024 • 12.6k • 75 • 6 Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset Viewer • Updated Jul 18 • 18.5M • 21.9k • 186
Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset Viewer • Updated Jul 18 • 18.5M • 21.9k • 186