LLMs-to-test Qwen/Qwen3-0.6B Text Generation • 0.8B • Updated Jul 26, 2025 • 27.8M • • 1.5k Qwen/Qwen3-1.7B Text Generation • 2B • Updated Jul 26, 2025 • 7.44M • • 514 Qwen/Qwen3-4B Text Generation • 4B • Updated Jul 26, 2025 • 4.41M • • 675 Qwen/Qwen3-8B Text Generation • 8B • Updated Jul 26, 2025 • 15M • • 1.28k
Datasets-ScaleLLM truthfulqa/truthful_qa Viewer • Updated Jan 4, 2024 • 1.63k • 106k • 289 allenai/qasc Viewer • Updated Jan 4, 2024 • 9.98k • 24.3k • 23 Anthropic/model-written-evals Viewer • Updated Dec 21, 2022 • 3.25k • 2.22k • 67 yesilhealth/Health_Benchmarks Viewer • Updated Apr 20, 2025 • 7.54k • 142 • 10
LLMs-to-test Qwen/Qwen3-0.6B Text Generation • 0.8B • Updated Jul 26, 2025 • 27.8M • • 1.5k Qwen/Qwen3-1.7B Text Generation • 2B • Updated Jul 26, 2025 • 7.44M • • 514 Qwen/Qwen3-4B Text Generation • 4B • Updated Jul 26, 2025 • 4.41M • • 675 Qwen/Qwen3-8B Text Generation • 8B • Updated Jul 26, 2025 • 15M • • 1.28k
Datasets-ScaleLLM truthfulqa/truthful_qa Viewer • Updated Jan 4, 2024 • 1.63k • 106k • 289 allenai/qasc Viewer • Updated Jan 4, 2024 • 9.98k • 24.3k • 23 Anthropic/model-written-evals Viewer • Updated Dec 21, 2022 • 3.25k • 2.22k • 67 yesilhealth/Health_Benchmarks Viewer • Updated Apr 20, 2025 • 7.54k • 142 • 10