We're releasing the BananaMind SLM Leaderboard! It offers a easier look at which models are actually good for your specific needs. Its primary metric, Intelligence index is a composite of BananaMind Base Bench, PIQA, Hellaswag, ARC Easy and Arithmark 3. It also allows you to see specific categories like Commonsense on a model.
176 models, 54 orgs, 5 benchmarks, and a whole community of support!
Thanks to everyone whoβs contributed models, reported issues, suggested benchmark improvements, or used the leaderboard to compare and evaluate small language models.
Itβs been awesome watching the leaderboard grow into a broader community resource for transparent and reproducible SLM evaluation.