AutoBench Leaderboard
Multi-run AutoBench leaderboard with historical navigation
The third public run of AutoBench has landed, setting a new benchmark for Large Language Model (LLM) evaluation with unmatched scale and precision. Ranking 33 models with over 300,000 individual ranks, it achieves correlations of 90%* with leading benchmarks and delivers granular insights into model quality, cost, and speed – all fully automated and open-source. And as a surprise, we’ve launched autobench.org, your new hub for transparent AI benchmarking.
Building on our first run and second run, this release pushes the limits of LLM evaluation. Here, we’ll unpack the methodology, dive into the third run’s massive stats, highlight top performers and efficiency insights, unveil the new website, acknowledge our partners, and connect to the broader ecosystem, including Bot Scanner. Let’s get started!
With thousands of LLMs flooding the AI landscape, choosing the right model is daunting. Traditional benchmarks are static, gameable, and often too broad to reveal domain-specific strengths. Human evaluations are slow, costly, and subjective, limiting scalability. AutoBench tackles these issues with its Collective-LLM-as-a-Judge methodology, using LLMs to dynamically generate questions, provide answers, and rank outputs. This creates an ungameable, scalable, and objective evaluation system, measuring performance against the AI ecosystem’s consensus.
The result? Accurate insights into LLM quality, real-world costs, and production speed, empowering developers, enterprises, and researchers to make informed choices and avoid costly errors in AI agent workflows.
AutoBench’s automated, iterative workflow ensures robustness and statistical significance:
This cycle repeats hundreds of times, producing aggregate and domain-specific ranks, plus efficiency metrics like cost per answer, average duration, and P99 latency. It’s all open-source and customizable – explore the details on autobench.org or our Hugging Face Space.
The third run, completed in August 2025, is our most ambitious yet:
This massive dataset, processed automatically, underscores AutoBench’s ability to handle the LLM explosion with precision and efficiency.
AutoBench correlations with leading benchmarks.
Our third run’s correlations with established benchmarks confirm AutoBench’s reliability:
These 85-95% correlations are the gold standard, proving our methodology captures true model capabilities without the pitfalls of static datasets or human bias.
The results showcase fierce competition and unexpected standouts:
Outlier Alert: Open-source model GPT OSS 120B is hitting state-of-the-art levels, democratizing high performance. .
Domain-specific insights reveal more:
Explore these nuances on our interactive leaderboard – filter by domain, sort by metrics, and compare runs.
AutoBench average ranks for most comon models
AutoBench goes beyond performance to deliver real-world metrics:
These insights are critical for AI agents, where even minor inefficiencies can cascade into failures.
Scatter plot: Models by cost vs. rank, highlighting value leaders.
The third run’s scale was anticipated, but here’s the surprise: autobench.org is now live! This is your hub for:
Run AutoBench yourself or reach out for custom solutions. This launch makes our mission of transparent AI evaluation more accessible than ever.
Screenshot: Hero section of the new site.
This run was powered by an incredible network:
AutoBench ties into our ecosystem, including Bot Scanner (botscanner.ai, @BotScanner_AI). Dubbed the "Skyscanner of LLM responses," it ranks answers from 40+ models in real-time using AutoBench’s methodology. This run leveraged Bot Scanner’s API for efficiency – try it with $3 free credits.
AutoBench is proudly open-source. Contributors are already shaping its future – join them! Full data, samples, and methodology are on Hugging Face. Fork the third run, run your own benchmarks, or add models.
More info and data in our Hugging Face Repository.
Engage with us on Hugging Face, X (@pwk), or the new website. Let’s build the future of AI evaluation together.
The third run cements AutoBench as a leader in LLM evaluation: 33 models, 300k ranks, 92%+ correlations, and actionable insights into performance and efficiency. In a chaotic AI landscape, we provide clarity for developers, enterprises, and researchers.
What’s your take? Comment below!
Visit autobench.org, explore the leaderboard, or reach out on our website contact form for custom evals. Let’s make AI transparent together.
#AutoBench #LLM #AIBenchmarking #OpenSourceAI #HuggingFace #Leaderboards @Benchmarks
Multi-run AutoBench leaderboard with historical navigation