File size: 897 Bytes
6042672 99a24d2 6042672 99a24d2 6042672 99a24d2 6042672 99a24d2 91984fa 99a24d2 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 | ---
title: LLM Benchmark Usage Explorer
emoji: π
colorFrom: blue
colorTo: red
sdk: gradio
sdk_version: 6.19.0
app_file: app.py
pinned: false
license: mit
---
# LLM Benchmark Usage Explorer
Explore [`SaylorTwift/llm-benchmark-usage`](https://huggingface.co/datasets/SaylorTwift/llm-benchmark-usage) β
which evaluation benchmarks 41 AI labs use, when they first showed up, and how the
mix of benchmark categories has shifted over time (2023β2026).
- **Latest releases** β the most recent models, the benchmarks they report, and their sources
- **Popularity & timeline** β most-used benchmarks, filterable by category/openness
- **Benchmark β Models** β every model that reports a given benchmark
- **Model β Benchmarks** β full evaluation suite for a given model
- **Category evolution** β how benchmark categories (coding, agentic, math, safety, ...) have shifted over time
|