--- title: LLM Benchmark Usage Explorer emoji: 📊 colorFrom: blue colorTo: red sdk: gradio sdk_version: 6.19.0 app_file: app.py pinned: false license: mit --- # LLM Benchmark Usage Explorer Explore [`SaylorTwift/llm-benchmark-usage`](https://huggingface.co/datasets/SaylorTwift/llm-benchmark-usage) — which evaluation benchmarks 41 AI labs use, when they first showed up, and how the mix of benchmark categories has shifted over time (2023–2026). - **Latest releases** — the most recent models, the benchmarks they report, and their sources - **Popularity & timeline** — most-used benchmarks, filterable by category/openness - **Benchmark → Models** — every model that reports a given benchmark - **Model → Benchmarks** — full evaluation suite for a given model - **Category evolution** — how benchmark categories (coding, agentic, math, safety, ...) have shifted over time