SaylorTwift HF Staff
Document Latest releases tab (also refreshes dataset snapshot)
91984fa verified A newer version of the Gradio SDK is available: 6.23.1
metadata
title: LLM Benchmark Usage Explorer
emoji: π
colorFrom: blue
colorTo: red
sdk: gradio
sdk_version: 6.19.0
app_file: app.py
pinned: false
license: mit
LLM Benchmark Usage Explorer
Explore SaylorTwift/llm-benchmark-usage β
which evaluation benchmarks 41 AI labs use, when they first showed up, and how the
mix of benchmark categories has shifted over time (2023β2026).
- Latest releases β the most recent models, the benchmarks they report, and their sources
- Popularity & timeline β most-used benchmarks, filterable by category/openness
- Benchmark β Models β every model that reports a given benchmark
- Model β Benchmarks β full evaluation suite for a given model
- Category evolution β how benchmark categories (coding, agentic, math, safety, ...) have shifted over time