SaylorTwift's picture
SaylorTwift HF Staff
Document Latest releases tab (also refreshes dataset snapshot)
91984fa verified
|
Raw
History Blame Contribute Delete
897 Bytes
---
title: LLM Benchmark Usage Explorer
emoji: πŸ“Š
colorFrom: blue
colorTo: red
sdk: gradio
sdk_version: 6.19.0
app_file: app.py
pinned: false
license: mit
---
# LLM Benchmark Usage Explorer
Explore [`SaylorTwift/llm-benchmark-usage`](https://huggingface.co/datasets/SaylorTwift/llm-benchmark-usage) β€”
which evaluation benchmarks 41 AI labs use, when they first showed up, and how the
mix of benchmark categories has shifted over time (2023–2026).
- **Latest releases** β€” the most recent models, the benchmarks they report, and their sources
- **Popularity & timeline** β€” most-used benchmarks, filterable by category/openness
- **Benchmark β†’ Models** β€” every model that reports a given benchmark
- **Model β†’ Benchmarks** β€” full evaluation suite for a given model
- **Category evolution** β€” how benchmark categories (coding, agentic, math, safety, ...) have shifted over time