SaylorTwift HF Staff
Document Latest releases tab (also refreshes dataset snapshot)
91984fa verified | title: LLM Benchmark Usage Explorer | |
| emoji: π | |
| colorFrom: blue | |
| colorTo: red | |
| sdk: gradio | |
| sdk_version: 6.19.0 | |
| app_file: app.py | |
| pinned: false | |
| license: mit | |
| # LLM Benchmark Usage Explorer | |
| Explore [`SaylorTwift/llm-benchmark-usage`](https://huggingface.co/datasets/SaylorTwift/llm-benchmark-usage) β | |
| which evaluation benchmarks 41 AI labs use, when they first showed up, and how the | |
| mix of benchmark categories has shifted over time (2023β2026). | |
| - **Latest releases** β the most recent models, the benchmarks they report, and their sources | |
| - **Popularity & timeline** β most-used benchmarks, filterable by category/openness | |
| - **Benchmark β Models** β every model that reports a given benchmark | |
| - **Model β Benchmarks** β full evaluation suite for a given model | |
| - **Category evolution** β how benchmark categories (coding, agentic, math, safety, ...) have shifted over time | |