File size: 897 Bytes
6042672
99a24d2
 
 
 
6042672
99a24d2
6042672
 
99a24d2
6042672
 
99a24d2
 
 
 
 
 
91984fa
99a24d2
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
---
title: LLM Benchmark Usage Explorer
emoji: πŸ“Š
colorFrom: blue
colorTo: red
sdk: gradio
sdk_version: 6.19.0
app_file: app.py
pinned: false
license: mit
---

# LLM Benchmark Usage Explorer

Explore [`SaylorTwift/llm-benchmark-usage`](https://huggingface.co/datasets/SaylorTwift/llm-benchmark-usage) β€”
which evaluation benchmarks 41 AI labs use, when they first showed up, and how the
mix of benchmark categories has shifted over time (2023–2026).

- **Latest releases** β€” the most recent models, the benchmarks they report, and their sources
- **Popularity & timeline** β€” most-used benchmarks, filterable by category/openness
- **Benchmark β†’ Models** β€” every model that reports a given benchmark
- **Model β†’ Benchmarks** β€” full evaluation suite for a given model
- **Category evolution** β€” how benchmark categories (coding, agentic, math, safety, ...) have shifted over time