SaylorTwift's picture
SaylorTwift HF Staff
Document Latest releases tab (also refreshes dataset snapshot)
91984fa verified
|
Raw
History Blame Contribute Delete
897 Bytes

A newer version of the Gradio SDK is available: 6.23.1

Upgrade
metadata
title: LLM Benchmark Usage Explorer
emoji: πŸ“Š
colorFrom: blue
colorTo: red
sdk: gradio
sdk_version: 6.19.0
app_file: app.py
pinned: false
license: mit

LLM Benchmark Usage Explorer

Explore SaylorTwift/llm-benchmark-usage β€” which evaluation benchmarks 41 AI labs use, when they first showed up, and how the mix of benchmark categories has shifted over time (2023–2026).

  • Latest releases β€” the most recent models, the benchmarks they report, and their sources
  • Popularity & timeline β€” most-used benchmarks, filterable by category/openness
  • Benchmark β†’ Models β€” every model that reports a given benchmark
  • Model β†’ Benchmarks β€” full evaluation suite for a given model
  • Category evolution β€” how benchmark categories (coding, agentic, math, safety, ...) have shifted over time