JacobPEvans's picture
Deploy from 1ad58abad9750ed5d30bb5f033c99a815a3813a2
b8182d6 verified
|
Raw
History Blame Contribute Delete
2.38 kB

A newer version of the Gradio SDK is available: 6.20.0

Upgrade
metadata
title: MLX Benchmarks Viewer
emoji: πŸ“Š
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 6.13.0
app_file: app.py
python_version: '3.11'
pinned: true
license: apache-2.0
short_description: Interactive viewer for the MLX Benchmarks dataset
tags:
  - benchmarks
  - mlx
  - apple-silicon
  - lm-eval
  - visualization
models:
  - mlx-community/Qwen3.5-9B-MLX-4bit
datasets:
  - JacobPEvans/mlx-benchmarks

Interactive Gradio viewer for the JacobPEvans/mlx-benchmarks dataset β€” a collection of benchmark runs of MLX-quantized and locally-hosted LLMs on Apple Silicon, all serialized under the envelope v1 schema.

Installation

Requires Python 3.11+.

cd space
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Usage

Launch the viewer locally:

python app.py

Gradio prints a local URL (default http://127.0.0.1:7860). The app pulls all data/*.parquet shards from the HF dataset, caches them for 10 minutes, and offers three tabs:

  • Bar chart β€” latest run β€” one bar per model for a given (suite, task, metric)
  • Trend β€” over time β€” score trajectory per model across runs
  • Summary table β€” pivot of models x tasks for the selected (suite, metric)

Hit Refresh data to invalidate the cache manually.

Deployment

This Space is synced automatically from the JacobPEvans/mlx-benchmarks GitHub repository via the deploy-space.yml workflow on every main push that touches space/.

Contributing

Issues and PRs live in the upstream GitHub repo. See the repository CONTRIBUTING.md for the full developer workflow.

API

The viewer is a Gradio UI β€” it has no stable programmatic API. To query the underlying data yourself, read the HF dataset directly:

import pandas as pd
from huggingface_hub import HfFileSystem

fs = HfFileSystem()
paths = fs.glob("datasets/JacobPEvans/mlx-benchmarks/data/*.parquet")
df = pd.concat([pd.read_parquet(f"hf://{p}") for p in paths], ignore_index=True)

License

Apache-2.0. See LICENSE.

Source

Full source, tests, and developer docs: https://github.com/JacobPEvans/mlx-benchmarks