Usage Examples
Practical examples for adding evaluations to HuggingFace model repositories using the .eval_results/ format.
Table of Contents
- Setup
- Add Single Benchmark (Recommended)
- Batch Process Trending Models
- Extract from README Tables
- Import from Artificial Analysis
- Common Workflows
Setup
Environment Variables
# Required for creating PRs
export HF_TOKEN="hf_your_write_token_here"
# Optional: for Artificial Analysis source
export AA_API_KEY="your_aa_api_key_here"
Or use a .env file:
cp examples/.env.example .env
# Edit .env with your tokens
Verify Installation
uv run scripts/evaluation_manager.py --help
Add Single Benchmark (Recommended)
The simplest way to add a specific benchmark score to a model.
Basic Usage
# Preview (default - prints YAML without uploading)
uv run scripts/evaluation_manager.py add-eval \
--benchmark HLE \
--repo-id "moonshotai/Kimi-K2-Thinking"
Output:
Looking up HLE score for moonshotai/Kimi-K2-Thinking from model_card...
Found: HLE = 23.9
Generated YAML:
- dataset:
id: cais/hle
task_id: default
value: 23.9
date: "2026-01-14"
source:
url: https://huggingface.co/moonshotai/Kimi-K2-Thinking
name: Model Card
From Artificial Analysis
uv run scripts/evaluation_manager.py add-eval \
--benchmark HLE \
--repo-id "MiniMaxAI/MiniMax-M2.1" \
--source aa
Create PR
# Always check for existing PRs first!
uv run scripts/evaluation_manager.py get-prs --repo-id "model/name"
# If no PRs exist, create one
uv run scripts/evaluation_manager.py add-eval \
--benchmark HLE \
--repo-id "model/name" \
--create-pr
Push Directly (Your Own Model)
uv run scripts/evaluation_manager.py add-eval \
--benchmark GPQA \
--repo-id "your-username/your-model" \
--apply
Provide Score Manually
uv run scripts/evaluation_manager.py add-eval \
--benchmark HLE \
--repo-id "model/name" \
--value 84.5 \
--create-pr
Batch Process Trending Models
Process multiple trending models at once.
Preview Mode (Dry Run)
uv run scripts/batch_eval_prs.py --limit 10 --benchmark HLE --dry-run
Output: ```
Batch Evaluation PR Creator
Benchmark: HLE Source: model_card Pipeline tag: text-generation Limit: 10 Sort: trending Dry run: True
Processing: LiquidAI/LFM2.5-1.2B-Instruct Not found: HLE score not available Processing: MiniMaxAI/MiniMax-M2.1 Found: HLE = 22.2 Status: Would create PR (dry run) ...
Summary: Success: 3 Not found: 7
### Create PRs
```bash
# From model cards
uv run scripts/batch_eval_prs.py --limit 10 --benchmark HLE
# From Artificial Analysis
uv run scripts/batch_eval_prs.py --limit 10 --benchmark HLE --source aa
Sort Options
# By downloads (more established models)
uv run scripts/batch_eval_prs.py --limit 20 --sort downloads --benchmark GPQA
# By likes
uv run scripts/batch_eval_prs.py --limit 10 --sort likes --benchmark MMLU-Pro
Filter by Pipeline Tag
# Only text-generation models (default)
uv run scripts/batch_eval_prs.py --limit 10 --benchmark HLE --pipeline-tag text-generation
# Image generation models
uv run scripts/batch_eval_prs.py --limit 10 --benchmark HLE --pipeline-tag text-to-image
Results Tracking
Results are saved to runs/{benchmark}_{date}_{hash}.json:
cat runs/hle_20260114_abc123.json
{
"benchmark": "HLE",
"source": "aa",
"source_url": "https://artificialanalysis.ai",
"created": "2026-01-14T08:00:00Z",
"results": [
{
"repo_id": "MiniMaxAI/MiniMax-M2.1",
"value": 22.2,
"status": "pr_created",
"source_url": "https://artificialanalysis.ai"
}
]
}
Extract from README Tables
For models with evaluation tables in their README.
Step 1: Inspect Tables
uv run scripts/evaluation_manager.py inspect-tables \
--repo-id "deepseek-ai/DeepSeek-V3"
This shows all tables with their structure, helping you identify which table to extract.
Step 2: Preview Extraction
uv run scripts/evaluation_manager.py extract-readme \
--repo-id "deepseek-ai/DeepSeek-V3" \
--table 1
Step 3: Create PR
uv run scripts/evaluation_manager.py extract-readme \
--repo-id "deepseek-ai/DeepSeek-V3" \
--table 1 \
--create-pr
Import from Artificial Analysis
Import all available benchmarks from Artificial Analysis API.
Preview
uv run scripts/evaluation_manager.py import-aa \
--creator-slug "anthropic" \
--model-name "claude-sonnet-4" \
--repo-id "your-username/claude-mirror"
Create PR
uv run scripts/evaluation_manager.py import-aa \
--creator-slug "anthropic" \
--model-name "claude-sonnet-4" \
--repo-id "your-username/claude-mirror" \
--apply --create-pr
Finding Creator Slug and Model Name
Visit Artificial Analysis and check the URL:
- URL:
https://artificialanalysis.ai/models/{creator-slug}/{model-name}
Common examples:
- Anthropic:
--creator-slug "anthropic" --model-name "claude-sonnet-4" - OpenAI:
--creator-slug "openai" --model-name "gpt-4-turbo" - Meta:
--creator-slug "meta" --model-name "llama-3-70b"
Common Workflows
Workflow 1: Add Missing Benchmark to Popular Model
# 1. Check for existing PRs
uv run scripts/evaluation_manager.py get-prs \
--repo-id "meta-llama/Llama-3.1-8B-Instruct"
# 2. Preview what we'd add
uv run scripts/evaluation_manager.py add-eval \
--benchmark HLE \
--repo-id "meta-llama/Llama-3.1-8B-Instruct"
# 3. Create PR if score found
uv run scripts/evaluation_manager.py add-eval \
--benchmark HLE \
--repo-id "meta-llama/Llama-3.1-8B-Instruct" \
--create-pr
Workflow 2: Batch Update Trending Models
# 1. Dry run to see which models have HLE scores
uv run scripts/batch_eval_prs.py --limit 20 --benchmark HLE --source aa --dry-run
# 2. Create PRs for models with scores
uv run scripts/batch_eval_prs.py --limit 20 --benchmark HLE --source aa
# 3. Check results
cat runs/hle_*.json | jq '.results[] | select(.status == "pr_created")'
Workflow 3: Update Your Own Model
# 1. Add HLE score from your model card
uv run scripts/evaluation_manager.py add-eval \
--benchmark HLE \
--repo-id "your-username/your-model" \
--apply
# 2. Add GPQA score manually
uv run scripts/evaluation_manager.py add-eval \
--benchmark GPQA \
--repo-id "your-username/your-model" \
--value 84.5 \
--apply
Output Format
Results are stored in .eval_results/*.yaml:
- dataset:
id: cais/hle # Hub Benchmark dataset ID
task_id: default # Optional task ID
value: 23.9 # Metric value
date: "2026-01-14" # ISO-8601 date
source: # Attribution
url: https://huggingface.co/model/name
name: Model Card
Supported Benchmarks
| Benchmark | Hub Dataset ID |
|---|---|
| HLE | cais/hle |
| GPQA | Idavidrein/gpqa |
| MMLU-Pro | TIGER-Lab/MMLU-Pro |
| GSM8K | openai/gsm8k |
To add a new benchmark, update examples/metric_mapping.json.
Troubleshooting
"AA_API_KEY not set"
export AA_API_KEY="your-key"
# or add to .env file
"Could not find benchmark in model card"
The benchmark name may be formatted differently in the README. Check the model card manually.
"Token does not have write access"
Generate a new token at https://huggingface.co/settings/tokens with Write scope.
Getting Help
uv run scripts/evaluation_manager.py --help
uv run scripts/evaluation_manager.py add-eval --help
uv run scripts/batch_eval_prs.py --help
For more information:
- HuggingFace Eval Results Documentation
- SKILL.md - Complete skill documentation