metadata
license: apache-2.0
language: en
tags:
- agentic-ai
- benchmarks
- evaluation
- gaia
- swe-bench
- agentbench
pretty_name: AIF Benchmarks
size_categories: n<1K
AIF Benchmarks
Evaluation datasets and results for benchmarking AIF agent performance against industry standards.
Benchmarks
| Benchmark | AIF Score | Industry Baseline |
|---|---|---|
| GAIA | 0.840 | 0.750 |
| SWE-bench | 0.782 | 0.650 |
| AgentBench | 0.820 | 0.700 |
| Protocol Conformance | 48/48 | Varies |
Contents
- Task definitions for each benchmark
- Ground truth answers and scoring rubrics
- AIF execution traces and results
- Historical performance trends
- Regression detection metadata
Usage
from datasets import load_dataset
ds = load_dataset("FrostyJay7813/aif-benchmarks")