aif-benchmarks / README.md
FrostyJay7813's picture
Upload README.md with huggingface_hub
804f3a8 verified
|
Raw
History Blame Contribute Delete
826 Bytes
metadata
license: apache-2.0
language: en
tags:
  - agentic-ai
  - benchmarks
  - evaluation
  - gaia
  - swe-bench
  - agentbench
pretty_name: AIF Benchmarks
size_categories: n<1K

AIF Benchmarks

Evaluation datasets and results for benchmarking AIF agent performance against industry standards.

Benchmarks

Benchmark AIF Score Industry Baseline
GAIA 0.840 0.750
SWE-bench 0.782 0.650
AgentBench 0.820 0.700
Protocol Conformance 48/48 Varies

Contents

  • Task definitions for each benchmark
  • Ground truth answers and scoring rubrics
  • AIF execution traces and results
  • Historical performance trends
  • Regression detection metadata

Usage

from datasets import load_dataset

ds = load_dataset("FrostyJay7813/aif-benchmarks")