# AgentEval
Measure what agents actually do — not just what they say.
---
## Evaluation for AI systems that act
**AgentEval** is an independent Hugging Face organization focused on evaluating AI agents as complete systems.
A strong agent should do more than produce a good-looking answer.
It should:
- understand the task
- choose the right tools
- use them correctly
- recover from failure
- stay within constraints
- complete the task
- do so efficiently
- produce a verifiable result
That requires a different evaluation mindset.
> **Outcome first. Trace second. Model score third.**
---
# What should an agent evaluation measure?
A useful agent evaluation can include several layers:
```text
TASK
↓
PLAN
↓
TOOL CHOICE
↓
TOOL EXECUTION
↓
RECOVERY
↓
FINAL RESULT
↓
VERIFICATION
```
AgentEval focuses on the full path.
---
## 01 · Task Success
The most important question:
> **Did the agent complete the task?**
Possible metrics:
- success / failure
- partial completion
- goal coverage
- final answer correctness
- constraint compliance
- completion consistency across repeated runs
A fluent answer is not enough if the task was not completed.
---
## 02 · Tool Use
Agents increasingly depend on external tools.
Evaluation can inspect:
- correct tool selection
- valid arguments
- unnecessary tool calls
- failed tool calls
- retries
- fallback behavior
- sequence of tool usage
- tool-result interpretation
---
## 03 · Traces
Agent behavior becomes easier to understand when the execution trace is visible.
Useful trace elements may include:
- steps
- tool calls
- model calls
- errors
- retries
- timestamps
- latency
- token usage
- cost
- intermediate state
- final result
A final answer can hide a bad process.
A trace makes the process inspectable.
---
## 04 · Reliability
One successful run does not prove reliability.
AgentEval is interested in repeated execution.
Possible metrics:
- pass rate
- variance across runs
- retry rate
- failure mode frequency
- deterministic vs. unstable behavior
- recovery success
- consistency under perturbation
---
## 05 · Efficiency
An agent can complete a task and still be inefficient.
Possible efficiency metrics:
- number of steps
- tool calls
- model calls
- latency
- token usage
- estimated cost
- redundant actions
- unnecessary retries
The best agent is not always the one with the highest raw capability.
Sometimes it is the one that completes the task with fewer resources.
---
## 06 · Safety & Constraint Following
Agent evaluations may also inspect whether a system stays within defined boundaries.
Examples:
- allowed tools only
- no unauthorized actions
- no secret leakage
- no unsafe command execution
- approval required before sensitive actions
- adherence to role or workflow constraints
---
# Possible Spaces
### Agent Task Evaluator
Compare expected outcomes with an agent's final result and trace.
### Tool-Call Grader
Check whether an agent selected the correct tool and valid parameters.
### Trace Inspector
Upload an agent trace and inspect steps, retries, latency, failures, and cost.
### Reliability Lab
Run repeated synthetic evaluations and compare consistency.
### Agent Efficiency Score
Measure task completion against steps, token usage, tool calls, and runtime.
### Failure Mode Explorer
Analyze why an agent failed and classify recurring failure patterns.
### Multi-Agent Evaluator
Compare collaboration quality between planner, executor, reviewer, and verifier roles.
### Agent Benchmark Builder
Create custom task sets with expected outputs and evaluation criteria.
### Tool-Use Benchmark
Score agents on structured tool-selection and execution tasks.
### Agent Regression Check
Compare two agent versions on the same evaluation set.
---
# Possible Datasets
AgentEval can host datasets designed specifically for evaluation.
Examples:
```text
agent-eval-tasks
tool-use-evals
agent-failure-cases
agent-traces
multi-agent-tasks
recovery-scenarios
constraint-following-tests
agent-regression-suite
```
A useful evaluation dataset may include:
- task
- environment
- available tools
- expected result
- allowed actions
- prohibited actions
- success criteria
- reference trace
- scoring rubric
---
# Possible Models
Models may support evaluation tasks such as:
- trace classification
- failure-mode detection
- tool-call validation
- task-success prediction
- reward modeling
- output grading
- execution-quality scoring
- anomaly detection in agent traces
---
# A simple evaluation record
```json
{
"task": "Find the cheapest valid route",
"success": true,
"tool_calls": 4,
"failed_tool_calls": 1,
"steps": 7,
"latency_seconds": 12.4,
"estimated_cost": 0.031,
"constraint_violations": 0
}
```
One record is useful.
Hundreds of repeated records become a benchmark.
---
# Core Evaluation Dimensions
| Dimension | Core Question |
|---|---|
| **Success** | Did the agent complete the task? |
| **Correctness** | Was the result right? |
| **Tool Use** | Were the right tools used correctly? |
| **Reliability** | Does it work repeatedly? |
| **Efficiency** | How much time, cost and work did it require? |
| **Recovery** | Can it recover from failure? |
| **Safety** | Did it respect constraints? |
| **Trace Quality** | Is the execution understandable and auditable? |
---
# AgentEval Scorecard
A practical evaluation may combine multiple signals:
```text
Task Success 40%
Tool Accuracy 20%
Reliability 15%
Efficiency 10%
Recovery 10%
Constraint Safety 5%
```
The exact weighting depends on the use case.
AgentEval does not promote one universal score.
Different agents require different evaluation criteria.
---
# Why agent evaluation is different
Traditional model evaluation often asks:
> How good is the answer?
Agent evaluation may need to ask:
> Did the system complete the task correctly, efficiently, safely, and repeatably?
That distinction matters.
An agent is a system.
Its quality depends on more than one model response.
---
# Principles
### Measure outcomes
A convincing explanation does not equal successful execution.
### Inspect traces
Execution history matters.
### Repeat evaluations
One run is not enough.
### Separate capability from reliability
An agent may be capable but unstable.
### Include cost and latency
Operational quality matters.
### Evaluate tool behavior
Tool misuse can invalidate an otherwise correct result.
### Make failure visible
Benchmarks should help explain failure, not hide it.
---
# Technology Directions
Projects may use:
- Hugging Face Spaces
- Hugging Face Datasets
- structured evaluation sets
- agent traces
- tool-calling logs
- JSON evaluation records
- Python
- JavaScript
- benchmark harnesses
- scoring pipelines
- LLM-as-a-judge experiments
- deterministic validators
- regression testing
- observability data
---
# Who is AgentEval for?
AgentEval may be useful for:
- agent developers
- AI engineers
- eval teams
- platform engineers
- researchers
- QA teams
- MLOps teams
- startups
- enterprise AI teams
- tool developers
- open-source contributors
---
# Important Note
Projects published here are primarily intended for:
- research
- development
- benchmarking
- education
- prototyping
- technical experimentation
Evaluation scores are not universal guarantees of:
- reliability
- safety
- production readiness
- correctness
- regulatory compliance
- suitability for high-impact use
Agent systems should be evaluated in the environment and context in which they are actually used.
---
# Independent Organization
**AgentEval is an independent Hugging Face community organization.**
It is not an official benchmark authority, certification body, standards organization, model provider, or Hugging Face organization.
The name **AgentEval** reflects the technical focus:
> **evaluation for AI systems that plan, act, use tools, and complete tasks.**
---
# AgentEval
### **Evaluate actions. Measure outcomes. Improve agents.**