dataagent-phase0-evals / CONTRACT.md
AdithyaSK's picture
AdithyaSK HF Staff
Harbor run viewer: Phase 0 eval, 15 harnesses x 50 tasks, pass@4
3d20eb8 verified
|
Raw
History Blame Contribute Delete
5.74 kB
# Information contract
Everything the viewer shows lives in **one directory**, identical whether it is a local folder or a
Hugging Face bucket mounted into a Space. A mounted bucket appears as an ordinary filesystem path inside
the container, so local and hosted use the same code path and there is nothing to keep in sync.
## Hierarchy
```
project "data-agent" β€” a body of work
└── dataset "eval-v1", "eval-easy50", "dabstep" β€” variations within it
└── task one ROW in the table
└── cell one (model, harness) pair
└── attempts k tries β†’ pass@k, each with a trace
```
Two axes cross at a cell, and **which one is the column is a view choice, not a storage choice** β€” so
cells are keyed `"<model>|<harness>"` and the viewer pivots either way. Storing per-axis would force a
rewrite to flip the table.
```
<DATA_DIR>/
└── projects/
└── <project_id>/
β”œβ”€β”€ project.json # REQUIRED β€” label, description, source, support tiers
β”œβ”€β”€ datasets/
β”‚ └── <dataset_id>/
β”‚ β”œβ”€β”€ dataset.json # optional label/notes
β”‚ β”œβ”€β”€ summary.json # tasks[] with cells, by_model, by_harness
β”‚ └── traces/<trace_id>.json # one attempt each; referenced from a cell
└── runs/
└── <run_id>/ # training runs belonging to this project
β”œβ”€β”€ run.json
└── train/metrics.jsonl
```
Discovery is by file presence: `project.json` makes a project, `summary.json` a dataset, `run.json` a
training run. Nothing else is required and the viewer never assumes a file exists.
## `project.json`
```json
{
"project_id": "data-agent",
"label": "Data-Agent Bench",
"description": "Verified data-analysis tasks over Kaggle datasets, graded exact / numeric / LLM-judge.",
"source": {"hf_dataset": "AdithyaSK/data_agent_rl_environment_eval"},
"support": {
"mini-swe-agent": {"tier": "stable"},
"opencode": {"tier": "experimental",
"caveats": ["no step limit; ~90 turns and ~17% of rollouts retry after exit 137"]}
}
}
```
`support` is per-harness and drives the stable / experimental badge and its warning. A tier with no
caveat is just a colour, so an experimental harness has to say what does not work.
## `datasets/<id>/summary.json`
```json
{
"k_max": 4,
"models": ["Qwen/Qwen3.5-2B"],
"harnesses": ["mini-swe-agent", "opencode"],
"summary": {"tasks_total": 50, "tasks_any_pass": 31, "attempts_total": 400,
"attempts_passed": 74, "n_all_infra": 0},
"by_model": [{"model": "Qwen/Qwen3.5-2B", "cells": 100, "pass@1": 0.17, "pass@4": 0.31,
"n_measured": 100, "mean_turns": 11.7}],
"by_harness": [{"harness": "mini-swe-agent", "cells": 100, "pass@1": 0.17, "pass@4": 0.31,
"n_measured": 100, "mean_turns": 11.7}],
"tasks": [
{"id": "0000_416_416942_qa_3", "index": 7, "difficulty_level": 0, "difficulty": "easy",
"question": "What is the median of ...?", "answer": "0.4056", "reward_mode": "numeric",
"cells": {
"Qwen/Qwen3.5-2B|mini-swe-agent": {
"passed_at": 3,
"attempts": [{"attempt": 1, "reward": 0.0, "n_turns": 12, "elapsed_sec": 41.2,
"trace": "traces/0000_416_416942_qa_3-2b-mini-1.json"}]
}
}}
]
}
```
Conventions that carry meaning rather than being formatting:
* **`reward: null` means the verifier never ran** β€” dead sandbox, no answer. Not a zero. Aggregates
exclude those attempts and count a task under `n_all_infra` when every attempt is null. Averaging them
in as zeros makes infrastructure failure look like a weak model: one run reported `pass@4 0.333` from
3 measured tasks out of 8 before that rule existed.
* **`passed_at`** is the first passing attempt, so `pass@1` and `pass@k` both derive from one structure
and cannot disagree.
* **`trace`** is a path relative to the dataset directory, so opening a cell needs no naming convention
the viewer has to guess.
## `traces/<trace_id>.json`
Free-form per attempt; rendered as whatever is present.
```json
{"task_id": "...", "model": "...", "harness": "...", "attempt": 1,
"reward": 0.0, "reward_method": "llm", "n_turns": 12, "elapsed_sec": 41.2,
"gold": "0.4056", "predicted": "0.41",
"messages": [{"role": "user", "content": "..."},
{"role": "assistant", "content": "...", "tool_calls": [...]},
{"role": "tool", "content": "..."}]}
```
`reward_method` is worth carrying: the grader is three-tier (exact β†’ numeric β†’ gpt-4o-mini judge), so
knowing a 0 came from the judge rather than exact match separates a wrong answer from a formatting
mismatch.
## `runs/<run_id>/train/metrics.jsonl`
One object per step, append-only so a live run is readable without rewriting.
```json
{"step": 3, "reward": 0.829, "reward_std": 0.146, "loss": 0.279, "ratio": 0.9999,
"kl": 0.032, "entropy": 0.51, "n_rollouts": 8, "n_unscorable": 1, "wall_s": 47.6}
```
`reward_std` sits next to `reward` because a step whose generations all score the same has zero
advantage and teaches nothing β€” a reward curve alone cannot tell learning from a flat run.
## Publishing
```sh
hf buckets create <ns>/harbor-runs --exist-ok
HF_HUB_DISABLE_XET=1 hf sync ./data hf://buckets/<ns>/harbor-runs
hf spaces volumes set <ns>/harbor-viewer -v hf://buckets/<ns>/harbor-runs:/data
```
`HF_HUB_DISABLE_XET=1` is not decoration: the Xet upload path hangs indefinitely on this cluster while
the plain HTTP path finishes in seconds.