Buckets:
| license: apache-2.0 | |
| language: | |
| - en | |
| size_categories: | |
| - n<1K | |
| tags: | |
| - reinforcement-learning | |
| - data-science | |
| - code-agent | |
| - jupyter | |
| - harbor | |
| - eval | |
| [](https://huggingface.co/spaces/AdithyaSK/harbor-visualiser?dataset=AdithyaSK/data_agent_rl_environment_eval) | |
| # data_agent_rl_environment_eval | |
| **The official verified eval suite for the data-agent RL pipeline.** 366 Harbor-format | |
| data-analysis tasks, each with an LLM-assigned difficulty label (L1βL5), a Kaggle | |
| dataset dependency, and a tested reward function. | |
| > π‘ **Browse this dataset in your browser** β click the badge above or open | |
| > [`AdithyaSK/harbor-visualiser`](https://huggingface.co/spaces/AdithyaSK/harbor-visualiser?dataset=AdithyaSK/data_agent_rl_environment_eval) | |
| > to inspect every task's spec, instruction, environment, tests, and difficulty. | |
| --- | |
| ## Reproduce the eval β end to end | |
| The dataset is **self-contained**: every task ships its own `task.toml`, `instruction.md`, | |
| `environment/Dockerfile`, `environment/pull_bucket.py`, `tests/test.sh`, and `tests/grader.py`. | |
| The only runtime dependencies are Docker, the [`harbor`](https://github.com/harbor-framework/harbor) CLI, | |
| and a `HF_TOKEN` (needed to fetch the Kaggle bucket at container start). | |
| ### 0. Prerequisites | |
| ```bash | |
| # Docker Desktop (or any docker daemon) β at least ~8 GiB of memory available | |
| docker info | |
| # Harbor CLI | |
| pip install harbor | |
| # Tokens β at minimum HF_TOKEN. ANTHROPIC_API_KEY / OPENAI_API_KEY if you use | |
| # those models for your agent. OPENAI_API_KEY also needed for tasks whose | |
| # reward_mode_initial == "llm_judge_long" (the grader uses gpt-4o-mini). | |
| export HF_TOKEN=hf_... | |
| export ANTHROPIC_API_KEY=sk-ant-... | |
| export OPENAI_API_KEY=sk-... | |
| ``` | |
| ### 1. Download the dataset | |
| ```bash | |
| # Option A β Harbor CLI (preferred; uses registry.json) | |
| harbor download data_agent_rl_environment_eval \ | |
| --registry-url https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_eval/resolve/main/registry.json \ | |
| --output ./eval | |
| # Option B β huggingface_hub Python API | |
| python -c " | |
| from huggingface_hub import snapshot_download | |
| snapshot_download( | |
| repo_id='AdithyaSK/data_agent_rl_environment_eval', | |
| repo_type='dataset', | |
| local_dir='./eval', | |
| ) | |
| " | |
| # Option C β just one task | |
| python -c " | |
| from huggingface_hub import snapshot_download | |
| snapshot_download( | |
| repo_id='AdithyaSK/data_agent_rl_environment_eval', | |
| repo_type='dataset', | |
| local_dir='./eval', | |
| allow_patterns=['tasks/0000_419_419825_qa_1/**'], | |
| ) | |
| " | |
| ``` | |
| ### 2. Bring an agent | |
| Implement Harbor's `BaseAgent` protocol β or use any community Harbor agent. | |
| The minimal contract: write the answer to `/workdir/answer.txt` inside the | |
| container. The grader reads it from there. | |
| A reference single-tool ("bash") implementation is below. Save as `my_agent.py` | |
| in the same directory you'll run `harbor` from. | |
| <details> | |
| <summary>Minimal `BashOnlyAgent` β Anthropic + bash tool (~80 lines)</summary> | |
| ```python | |
| """Minimal Harbor agent that exposes a single `bash` tool.""" | |
| from __future__ import annotations | |
| import json, os, time | |
| from harbor.agents.base import BaseAgent | |
| from harbor.environments.base import BaseEnvironment | |
| from harbor.models.agent.context import AgentContext | |
| SYSTEM_PROMPT = ( | |
| "You are an autonomous data-analysis agent in a sandboxed Linux container. " | |
| "Your only tool is `bash`. Dataset files are under /home/user/input/. " | |
| "Python 3, pandas, numpy, scikit-learn and scipy are pre-installed.\n\n" | |
| "Work step-by-step: inspect (ls, head, shape, dtypes), then plan, then " | |
| "compute. To submit your final answer, write it to /workdir/answer.txt " | |
| "with `echo -n \"<value>\" > /workdir/answer.txt`. After the file is " | |
| "written, stop calling tools." | |
| ) | |
| TOOLS = [{ | |
| "type": "function", | |
| "function": { | |
| "name": "bash", | |
| "description": "Run a shell command and return its combined stdout+stderr.", | |
| "parameters": { | |
| "type": "object", | |
| "properties": {"command": {"type": "string"}}, | |
| "required": ["command"], | |
| }, | |
| }, | |
| }] | |
| class BashOnlyAgent(BaseAgent): | |
| SUPPORTS_WINDOWS = False | |
| @staticmethod | |
| def name() -> str: return "bash-only" | |
| def version(self) -> str: return "0.1.0" | |
| async def setup(self, env: BaseEnvironment) -> None: | |
| await env.exec("mkdir -p /workdir", timeout_sec=10) | |
| async def run(self, instruction: str, env: BaseEnvironment, context: AgentContext) -> None: | |
| from anthropic import Anthropic | |
| client = Anthropic() # reads ANTHROPIC_API_KEY from environment | |
| messages = [{"role": "user", "content": instruction}] | |
| for _ in range(25): | |
| resp = client.messages.create( | |
| model=self.model_name, | |
| max_tokens=4096, | |
| system=SYSTEM_PROMPT, | |
| tools=TOOLS, | |
| messages=messages, | |
| ) | |
| tool_calls = [b for b in resp.content if b.type == "tool_use"] | |
| messages.append({"role": "assistant", "content": resp.content}) | |
| if not tool_calls: break | |
| tool_results = [] | |
| for tc in tool_calls: | |
| cmd = (tc.input.get("command") or "").strip() | |
| r = await env.exec(cmd, timeout_sec=180) | |
| out = (r.stdout or "") + (r.stderr or "") | |
| if len(out) > 8000: out = out[:8000] + "\n... [truncated]" | |
| tool_results.append({"type": "tool_result", "tool_use_id": tc.id, "content": out or f"(empty rc={r.return_code})"}) | |
| messages.append({"role": "user", "content": tool_results}) | |
| ``` | |
| </details> | |
| ### 3. Run a single task | |
| ```bash | |
| harbor run \ | |
| -p ./eval/tasks \ | |
| -i 0000_419_419825_qa_1 \ | |
| --env docker \ | |
| --agent-import-path my_agent:BashOnlyAgent \ | |
| --model anthropic/claude-sonnet-4-6 \ | |
| --ae HF_TOKEN="$HF_TOKEN" \ | |
| --ae ANTHROPIC_API_KEY="$ANTHROPIC_API_KEY" \ | |
| --ve OPENAI_API_KEY="$OPENAI_API_KEY" \ | |
| --yes -n 1 \ | |
| --jobs-dir ./eval_jobs | |
| ``` | |
| Flag notes: | |
| - `-p ./eval/tasks` β directory that *contains* per-task folders | |
| - `-i 0000_419_419825_qa_1` β which task (the folder name) | |
| - `--ae KEY=val` β env var forwarded into the **agent** container | |
| - `--ve KEY=val` β env var forwarded into the **verifier** container (only `OPENAI_API_KEY` is needed there, and only for `llm_judge_long` tasks) | |
| - `-n 1` β single trial. Bump to `-n 3` for majority-vote variance. | |
| **Expected output:** | |
| ``` | |
| [turn 0] requesting completion (model=claude-sonnet-4-6) | |
| [turn 0] bash chars=337 | |
| ... | |
| bash-only agent done: provider=anthropic model=claude-sonnet-4-6 calls=3 tokens=5028+303 cost_usd=0.019629 elapsed=7.4s | |
| 1/1 Mean: 1.000 | |
| Reward | |
| 1.0 1 | |
| ``` | |
| The reward lands at `./eval_jobs/<timestamp>/<task_id>__<random>/verifier/reward.txt`. | |
| ### 4. Run the full 366 | |
| ```bash | |
| # Pre-filter β pick e.g. only L3-L5 numeric tasks for a focused eval | |
| python -c " | |
| import pandas as pd | |
| df = pd.read_parquet('hf://datasets/AdithyaSK/data_agent_rl_environment_eval/manifest.parquet') | |
| ids = df[(df.difficulty_level >= 3) & (df.reward_mode_initial == 'numeric')]['task_dir'].tolist() | |
| open('ids.txt','w').write('\n'.join(ids)) | |
| print(f'selected {len(ids)} tasks') | |
| " | |
| # Run them in parallel (concurrency depends on your Docker VM size β 12-30 typical) | |
| harbor run -p ./eval/tasks -i $(paste -sd, ids.txt) \ | |
| --env docker -n 1 -j 16 \ | |
| --agent-import-path my_agent:BashOnlyAgent \ | |
| --model anthropic/claude-sonnet-4-6 \ | |
| --ae HF_TOKEN="$HF_TOKEN" --ae ANTHROPIC_API_KEY="$ANTHROPIC_API_KEY" \ | |
| --ve OPENAI_API_KEY="$OPENAI_API_KEY" \ | |
| --yes --jobs-dir ./eval_jobs | |
| ``` | |
| ### 5. Aggregate results | |
| ```python | |
| # Walk the verifier reward files and compute aggregate stats | |
| import json, glob | |
| from pathlib import Path | |
| import pandas as pd | |
| manifest = pd.read_parquet('hf://datasets/AdithyaSK/data_agent_rl_environment_eval/manifest.parquet') | |
| rewards = {} | |
| for reward_file in glob.glob('./eval_jobs/**/verifier/reward.txt', recursive=True): | |
| trial_dir = Path(reward_file).parent.parent | |
| task_dir = trial_dir.name.split('__', 1)[0] # e.g. 0000_419_419825_qa_1 | |
| rewards[task_dir] = float(open(reward_file).read().strip()) | |
| results = manifest[['task_dir', 'difficulty_level', 'reward_mode_initial']].copy() | |
| results['reward'] = results['task_dir'].map(rewards) | |
| print(f'pass rate overall: {(results.reward == 1.0).mean():.1%}') | |
| print(f'pass rate by difficulty:') | |
| print(results.groupby('difficulty_level')['reward'].agg(['count','mean'])) | |
| ``` | |
| --- | |
| ## Pipeline that produced these 366 | |
| Starting from a 500-task eval pool stratified across `(reward_mode_initial Γ package_tier)`: | |
| - **Stage 1** (Sonnet anchor): one + one retry; tasks that pass go straight to difficulty labeling. | |
| - **Stage 2** (Doctor): for Stage-1 failures, Sonnet's "doctor" agent calls `probe(model)` on `nano`/`gpt-5.5` to cross-check the gold, `rewrite_spec()` (e.g. numericβflexible), `correct_gold()` if the original gold is wrong, or `drop()` if genuinely unverifiable. | |
| **Verdict mix:** | |
| | Verdict | Count | % | Means | | |
| |---|---:|---:|---| | |
| | `verified` | 273 | 75% | Sonnet passed against the original gold (Phase B) | | |
| | `verified_gold_corrected` | 57 | 16% | Doctor's probes converged on a NEW answer; gold was wrong | | |
| | `verifiable_judge` | 20 | 5% | LLM judge agreed agent's answer β‘ gold | | |
| | `verified_after_rewrite` | 16 | 4% | Doctor relaxed `reward_mode` (e.g. numeric β flexible); re-run passed | | |
| (Of the 500-task pool, 127 were dropped as unverifiable, 7 became `phase_b_failed` residue; only verified-class tasks are published here.) | |
| ## Difficulty distribution | |
| | Level | Count | Typical pattern | | |
| |---|---:|---| | |
| | **L1** | 75 | one-line filter / aggregation | | |
| | **L2** | 151 | filter + groupby + aggregate (2-4 turns) | | |
| | **L3** | 71 | multi-step pandas, joins, light feature work | | |
| | **L4** | 68 | ML training / non-trivial pipelines / complex statistical reasoning | | |
| | **L5** | 1 | extreme complexity | | |
| Categorize was an LLM rubric (Sonnet) reading the passing trajectory. | |
| ## Layout | |
| ``` | |
| tasks/ | |
| βββ <task_dir>/ # e.g. 0000_419_419825_qa_1 | |
| βββ task.toml # Harbor task spec β gold_answer, reward_mode, difficulty_level | |
| βββ instruction.md # natural-language question for the agent | |
| βββ environment/ | |
| β βββ Dockerfile # base image | |
| β βββ pull_bucket.py # downloads task's slice from hf://buckets/AdithyaSK/jupyter-agent-kaggle-all | |
| βββ tests/ | |
| βββ test.sh # verifier entrypoint | |
| βββ grader.py # mode-aware grader (exact / numeric / flexible / llm_judge) | |
| manifest.parquet # per-task: task_id, verdict, difficulty, gold, kaggle dataset, question, cost | |
| registry.json # Harbor visualizer index | |
| ``` | |
| `manifest.parquet` is the easiest entry point for filtering: | |
| ```python | |
| import pandas as pd | |
| df = pd.read_parquet('hf://datasets/AdithyaSK/data_agent_rl_environment_eval/manifest.parquet') | |
| # all L1 tasks (easy) | |
| easy = df[df.difficulty_level == 1] | |
| # all gold-corrected tasks (the doctor's gold rewrites β interesting failure modes) | |
| fixed = df[df.verdict == 'verified_gold_corrected'] | |
| ``` | |
| ## Reward modes | |
| Each task's `task.toml` declares `reward_mode_initial` in `[metadata]`. The grader at | |
| `tests/grader.py` is mode-aware: | |
| | Mode | Logic | Pass condition | | |
| |---|---|---| | |
| | `exact_short` | string equality (case-folded, stripped) | answer β‘ gold | | |
| | `numeric` | float parse + atol/rtol tolerance | abs(answer β gold) β€ tol | | |
| | `exact_bool` | yes/no/true/false coercion | bool(answer) β‘ bool(gold) | | |
| | `flexible` | numeric-aware partial-match | answer contains the gold value | | |
| | `list` / `list_csv` | set or ordered list comparison | elements match | | |
| | `llm_judge_long` | gpt-4o-mini judge | judge says yes | | |
| The `verified_gold_corrected` cohort has had `gold_answer` overwritten by Stage-2 cross-model consensus; the original is preserved in `manifest.parquet`'s `gold_original` column. | |
| ## Troubleshooting | |
| | Symptom | Cause | Fix | | |
| |---|---|---| | |
| | Container fails healthcheck | `HF_TOKEN` not in agent env | add `--ae HF_TOKEN="$HF_TOKEN"` | | |
| | Grader writes 0.0 but agent's answer looks right | `reward_mode_initial = numeric` and your answer included units/text | strip to numeric in `answer.txt` | | |
| | `llm_judge_long` always 0.0 | `OPENAI_API_KEY` not in verifier env | add `--ve OPENAI_API_KEY="$OPENAI_API_KEY"` | | |
| | Docker daemon OOMs at high concurrency | Per-container `--memory` Γ concurrency exceeds Docker Desktop VM | bump Docker Desktop's Memory limit (Settings β Resources) or lower `-j` | | |
| | Lots of "ghost" containers after a crash | Harbor's cleanup is best-effort on host-process death | `docker container prune -f` between runs | | |
| ## Citation | |
| ```bibtex | |
| @dataset{adithya_data_agent_rl_eval_2026, | |
| author = {Adithya S Kolavi}, | |
| title = {data_agent_rl_environment_eval: a 366-task verified data-analysis benchmark for Harbor}, | |
| year = 2026, | |
| publisher = {Hugging Face}, | |
| url = {https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_eval} | |
| } | |
| ``` | |
| ## Related | |
| - [`AdithyaSK/data_agent_rl`](https://huggingface.co/datasets/AdithyaSK/data_agent_rl) β source-of-truth eval/train split manifests (500 eval / 29055 train, parquet-only) | |
| - [`AdithyaSK/jupyter-agent-kaggle-all`](https://huggingface.co/datasets/AdithyaSK/jupyter-agent-kaggle-all) β the Kaggle bucket pulled by `pull_bucket.py` | |
| - [`AdithyaSK/harbor-visualiser`](https://huggingface.co/spaces/AdithyaSK/harbor-visualiser) β Gradio Space for browsing this dataset (and any Harbor-format dataset) | |
Xet Storage Details
- Size:
- 13.8 kB
- Xet hash:
- 4face540cd6270f6123c8cf29b3ca93203f8d0130954211ceaa44dc7d4856694
Β·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.