Buckets:
| license: apache-2.0 | |
| language: | |
| - en | |
| size_categories: | |
| - 1K<n<10K | |
| tags: | |
| - reinforcement-learning | |
| - data-science | |
| - code-agent | |
| - jupyter | |
| - harbor | |
| - training-data | |
| - sft | |
| [](https://huggingface.co/spaces/AdithyaSK/harbor-visualiser?dataset=AdithyaSK/data_agent_rl_environment_train) | |
| # data_agent_rl_environment_train | |
| **The official verified training suite for the data-agent RL pipeline.** | |
| 2238 Harbor-format data-analysis tasks, each with: | |
| - An LLM-assigned difficulty label (L1-L5) | |
| - A Kaggle dataset dependency (fetched at container start) | |
| - A tested reward function | |
| This is the **training-data counterpart** to | |
| [`AdithyaSK/data_agent_rl_environment_eval`](https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_eval). | |
| For your held-out eval split, use that one. | |
| > ๐ก **Browse in your browser** โ click the badge above or open | |
| > [`AdithyaSK/harbor-visualiser`](https://huggingface.co/spaces/AdithyaSK/harbor-visualiser?dataset=AdithyaSK/data_agent_rl_environment_train) | |
| > to inspect every task's spec, instruction, environment, tests, and difficulty. | |
| ## Why "training" vs "eval" | |
| | | This dataset (`_train`) | Eval (`_eval`) | | |
| |---|---|---| | |
| | Pipeline run | **Stage 1 only** (Sonnet anchor + categorize on pass) | Stage 1 + Stage 2 (doctor rescue) | | |
| | Verdicts | 100% pure `verified` | mix of `verified` + `gold_corrected` + `verifiable_judge` + `verified_after_rewrite` | | |
| | Pass rate of attempted pool | ~45% (cheap, high signal-quality) | ~73% (expensive, broader coverage) | | |
| | Per verified-task cost | **~$0.17** | ~$0.20 | | |
| | Intended use | SFT / RL training | held-out eval, benchmarking | | |
| The "Stage 1 only" choice for training data is deliberate: a clean `verified` | |
| verdict means the agent (Sonnet) passed against the **original** gold without | |
| any doctor-driven rewrite. That's exactly the signal you want for SFT/RL โ | |
| the gold answer is canonical, no learner gets confused by post-hoc gold | |
| corrections. | |
| ## Production stats | |
| - **Pool**: stratified sample from `AdithyaSK/data_agent_rl`'s 29k-task train split | |
| - **Stratification**: by `(reward_mode_initial ร package_tier)`, seed=42 (batch 1) & seed=43 (batch 2) | |
| - **Bucket-covered**: all task Kaggle datasets exist in [`AdithyaSK/jupyter-agent-kaggle-all`](https://huggingface.co/datasets/AdithyaSK/jupyter-agent-kaggle-all) | |
| - **Attempted**: 4990 tasks (across two Sonnet+seta sweeps) | |
| - **Verified**: **2238 (45% pass rate)** | |
| - **Total spend across attempted pool**: $376.68 (4990 tasks) | |
| - **Per task attempted** (Stage-1-only, amortized): $0.0755 | |
| - **Per verified task** (cost-to-produce, amortized): $0.1683 ($141.75 spent on the 2238 successes + $234.93 spent on the 2752 failures/drops that you have to attempt to find the successes) | |
| ## Difficulty distribution | |
| | Level | Count | % | | |
| |---|---:|---:| | |
| | **L0** | 1 | 0% | | |
| | **L1** | 544 | 24% | | |
| | **L2** | 989 | 44% | | |
| | **L3** | 335 | 14% | | |
| | **L4** | 358 | 15% | | |
| | **L5** | 11 | 0% | | |
| | Level | Typical pattern | | |
| |---|---| | |
| | L1 | one-line filter / aggregation | | |
| | L2 | filter + groupby + aggregate (2-4 turns) | | |
| | L3 | multi-step pandas, joins, light feature work | | |
| | L4 | ML training, complex stats, non-trivial pipelines | | |
| | L5 | extreme complexity (rare) | | |
| Categorize was an LLM rubric (Sonnet) reading each passing trajectory. | |
| ## Layout | |
| ``` | |
| tasks/ | |
| โโโ <task_dir>/ # e.g. 0114_986_114986805_qa_2 | |
| โโโ task.toml # Harbor task spec โ gold_answer, reward_mode, difficulty_level | |
| โโโ instruction.md # natural-language question | |
| โโโ environment/ | |
| โ โโโ Dockerfile # container image | |
| โ โโโ pull_bucket.py # fetches task's Kaggle slice at startup | |
| โโโ tests/ | |
| โโโ test.sh # verifier entrypoint | |
| โโโ grader.py # mode-aware grader | |
| manifest.parquet # per-task: task_id, verdict, difficulty, gold, kaggle, question, cost, trials | |
| registry.json # Harbor visualizer index | |
| ``` | |
| ## Reproduce a task end-to-end | |
| ```bash | |
| # Prereqs | |
| pip install harbor | |
| export HF_TOKEN=hf_... # to fetch the Kaggle bucket | |
| export ANTHROPIC_API_KEY=sk-ant-... # or your model of choice | |
| export OPENAI_API_KEY=sk-... # only for tasks whose reward_mode_initial == 'llm_judge_long' | |
| # Download (just one task as a smoke test) | |
| python -c " | |
| from huggingface_hub import snapshot_download | |
| snapshot_download( | |
| repo_id='AdithyaSK/data_agent_rl_environment_train', repo_type='dataset', | |
| local_dir='./train', allow_patterns=['tasks/0114_986_114986805_qa_2/**'], | |
| )" | |
| # Run one task with the bash-only reference agent + Docker | |
| harbor run \ | |
| -p ./train/tasks \ | |
| -i 0114_986_114986805_qa_2 \ | |
| --env docker \ | |
| --agent-import-path my_agent:BashOnlyAgent \ | |
| --model anthropic/claude-sonnet-4-6 \ | |
| --ae HF_TOKEN="$HF_TOKEN" \ | |
| --ae ANTHROPIC_API_KEY="$ANTHROPIC_API_KEY" \ | |
| --ve OPENAI_API_KEY="$OPENAI_API_KEY" \ | |
| --yes -n 1 --jobs-dir ./jobs | |
| ``` | |
| `manifest.parquet` is the easiest entry point for filtering: | |
| ```python | |
| import pandas as pd | |
| df = pd.read_parquet('hf://datasets/AdithyaSK/data_agent_rl_environment_train/manifest.parquet') | |
| # only L3-L5 numeric tasks | |
| sub = df[(df.difficulty_level >= 3) & (df.reward_mode_initial == 'numeric')] | |
| ``` | |
| ## Reward modes | |
| Each task's `task.toml` declares `reward_mode_initial` in `[metadata]`: | |
| | Mode | Logic | Pass condition | | |
| |---|---|---| | |
| | `exact_short` | string equality (case-folded, stripped) | answer โก gold | | |
| | `numeric` | float parse + atol/rtol tolerance | abs(answer โ gold) โค tol | | |
| | `exact_bool` | yes/no/true/false coercion | bool(answer) โก bool(gold) | | |
| | `flexible` | numeric-aware partial-match | answer contains the gold value | | |
| | `list` / `list_csv` | set or ordered list comparison | elements match | | |
| | `llm_judge_long` | gpt-4o-mini judge | judge says yes | | |
| ## Citation | |
| ```bibtex | |
| @dataset{adithya_data_agent_rl_train_2026, | |
| author = {Adithya S Kolavi}, | |
| title = {data_agent_rl_environment_train: a 2238-task verified training suite for data-agent RL}, | |
| year = 2026, | |
| publisher = {Hugging Face}, | |
| url = {https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_train} | |
| } | |
| ``` | |
| ## Related | |
| - [`AdithyaSK/data_agent_rl_environment_eval`](https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_eval) โ matching held-out eval (366 tasks, Stage 1 + Stage 2) | |
| - [`AdithyaSK/data_agent_rl`](https://huggingface.co/datasets/AdithyaSK/data_agent_rl) โ source-of-truth train/eval split manifest (~29k train, ~500 eval) | |
| - [`AdithyaSK/jupyter-agent-kaggle-all`](https://huggingface.co/datasets/AdithyaSK/jupyter-agent-kaggle-all) โ Kaggle bucket pulled at container start | |
| - [`AdithyaSK/harbor-visualiser`](https://huggingface.co/spaces/AdithyaSK/harbor-visualiser) โ Gradio Space for browsing this dataset | |
Xet Storage Details
- Size:
- 7 kB
- Xet hash:
- 8a3c806c15dfdf58955b35eeb02e8061646577adc27914d77615bedf26efebc1
ยท
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.