| pretty_name: Claude Fable 5 Agent Traces | |
| license: cc-by-4.0 | |
| language: | |
| - en | |
| annotations_creators: | |
| - machine-generated | |
| task_categories: | |
| - text-generation | |
| size_categories: | |
| - 10K<n<100K | |
| tags: | |
| - "traces" | |
| - "code" | |
| - "agentic" | |
| - "tool-use" | |
| - "coding-agent" | |
| - "coding-agents" | |
| - "agent-traces" | |
| - "instruction-following" | |
| - "function-calling" | |
| - "parallel-tool-calling" | |
| - "multi-turn" | |
| - "sft" | |
| - "distillation" | |
| - "reasoning" | |
| - "chain-of-thought" | |
| - "cot" | |
| # Claude Fable 5 Agent Traces | |
|  | |
| <div align="center"> | |
| <h2>2,161 TRAJECTORIES · 11,235 TRAINING ROWS · 11 MB PARQUET · 656 MB JSONL</h2> | |
| </div> | |
| > Generated by **[moonshiner](https://github.com/greghavens/moonshiner)** — an open harness for | |
| > distilling verified instruction-following, tool-use, and agentic coding traces. | |
| Behavior-preserving **instruction-following, tool-use, and agent trajectories** | |
| from **Claude Fable 5** (`anthropic/claude-fable-5`). The category and row-share tables | |
| below describe the actual mix seen during training rather than assuming a | |
| particular task domain. | |
| **This is an actively growing dataset. More is coming**: additional training | |
| programs and substantially more sessions will be added to this same repo. | |
| ## What makes it different | |
| - **All real model trajectories.** Every row is a cumulative prefix of a genuine | |
| Claude Fable 5 session captured end-to-end. The agent's causal exploration, tool | |
| arguments, results, corrections, and final responses are retained. | |
| - **One next step per row.** A trajectory with N assistant turns produces N | |
| rows. Row k contains the complete context through assistant turn k; that final | |
| assistant message is the sole training target. | |
| - **Runtime-normalized.** Runtime plumbing, UI decoration, control sequences, | |
| and verbose success boilerplate are removed or canonicalized while causal | |
| context remains. | |
| - **Independently verified.** Coding sessions must pass deterministic tests and | |
| protected-file checks. Instruction-following sessions must pass deterministic | |
| tool-call, staging, argument, and response-constraint checks. Every retained | |
| trajectory also clears independent review. | |
| - **Reasoning-effort step-down.** Failed trace attempts proceed through `xhigh → medium → low` (up to the configured attempt count) and stop at the first judge-accepted trace. If higher reasoning fails a task that lower reasoning succeeds on, the lower-effort trace is retained. | |
| ## Task mix | |
| High-level training programs, calculated from accepted trajectories using the | |
| same program mapping published in the seed catalog: | |
| | kind | trajectories | share | row share | flavor | | |
| |---|---:|---:|---:|---| | |
| | Tool calling | 640 | 29.6% | 12.4% | Select tools, construct grounded arguments, run independent calls together, and stage dependent calls. | | |
| | Instruction following | 545 | 25.2% | 12.4% | Honor constraints, formats, corrections, state, context, memory, relevance, and abstention. | | |
| | Building | 286 | 13.2% | 14.0% | Implement complete libraries, services, CLIs, workflows, and systems from specifications. | | |
| | Debugging | 230 | 10.6% | 10.5% | Diagnose failures, repair defects, resolve compiler/runtime issues, and preserve regressions. | | |
| | Project & integration | 138 | 6.4% | 11.6% | Coordinate multi-file, repository-scale, migration, and integration work. | | |
| | Clarification | 101 | 4.7% | 1.7% | Recognize missing information, ask only when required, and never invent parameters. | | |
| | Feature development | 81 | 3.7% | 5.3% | Extend working systems while preserving existing behavior. | | |
| | Error recovery | 70 | 3.2% | 1.9% | Recover from tool failures, partial results, retries, and idempotency hazards. | | |
| | Seed authoring | 49 | 2.3% | 29.1% | Accepted trajectories grouped by their explicit published dataset category. | | |
| | Refactoring & performance | 21 | 1.0% | 1.0% | Restructure safely and improve measured performance without behavior drift. | | |
| ## Languages (current drop) | |
| `English`, `Python`, `TypeScript`, `Go`, `PowerShell`, `Java`, `Rust`, `C#`, `Ruby`, `Bash`, `Zsh`, `JavaScript`, `C`, `C++` | |
| ## Schema | |
| Each row: | |
| | column | type | contents | | |
| |---|---|---| | |
| | `task` | string | stable task id | | |
| | `lang` | string | English (`en`) or primary programming language | | |
| | `category` | string | detailed recipe category | | |
| | `split` | string | trajectory-disjoint `train` or `val` partition | | |
| | `assistant_step` | int | 1-based target assistant turn | | |
| | `assistant_steps` | int | assistant turns in the source trajectory | | |
| | `target_message_index` | int | index of the final assistant target | | |
| | `n_messages` | int | cumulative message count through the target | | |
| | `messages` | list of objects | cumulative context ending at the target | | |
| `messages` is native JSON. | |
| ## Layout | |
| The Hugging Face viewer reads the validated active Parquet shards listed in `dataset-manifest.json`. The equivalent canonical `traces.jsonl` is also published for direct download and conversion. It currently contains | |
| 11,235 cumulative next-step rows derived from 2,161 accepted | |
| trajectories over disjoint train and validation tasks. | |
| When training from the cumulative view, supervise only the final assistant | |
| message in each row. Supervising every assistant span would repeatedly | |
| overweight early steps because those spans recur as context in later prefixes. | |
| ## Intended use | |
| Supervised fine-tuning of instruction-following, tool-calling, and coding | |
| agents, plus analysis of multi-step planning, parallel calls, tool selection, | |
| state tracking, build-test-fix loops, and verification-driven completion. | |
| ## Provenance | |
| Generated with Claude Fable 5 (`anthropic/claude-fable-5`), then filtered by deterministic verification and | |
| explicit Codex acceptance review. Codex supplies no demonstration content; it | |
| only judges or requests replacement of candidate traces. Provider credentials, | |
| user keys, and host-identifying data are scrubbed before publication. | |
| ## License | |
| [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) — free for training, | |
| research, commercial products, modification, redistribution, and inclusion in | |
| other datasets or corpora, with attribution. | |
| Suggested attribution: | |
| > Claude Fable 5 Agent Traces — | |
| > https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces | |
Xet Storage Details
- Size:
- 6.38 kB
- Xet hash:
- 6891583fb87642663506b1b403612dd15469e7c05f634787697e44de008924d9
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.