679 MB
252 files
Updated 8 days ago
Name
Size
data
traces.jsonl663 MB
xet
moonshiner-dataset-banner.png1.58 MB
xet
dataset-manifest.json136 kB
xet
README.md15.8 kB
xet
.gitattributes3.4 kB
xet
README.md

Claude Fable 5 Agent Traces

Moonshiner — Claude Fable 5 instruction following, tool use, and coding

2,371 TRAJECTORIES · 12,439 TRAINING ROWS · 14 MB PARQUET · 662 MB JSONL

Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces.

Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain.

This is an actively growing dataset. More is coming: additional training programs and substantially more sessions will be added to this same repo.

What makes it different

  • All real model trajectories. Every row is a cumulative prefix of a genuine Claude Fable 5 session captured end-to-end. The agent's causal exploration, tool arguments, results, corrections, and final responses are retained.

  • One next step per row. A trajectory with N assistant turns produces N rows. Row k contains the complete context through assistant turn k; that final assistant message is the sole training target.

  • Runtime-normalized. Runtime plumbing, UI decoration, control sequences, and verbose success boilerplate are removed or canonicalized while causal context remains.

  • Independently verified. Coding sessions must pass deterministic tests and protected-file checks. Instruction-following sessions must pass deterministic tool-call, staging, argument, and response-constraint checks. Every retained trajectory also clears independent review.

  • Reasoning-effort step-down. Failed trace attempts proceed through xhigh → medium → low (up to the configured attempt count) and stop at the first judge-accepted trace. If higher reasoning fails a task that lower reasoning succeeds on, the lower-effort trace is retained.

Task mix

High-level training programs, calculated from accepted trajectories using the same program mapping published in the seed catalog:

kind trajectories share row share flavor
Tool calling 694 29.3% 14.3% Select tools, construct grounded arguments, run independent calls together, and stage dependent calls.
Instruction following 556 23.5% 11.5% Honor constraints, formats, corrections, state, context, memory, relevance, and abstention.
Building 312 13.2% 13.6% Implement complete libraries, services, CLIs, workflows, and systems from specifications.
Debugging 280 11.8% 11.5% Diagnose failures, repair defects, resolve compiler/runtime issues, and preserve regressions.
Project & integration 199 8.4% 13.4% Coordinate multi-file, repository-scale, migration, and integration work.
Clarification 101 4.3% 1.6% Recognize missing information, ask only when required, and never invent parameters.
Feature development 81 3.4% 4.8% Extend working systems while preserving existing behavior.
Error recovery 75 3.2% 2.0% Recover from tool failures, partial results, retries, and idempotency hazards.
Seed authoring 49 2.1% 26.3% Accepted trajectories grouped by their explicit published dataset category.
Refactoring & performance 21 0.9% 0.9% Restructure safely and improve measured performance without behavior drift.
Security 3 0.1% 0.1% Enforce authorization, resource, path, boundary, and adversarial-input safety in defensive systems and repairs.

Languages (current drop)

English, Python, TypeScript, Go, Java, PowerShell, C#, Rust, Bash, C, C++, Ruby, Zsh, JavaScript, Assembly

Schema

Each row:

column type contents
task string stable task id
lang string English (en) or primary programming language
category string detailed recipe category
split string trajectory-disjoint train or val partition
assistant_step int 1-based target assistant turn
assistant_steps int assistant turns in the source trajectory
target_message_index int index of the final assistant target
n_messages int cumulative message count through the target
messages list of objects cumulative context ending at the target

messages is native JSON.

Layout

The Hugging Face viewer reads the validated active Parquet shards listed in dataset-manifest.json. The equivalent canonical traces.jsonl is also published for direct download and conversion. It currently contains 12,439 cumulative next-step rows derived from 2,371 accepted trajectories over disjoint train and validation tasks.

When training from the cumulative view, supervise only the final assistant message in each row. Supervising every assistant span would repeatedly overweight early steps because those spans recur as context in later prefixes.

Intended use

Supervised fine-tuning of instruction-following, tool-calling, and coding agents, plus analysis of multi-step planning, parallel calls, tool selection, state tracking, build-test-fix loops, and verification-driven completion.

Provenance

Generated with Claude Fable 5 (anthropic/claude-fable-5), then filtered by deterministic verification and explicit Codex acceptance review. Codex supplies no demonstration content; it only judges or requests replacement of candidate traces. Provider credentials, user keys, and host-identifying data are scrubbed before publication.

License

CC BY 4.0 — free for training, research, commercial products, modification, redistribution, and inclusion in other datasets or corpora, with attribution.

Suggested attribution:

Claude Fable 5 Agent Traces — https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces

Total size
679 MB
Files
252
Last updated
Jul 29
Pre-warmed CDN
US EU US EU

Contributors