Buckets:
| pretty_name: GPT-5.6 Sol Coding & Debugging Traces | |
| license: cc-by-4.0 | |
| language: | |
| - en | |
| tags: | |
| - traces | |
| - code | |
| - agentic | |
| - tool-use | |
| - coding-agent | |
| - openai | |
| - codex | |
| - codex-cli | |
| - gpt | |
| - gpt-5.6-sol | |
| - sft | |
| - agent-traces | |
| - coding-agents | |
| - reasoning | |
| - chain-of-thought | |
| - cot | |
| - cybersecurity | |
| - security | |
| - secure-coding | |
| - defensive-security | |
| - vulnerability-detection | |
| - static-analysis | |
| - code-review | |
| - owasp | |
| - cwe | |
| - data-engineering | |
| - training-pipeline | |
| - llm-training | |
| task_categories: | |
| - text-generation | |
| - question-answering | |
| size_categories: | |
| - 10K<n<100K | |
| # GPT-5.6 Sol Coding & Debugging Traces | |
| Verified software-engineering, independent model-judging, seed-authoring, | |
| defensive-security, and training-harness trajectories from | |
| **GPT-5.6 Sol** (`gpt-5.6-sol`) running through the Codex CLI as an | |
| autonomous coding agent. Sessions show the observable development loop: | |
| inspecting repositories, reproducing failures, explaining evidence, editing | |
| files, running compilers and test suites, correcting mistakes, and verifying | |
| the completed result. Security sessions additionally show secure-code reasoning, | |
| classification, remediation, vulnerability chaining, and full static repository audits. | |
| Harness sessions show how the collectors, verifiers, isolation boundary, rolling services, | |
| Git repository, dataset schema, model card, and public Hub publication were built and debugged. | |
| Seed-authoring sessions show the model translating binding use-case catalogs into exact, | |
| hermetic coding fixtures, protected oracles, and reversible reference solutions. | |
| Model-judge sessions show `gpt-5.6-sol` independently inspecting code written by | |
| other models, reproducing acceptance checks, and returning evidence-backed accept | |
| or reject verdicts. Every clean judging attempt is retained, including repeated | |
| reviews after a candidate changes. | |
| <!-- screened-snapshot:start --> | |
| **Current published snapshot:** 17,939 cumulative next-step rows from 1,474 accepted source trajectories, including 4 seed-authoring and 629 independent model-judge trajectories. This canonical file is rebuilt and replaced after each newly accepted implementation, seed-authoring, or judge trace. | |
| <!-- screened-snapshot:end --> | |
| **This is an early, actively growing dataset.** The coding program continuously | |
| adds independently screened implementation traces and now also preserves accepted | |
| seed-authoring trajectories for the binding Wave 10–14 catalog. A separate security | |
| program regenerates answer-backed use cases from `fable-secure` as fresh Codex traces | |
| and performs blind whole-repository reviews against held-out findings. These task | |
| sessions are supplemented by native, turn-bounded traces of the `gpt-5.6-sol` work | |
| that constructed and operated this training pipeline itself. | |
| ## What makes it different | |
| - **Actual agent runs, not reconstructed dialogue.** Every row comes from a | |
| live `codex exec` session operating on a real Git workspace. Tool calls, | |
| command output, edits, visible commentary, corrections, and the final answer | |
| retain their original order. No assistant turns are back-filled after the | |
| run. | |
| - **Cumulative context with one next action.** Whole normalized source | |
| trajectories are retained in the project repository, then expanded into one | |
| exact cumulative prefix per assistant action. Every published row ends at | |
| the single next assistant message to supervise; earlier actions and tool | |
| results remain context. No prefix is summarized or token-capped. | |
| - **Independently verified and reviewed.** After Codex exits, the pipeline | |
| applies its captured patch in a fresh workspace, preserves every protected | |
| test, passes the acceptance command twice, and scans recorded actions for | |
| scope violations. A separate read-only Codex session must also clear scope | |
| creep, missed requirements, regressions, bugs, and non-working code. Review | |
| decisions pin raw/diff hashes and stale decisions fail closed. | |
| - **Exact-model judge rollouts.** Standalone code-review sessions are imported | |
| from Codex's persisted session logs only when every turn attests exactly | |
| `gpt-5.6-sol`, the rollout contains one uncorrected judge prompt, its tool | |
| calls and results balance, and its final structured verdict validates. Both | |
| accepts and rejects are retained; incomplete or interactive sessions are not. | |
| - **Seed authoring is independently screened too.** A seed chunk must match its | |
| exact catalog slots, fail twice at baseline, pass twice with its production-only | |
| reference patch, reverse cleanly to the failing baseline, and survive a separate | |
| semantic review for missed or extra requirements, fixture/reference bugs, weak | |
| tests, nondeterminism, and scope creep. Reviews pin the complete seed tree plus | |
| the author prompt and raw trajectory before those trajectories enter this file. | |
| - **Hidden-reference security grading.** The 2,391 security teaching prompts are | |
| solved without their supplied answers. A separate post-run judge compares the | |
| candidate to the hidden reference, with structured CWE/OWASP presence gates on | |
| classification cases. Justified equivalent taxonomy mappings are allowed and exact | |
| keyed-label differences remain in private verdict provenance. The reference never | |
| enters the teacher workspace. | |
| - **Blind repository audits.** Eighteen pinned deliberately vulnerable projects are | |
| reviewed from source. A host-side oracle matches `findings.json` against 196 | |
| held-out file/line findings, enforces minimum recall, and rejects finding spray. | |
| - **Hostile-repository containment.** Security teachers run in disposable copies | |
| under an outer filesystem namespace plus Codex's command sandbox. The real home, | |
| answer keys, sibling repositories, and saved credential are inaccessible; the | |
| short-lived credential file is removed before the first model tool can run. | |
| - **Lossless Codex tool source.** Codex's public JSON event stream summarizes | |
| file changes, so the collector also resolves the session's persisted rollout | |
| from its documented thread id. The dataset preserves the exact outer `exec` | |
| JavaScript orchestration source—including shell and patch calls—and its | |
| returned output. | |
| - **Privacy-aware conversion.** Codex base/developer instructions, personal | |
| configuration, injected environment context, and encrypted hidden reasoning | |
| are not training targets. Workspace paths become relative, `$HOME` becomes | |
| `~`, and likely API keys or private keys cause the entire row to be rejected. | |
| - **Real toolchains and acceptance suites.** Tasks use Python, Node, | |
| TypeScript, Go, Rust, Java, Ruby, C#, Bash, Zsh, C, C++, PowerShell, and GNU | |
| x86-64 assembly toolchains. Dependency-oriented seeds use pinned project | |
| manifests; networked tasks use hermetic loopback fixtures rather than public | |
| services. | |
| ## Task program | |
| The catalog spans software construction and defensive security: | |
| | mode | examples | | |
| |---|---| | |
| | build | implement libraries, CLIs, parsers, services, games, and complete mini-projects from acceptance tests | | |
| | debug | diagnose symptom-first failures involving concurrency, state, encoding, APIs, resources, and boundaries | | |
| | feature | extend working repositories while preserving their existing behavior | | |
| | refactor / performance | restructure or optimize under pinned behavioral and measurement contracts | | |
| | security knowledge | explain, classify, detect, locate, remediate, evaluate, and apply 407 secure-review techniques | | |
| | security chains | connect exploitable primitives into realistic higher-severity paths and identify break points | | |
| | whole-repo security review | inspect pinned authorized repositories, confirm exact evidence, prioritize findings, and prescribe fixes | | |
| | seed authoring | convert binding catalog items into scoped baseline fixtures, protected acceptance contracts, and reversible reference fixes | | |
| | harness construction | design collectors/verifiers, debug isolation, configure persistent services, build datasets/cards, and verify Git/Hub publication | | |
| Domains include data structures, parsers and codecs, state machines, async and | |
| concurrency primitives, HTTP clients and servers, storage, workflow engines, | |
| dependency management, frontend components, reporting, systems programming, | |
| PowerShell administration, deterministic games, and assembly routines. | |
| ## Languages | |
| The growing corpus covers: | |
| - Python | |
| - TypeScript and JavaScript | |
| - Go | |
| - Rust | |
| - Java | |
| - Ruby | |
| - C# | |
| - Bash and Zsh | |
| - C and C++ | |
| - PowerShell | |
| - x86-64 GNU assembly | |
| Early snapshots contain only the languages whose trace batches have completed; | |
| the list above describes the committed and queued catalog. | |
| ## Schema | |
| Each line in `traces.jsonl` is one cumulative trajectory prefix ending at the | |
| assistant action identified by `assistant_step`: | |
| | column | type | contents | | |
| |---|---|---| | |
| | `task` | string | stable implementation, seed-authoring, or judge trajectory id | | |
| | `source_trajectory_id` | string | fingerprint grouping every prefix from one source row | | |
| | `lang` | string | primary implementation language, `seed-authoring`, or reviewed language | | |
| | `category` | string | task-family label such as `debug-concurrency` or `build-cli` | | |
| | `kind` | nullable string | specialized source kind such as `seed-authoring` or `model-judge` | | |
| | `domain` | string | `coding`, `security`, or `harness` | | |
| | `security_task` | nullable string | security task family; null for ordinary coding rows | | |
| | `verifier` | string | independent acceptance, hidden-reference, or held-out finding gate | | |
| | `split` | string | deterministic task-disjoint `train` or `val` partition | | |
| | `teacher_model` | string | exact teacher model id (`gpt-5.6-sol`) | | |
| | `judge_model` | nullable string | exact judging model id; always `gpt-5.6-sol` for model-judge rows | | |
| | `judge_verdict` | nullable string | `accept`, `reject`, or `needs_human` for a judge trajectory | | |
| | `reviewed_task` | nullable string | task whose model-generated code was independently judged | | |
| | `session_id` | nullable string | persisted Codex judge session identity used for deduplication | | |
| | `reasoning_effort` | string | configured Codex reasoning level | | |
| | `trace_format` | string | `codex-rollout`, or the documented event-stream fallback | | |
| | `tools_used` | list<string> | outer tools invoked during the session | | |
| | `trace_part` | int | one-based part number; `1` for unsplit task sessions | | |
| | `trace_parts` | int | total exact parts for the logical turn; `1` for unsplit task sessions | | |
| | `continuation` | bool | true when this row continues a preceding harness part | | |
| | `derivation` | string | cumulative next-assistant derivation version | | |
| | `assistant_step` | int | one-based target action within the source trajectory | | |
| | `assistant_steps` | int | total assistant actions from the source trajectory | | |
| | `target_message_index` | int | final message index and sole supervised target | | |
| | `original_n_messages` | int | message count in the untouched source trajectory | | |
| | `n_messages` | int | message count in this cumulative prefix | | |
| | `messages` | list of objects | ordered prefix ending at the target assistant action | | |
| | `tools` | string (JSON) | tool schemas declared for the row | | |
| `messages` remains native JSON. `tools` is JSON-encoded as a string because | |
| tool parameter schemas can vary across runtime versions; load it with | |
| `json.loads(row["tools"])`. | |
| Train/validation assignment happens at source-trajectory level before | |
| expansion, so related prefixes never cross splits. The two separately | |
| designated evaluation holdouts never enter this file. | |
| ## Tool-call representation | |
| Codex records `exec` as a free-form custom tool. Chat templates used by common | |
| open models expect object-valued function arguments, so the source string is | |
| represented losslessly as: | |
| ```json | |
| { | |
| "name": "exec", | |
| "arguments": { | |
| "source": "const r = await tools.exec_command(...); text(r.output);" | |
| } | |
| } | |
| ``` | |
| A fine-tuned agent needs either a compatible JavaScript orchestration runtime | |
| or an explicit transformation from this contract to its deployment tools. | |
| ## Layout | |
| The dataset ships as one file, `traces.jsonl`, with one next-assistant-action | |
| example per line. The `split` column selects the trajectory-disjoint | |
| train/validation partition. The project Git repository separately keeps the | |
| original one-row-per-source export for provenance; it is not uploaded here. | |
| ## Intended uses | |
| - Supervised fine-tuning of tool-using coding agents | |
| - Supervised fine-tuning of defensive secure-code-review agents | |
| - Training or evaluating build-test-fix behavior | |
| - Training classification, remediation, chain analysis, and evidence-backed static review | |
| - Training agents to construct, verify, publish, and operate trace-distillation harnesses | |
| - Research on long-horizon repository work and tool selection | |
| - Analysis of verification-driven completion and self-correction | |
| - Converting Codex trajectories to other agent tool protocols | |
| ## Limitations | |
| - The export is success-filtered by design and therefore contains fewer | |
| terminal failures than real-world agent traffic. | |
| - Reference-graded answers use a model judge plus deterministic structured-label | |
| checks; model judging can still introduce bias. Whole-repository reviews use the | |
| stronger location oracle, but the default passing recall threshold is 0.6 rather | |
| than a claim that every planted issue appears in every accepted trace. | |
| - The repository-audit corpus is intentionally vulnerable. Its trajectories are for | |
| authorized defensive review and do not establish that generated advice is safe to | |
| execute against third-party systems. | |
| - Tasks are purpose-built, deterministic engineering environments rather than | |
| a random sample of public repositories. | |
| - Hidden model reasoning is encrypted by the provider and excluded. The | |
| reasoning signal consists of visible commentary, hypotheses, tool choices, | |
| corrections, and any plaintext reasoning summaries the runtime emits. | |
| - The current corpus reflects one teacher model, one high reasoning setting, | |
| and the Codex CLI behavior recorded in each row's provenance fields. | |
| - Cumulative prefixes can be long. Consumers must choose context limits | |
| explicitly; truncation can remove context required for the target action. | |
| - Training must apply loss only to the final assistant action. Supervising | |
| earlier assistant spans would overweight them because they recur as context | |
| in later prefixes. | |
| - Users remain responsible for reviewing generated code and applicable | |
| third-party licenses before deploying a trained model. | |
| ## Provenance and reproducibility | |
| The collection pipeline retains its complete seed, raw-trace, verification, | |
| and final-diff provenance in the project repository. Each task begins as | |
| `tasks/seeds/<id>/{task.json,files/,reference_fix.patch}`. The reference patch | |
| proves solvability but is never placed in the teacher workspace. Codex runs | |
| against a fresh baseline commit; the pipeline then records its final diff, | |
| replays verification twice, checks protected-file hashes and recorded scope, | |
| requires an independent acceptance review, scrubs the accepted rollout, and | |
| deterministically assigns the source trajectory split before expansion. | |
| Accepted seed-authoring sources additionally retain the exact catalog prompt, | |
| author and reviewer event streams, reversible validation evidence, per-requirement | |
| semantic verdict, and hashes for every byte in every authored seed directory. | |
| Imported model-judge sources retain the byte-exact persisted Codex rollout and | |
| a hash-pinned manifest. Admission is derived solely from the session log: exact | |
| model attestation, single-purpose prompt shape, completion event, balanced tools, | |
| and schema-valid verdict. Source repositories are not consulted during import. | |
| Harness-construction provenance is pinned by Codex thread id. Its full active rollout is | |
| snapshotted only to the private project repository. Source conversion emits a completed | |
| logical user turn as one source row, or as exact numbered continuations when the turn | |
| exceeds the local-model context target; cumulative expansion then derives the public | |
| next-action examples. It excludes provider base/developer messages, environment-only | |
| messages, encrypted reasoning, machine paths, likely secrets, and unfinished turns. | |
| Teacher configuration for this release: | |
| - model: `gpt-5.6-sol` | |
| - reasoning effort: `xhigh` | |
| - interface: non-interactive Codex CLI JSONL | |
| - coding workspace policy: unrestricted execution inside trusted disposable seed workspaces | |
| - security workspace policy: Bubblewrap filesystem isolation plus Codex workspace sandbox; | |
| static inspection only for deliberately vulnerable repositories | |
| - user configuration and exec rules: ignored for capture consistency | |
| ## License | |
| [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) — use, modify, | |
| redistribute, and train on the dataset with attribution. | |
| Suggested attribution: | |
| > GPT-5.6 Sol Coding & Debugging Traces — | |
| > https://huggingface.co/datasets/greghavens/gpt-5.6-sol-coding-and-debugging-traces | |
Xet Storage Details
- Size:
- 17 kB
- Xet hash:
- 04e7f5945bcc85eed274688faf749d830dec0604b09259d0346c5981c5794a1c
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.