Buckets:
| license: mit | |
| language: | |
| - en | |
| - multilingual | |
| task_categories: | |
| - text-generation | |
| - other | |
| tags: | |
| - distillation | |
| - instruction-tuning | |
| - sft | |
| - reasoning | |
| - coding | |
| - code-repositories | |
| - cybersecurity | |
| - attack | |
| - defense | |
| - exploit | |
| - penetration-testing | |
| - red-team | |
| - blue-team | |
| - open-source | |
| - collection | |
| - fable-5 | |
| - gpt-5.5 | |
| - claude | |
| - gemini | |
| - grok | |
| - kimi | |
| - deepseek | |
| - Trace | |
| - qwen | |
| - biology | |
| - science | |
| - Llm | |
| - Open-source | |
| - Math | |
| - cyber | |
| - security | |
| - cyber-security | |
| - cyber security | |
| size_categories: | |
| - 10M<n<100M | |
| pretty_name: "The Open Distillation Codex" | |
| configs: | |
| - config_name: default | |
| data_files: | |
| - split: train | |
| path: | |
| - "data/applied/*/*.jsonl" | |
| - "data/coding/*/*.jsonl" | |
| - "data/cybersecurity/high_quality_cybersecurity/shard-*.jsonl" | |
| - "data/cybersecurity/clydeiii_cybersecurity/shard-*.jsonl" | |
| - "data/cybersecurity/fenrir_v2_1/shard-*.jsonl" | |
| - "data/cybersecurity/precinct6_cybersecurity/shard-*.jsonl" | |
| - "data/cybersecurity/savani_cyber_attack/shard-*.jsonl" | |
| - "data/distilled/*/*.jsonl" | |
| - "data/humanities/*/*.jsonl" | |
| - "data/index/*/*.jsonl" | |
| - "data/instruction/*/*.jsonl" | |
| - "data/science/*/*.jsonl" | |
| <div align="center"> | |
| <img src="https://img.shields.io/badge/Version-8.2-blue?style=for-the-badge" alt="Version"> | |
| <img src="https://img.shields.io/badge/Storage-76GB%2B-green?style=for-the-badge" alt="Storage"> | |
| <img src="https://img.shields.io/badge/Sources-73-orange?style=for-the-badge" alt="Sources"> | |
| <img src="https://img.shields.io/badge/License-MIT-yellow?style=for-the-badge" alt="License"> | |
| <img src="https://img.shields.io/badge/Samples-18M%2B-red?style=for-the-badge" alt="Samples"> | |
| <img src="https://img.shields.io/badge/Cybersecurity-6%20Sources-purple?style=for-the-badge" alt="Cybersecurity"> | |
| <br><br> | |
| # ๐ The Open Distillation Codex | |
| ### ๐ *The Ultimate Open-Source Distillation Dataset โ No Skip, Full, with Attack & Defense* ๐ | |
| **Where 73 open-source minds converge into one unified stream of intelligence** | |
| `18M+ Distilled Signals` ยท `7,090 Raw GitHub Repositories` ยท `8 Curated Categories` ยท `~76 GB+` | |
| <br> | |
| > *"We did not write this dataset. We assembled it.* | |
| > *Every line is an echo โ of a model thinking, a coder drafting, a tutor explaining, a repo breathing.* | |
| > *Seventy-three sources. Eight categories. Zero gatekeeping. No skipping. Fully processed. Now fortified with real-world cybersecurity confrontations."* | |
| <br> | |
| </div> | |
| --- | |
| ## ๐ Table of Contents | |
| | # | Section | Description | | |
| |---|---|---| | |
| | 1 | [๐ Dataset Summary](#-dataset-summary) | High-level overview & value proposition | | |
| | 2 | [๐๏ธ Directory Structure](#๏ธ-directory-structure) | ASCII tree + folder explanation | | |
| | 3 | [๐ Data Sources](#-data-sources--provenance) | All 73 sources with full attribution | | |
| | 4 | [๐ก๏ธ Cybersecurity Deep Dive: Attack & Defense](#๏ธ-cybersecurity-deep-dive-attack--defense) | Importance, attack traces, defense, exploit analysis | | |
| | 5 | [๐ ๏ธ How to Use & Train](#๏ธ-how-to-use--train) | Loading, streaming, training scripts | | |
| | 6 | [๐ Licensing & Limitations](#-licensing--limitations) | License, intended use, limitations | | |
| | 7 | [๐ Changelog](#-changelog) | Version history | | |
| --- | |
| ## ๐ Dataset Summary | |
| <div align="center"> | |
| ### ๐ฏ The Numbers That Matter | |
| | Metric | Value | Status | | |
| |:---:|:---:|:---:| | |
| | **Total Storage** | `76 GB+` | โ Verified | | |
| | **JSONL Data Shards** | `516` | โ Verified | | |
| | **Archive Files (tar.gz)** | `7,090` | โ Verified | | |
| | **Source Datasets** | `73` | โ Verified | | |
| | **Categories** | `8` | โ Verified | | |
| | **Total Samples** | `18M+` | โ Verified | | |
| | **Largest Source** | `8.15M` (Vibe-Coding-Instruct-V2) | โ | | |
| | **Archive Size** | `~64 GB` (compressed GitHub repos) | โ | | |
| | **Cybersecurity Sources** | `6` | โ | | |
| | **Cybersecurity Data Size** | `~2.6 GB` | โ | | |
| </div> | |
| <br> | |
| ### ๐ Why "Ultimate Distilled"? | |
| This dataset is not a raw scrape. Every sample has been **distilled through a unified extraction pipeline**: | |
| ``` | |
| โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ | |
| โ UNIFIED EXTRACTION PIPELINE โ | |
| โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค | |
| โ โ | |
| โ 73 Upstream Sources (ALL FULLY PROCESSED, NO SKIP) โ | |
| โ โโโโโโโ โโโโโโโ โโโโโโโ โโโโโโโ โโโโโโโ โโโโโโโ โ | |
| โ โ HF โ โ HF โ โ HF โ โ GH โ โ HF โ โ ... โ โ | |
| โ โโโโฌโโโ โโโโฌโโโ โโโโฌโโโ โโโโฌโโโ โโโโฌโโโ โโโโฌโโโ โ | |
| โ โ โ โ โ โ โ โ | |
| โ โโโโโโโโโดโโโโโโโโดโโโโโโโโผโโโโโโโโดโโโโโโโโ โ | |
| โ โ โ | |
| โ โโโโโโผโโโโโ โ | |
| โ โ EXTRACT โ โ Field normalization โ | |
| โ โโโโโโฌโโโโโ (instruction/response) โ | |
| โ โ โ | |
| โ โโโโโโผโโโโโ โ | |
| โ โCATEGORIZEโ โ 8 semantic categories โ | |
| โ โโโโโโฌโโโโโ โ | |
| โ โ โ | |
| โ โโโโโโผโโโโโ โ | |
| โ โ SHARD โ โ 20K samples per shard โ | |
| โ โโโโโโฌโโโโโ โ | |
| โ โ โ | |
| โ โโโโโโผโโโโโ โ | |
| โ โ UPLOAD โ โ Batch commits to HF โ | |
| โ โโโโโโโโโโโ โ | |
| โ โ | |
| โ STATUS: ALL 73 SOURCES COMPLETE. NO SKIPPING. 18M+ ROWS. โ | |
| โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ | |
| ``` | |
| <br> | |
| ### ๐ Value to the Open-Source AI Community | |
| | ๐ฏ For... | ๐ฆ This dataset provides... | | |
| |---|---| | |
| | **Model Trainers** | Single `load_dataset()` call to stream 18M+ SFT-ready samples | | |
| | **Coding Agent Researchers** | 11M+ agentic coding traces from Fable-5, Vibe-Coding, Royal Ghost, Kimi, DeepSeek | | |
| | **Code Pretraining** | 7,090 full GitHub repository snapshots (64 GB compressed) | | |
| | **Reasoning Researchers** | 2.7M+ distilled reasoning traces from Claude, Gemini, Grok, GPT-5.5, Opus 4.8 | | |
| | **Domain Specialists** | 25K-sample sweeps across 29 disciplines | | |
| | **Cybersecurity Researchers** | Dedicated cybersecurity category with attack/defense/exploit traces, red/blue team dialogues, and incident reports | | |
| | **Red Team / Blue Team Trainers** | Realistic attack scenarios, defense strategies, exploit code, and post-mortem analysis | | |
| --- | |
| ## ๐๏ธ Directory Structure | |
| ``` | |
| ๐ Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset/ | |
| โ | |
| โโโ ๐ฆ archives/ # ~64 GB โ 7,090 compressed GitHub repos | |
| โ โโโ 0-chi__sonaure-lp.tar.gz | |
| โ โโโ 00MB__bitcoin_trading_bot.tar.gz | |
| โ โโโ 0101-agents__plugins.tar.gz | |
| โ โโโ ... (7,090 files total) | |
| โ โโโ zznmg1__playable-survivor-ad.tar.gz | |
| โ | |
| โโโ ๐ data/ # ~12 GB โ 516 JSONL shards (18M+ samples) | |
| โ โ | |
| โ โโโ ๐ป coding/ # 28 sources ยท ~11M+ samples | |
| โ โ โโโ vibe_instruct_v2/ # 8,152,510 samples | |
| โ โ โโโ fable5_2m/ # 2,006,487 samples | |
| โ โ โโโ vibe_instruct_v1/ # 1,100,000 samples | |
| โ โ โโโ vibe_coding/ # 1,100,000 samples | |
| โ โ โโโ royal_ghost_1m/ # 1,000,000 samples | |
| โ โ โโโ citation_ground/ # 980,064 samples | |
| โ โ โโโ royal_ghost_501k/ # 703,449 samples | |
| โ โ โโโ fable5_repos_full/ # 7,090 archive pointers | |
| โ โ โโโ fable5_agentic_sft/ # 159,972 samples | |
| โ โ โโโ gpt55_codex/ # 119,436 samples โญ FULL | |
| โ โ โโโ alpca_gpt55/ # 49,099 samples | |
| โ โ โโโ deepseek_v4_pro_agent/ # 96,597 samples โญ FULL | |
| โ โ โโโ fable5_traces/ # 49,544 samples โญ FULL | |
| โ โ โโโ genesis_code_100k/ # 68,000 samples | |
| โ โ โโโ genesis_code/ # 49,000 samples | |
| โ โ โโโ kimi_coding/ # 9,014 samples | |
| โ โ โโโ mimo_claude_code_traces/ # 15,046 samples โญ FULL | |
| โ โ โโโ kimi_k26_claude_code_traces/ # 7,438 samples | |
| โ โ โโโ genesis_code_10k/ # 9,800 samples | |
| โ โ โโโ legend_python/ # 5,000 samples | |
| โ โ โโโ autonomy/ # 10,000 samples | |
| โ โ โโโ genesis_code_demo/ # 1,000 samples | |
| โ โ โโโ god_coder/ # โญ FULL raw recovery | |
| โ โ โโโ python_god_coder/ # โญ FULL raw recovery | |
| โ โ โโโ elite_god_coder/ # โญ FULL raw recovery | |
| โ โ โโโ omega_genesis/ # โญ FULL raw recovery | |
| โ โ โโโ open_tool_trace/ # 48 samples | |
| โ โ โโโ genesis_v11/ # partial recovery | |
| โ โ | |
| โ โโโ ๐งฎ math/ # 2 sources | |
| โ โ โโโ math_25k/ | |
| โ โ โโโ deepseek_prover_v1/ # 27,503 Lean theorem proofs | |
| โ โ | |
| โ โโโ ๐ฌ science/ # 7 sources | |
| โ โ โโโ science_25k/ | |
| โ โ โโโ physics_25k/ | |
| โ โ โโโ chemistry_25k/ | |
| โ โ โโโ biology_25k/ | |
| โ โ โโโ medical_25k/ | |
| โ โ โโโ cs_25k/ | |
| โ โ โโโ biology_r2med/ # โญ NEW | |
| โ โ | |
| โ โโโ โ๏ธ applied/ # 8 sources | |
| โ โ โโโ robotics_25k/ | |
| โ โ โโโ nano_25k/ | |
| โ โ โโโ materials_25k/ | |
| โ โ โโโ earth_climate_25k/ | |
| โ โ โโโ renewable_energy_25k/ | |
| โ โ โโโ evolution_25k/ | |
| โ โ โโโ universe_25k/ | |
| โ โ โโโ kardashev_25k/ | |
| โ โ | |
| โ โโโ ๐ humanities/ # 8 sources | |
| โ โ โโโ psychology_25k/ | |
| โ โ โโโ economics_25k/ | |
| โ โ โโโ law_25k/ | |
| โ โ โโโ statistics_25k/ | |
| โ โ โโโ sports_25k/ | |
| โ โ โโโ human_25k/ | |
| โ โ โโโ conscience_25k/ | |
| โ โ โโโ supernatural_25k/ | |
| โ โ | |
| โ โโโ ๐ง distilled/ # 9 sources ยท frontier distillations | |
| โ โ โโโ claude_mythos/ | |
| โ โ โโโ gemini35/ | |
| โ โ โโโ fable5_cleaned/ | |
| โ โ โโโ grok44/ | |
| โ โ โโโ gemini_pro32/ | |
| โ โ โโโ gpt55_thinking/ | |
| โ โ โโโ gpt55_distilled/ | |
| โ โ โโโ claude_opus_48_distill/ # โญ NEW | |
| โ โ โโโ claude_opus_48_max_thinking/ # โญ NEW | |
| โ โ | |
| โ โโโ ๐ instruction/ # 3 sources | |
| โ โ โโโ alpaca/ # 52,002 samples | |
| โ โ โโโ oasst/ # 32,141 samples | |
| โ โ โโโ dolly/ # 15,011 samples | |
| โ โ | |
| โ โโโ ๐ cybersecurity/ # 6 sources | |
| โ โ โโโ high_quality_cybersecurity/ | |
| โ โ โโโ heimdall_v1_1/ # โญ NEW โ 78 MB conversations | |
| โ โ โโโ fenrir_v2_1/ # โญ NEW โ 411 MB (2.1M+ entries) | |
| โ โ โโโ clydeiii_cybersecurity/ # โญ NEW โ 20 MB yearly corpus | |
| โ โ โโโ precinct6_cybersecurity/ # โญ NEW โ 2.1 GB (graph+signals+ref) | |
| โ โ โโโ savani_cyber_attack/ # โญ NEW โ 17 MB attack CSV | |
| โ โ | |
| โ โโโ ๐ index/ # 2 sources | |
| โ โโโ species_25k/ | |
| โ โโโ transport_25k/ | |
| โ | |
| โโโ ๐ README.md | |
| โโโ ๐ dataset_info.json | |
| ``` | |
| ### ๐ค Why is `archives/` kept compressed? | |
| | Reason | Explanation | | |
| |---|---| | |
| | **๐พ Space Efficiency** | Uncompressed would exceed 200+ GB. Compressed = 64 GB (3ร saving) | | |
| | **๐ฏ On-Demand Access** | Download only specific repositories you need | | |
| | **๐ Preservation Fidelity** | tar.gz preserves exact file permissions, directory structure, binaries | | |
| > ๐ก **Tip**: For training on code content, use `data/coding/fable5_repos_full/` (475K samples, each a file extracted from archives, capped at 4KB). For full untruncated file access, stream directly from `archives/`. | |
| --- | |
| ## ๐ Data Sources & Provenance | |
| <div align="center"> | |
| ### ๐บ๏ธ 73 Sources Across 8 Categories | |
| | Category | Sources | Samples | Description | | |
| |:---:|:---:|:---:|:---| | |
| | ๐ป `coding` | 28 | ~11M+ | Agentic traces, code repos, coder distillations | | |
| | ๐ง `distilled` | 9 | ~200K | Frontier model distillations | | |
| | โ๏ธ `applied` | 8 | ~200K | Robotics, nano, materials, climate, energy | | |
| | ๐ `humanities` | 8 | ~200K | Psychology, economics, law, statistics | | |
| | ๐ฌ `science` | 7 | ~175K | Physics, chemistry, biology, medical, CS | | |
| | ๐ `instruction` | 3 | ~99K | Classic instruction (alpaca, oasst, dolly) | | |
| | ๐ `index` | 2 | ~50K | Species index, transport | | |
| | ๐ `cybersecurity` | 6 | ~2.6 GB | High-quality attack, defense, exploit traces | | |
| | ๐งฎ `math` | 2 | ~52K | Math + Lean theorem proofs | | |
| </div> | |
| <br> | |
| ### ๐ป Coding Category (28 sources โ ALL FULLY PROCESSED โญ) | |
| | Source Slug | Upstream Dataset | Type | Samples | | |
| |---|---|---|---:| | |
| | `vibe_instruct_v2` | `CodeDevX/Vibe-Coding-Instruct-V2` | Agentic coding | 8,152,510 | | |
| | `fable5_2m` | `Crownelius/Complete-FABLE.5-traces-2M` | Fable-5 traces | 2,006,487 | | |
| | `vibe_instruct_v1` | `CodeDevX/Vibe-Coding-Instruct` | Agentic coding | 1,100,000 | | |
| | `vibe_coding` | `attentionAllYouNeed/Vibe-Coding-Claude-Fable-5` | Claude coding | 1,100,000 | | |
| | `royal_ghost_1m` | `WithinUsAI/Royal_Ghost_Coder_1M` | Ghost coder | 1,000,000 | | |
| | `citation_ground` | `WithinUsAI/CitationGround-1M` | Citation-grounded | 980,064 | | |
| | `royal_ghost_501k` | `WithinUsAI/Royal_Ghost_Coder_501k` | Ghost coder | 703,449 | | |
| | `fable5_repos_full` | `notune/fable5-repos` | 7,090 repo pointers | 7,090 | | |
| | `fable5_agentic_sft` | `Nexlab/fable5-agentic-coding-sft` | Agentic SFT | 159,972 | | |
| | `gpt55_codex` | `AletheiaResearch/GPT-5.5-Codex` | GPT-5.5 Codex | 119,436 | | |
| | `alpca_gpt55` | `GabrielFreeze-2/alpca-mlt-gpt-5.5_chatml` | GPT-5.5 chatml | 49,099 | | |
| | `deepseek_v4_pro_agent` | `TeichAI/DeepSeek-v4-Pro-Agent` | DeepSeek v4 | 96,597 | | |
| | `fable5_traces` | `Glint-Research/Fable-5-traces` | Fable-5 traces | 49,544 | | |
| | `genesis_code_100k` | `WithinUsAI/Genesis_AI_Code_100k` | Genesis code | 68,000 | | |
| | `genesis_code` | `WithinUsAI/Genesis_AI_Code_50k` | Genesis code | 49,000 | | |
| | `kimi_coding` | `trjxter/Kimi-K2.7-CodingTraces-9000x` | Kimi K2.7 | 9,014 | | |
| | `mimo_claude_code_traces` | `choucsan/mimo-claude-code-traces-1k` | Mimo Claude | 15,046 | | |
| | `kimi_k26_claude_code_traces` | `armand0e/kimi-k2.6-claude-code-traces` | Kimi K2.6 | 7,438 | | |
| | `genesis_code_10k` | `WithinUsAI/Genesis_AI_Code_10k` | Genesis code | 9,800 | | |
| | `legend_python` | `WithinUsAI/Legend_Python_CoderV.1` | Python coder | 5,000 | | |
| | `autonomy` | `WithinUsAI/The_Autonomy_From_WithIn_10k` | Autonomy | 10,000 | | |
| | `genesis_code_demo` | `WithinUsAI/Genesis_AI_Code_1k_Demo` | Genesis demo | 1,000 | | |
| | `god_coder` | `WithinUsAI/GOD_Coder_100k` | GOD coder | FULL โญ | | |
| | `python_god_coder` | `WithinUsAI/python_GOD_coder_100k` | Python GOD | FULL โญ | | |
| | `elite_god_coder` | `WithinUsAI/Elite_GOD_Coder_100k` | Elite GOD | FULL โญ | | |
| | `omega_genesis` | `WithinUsAI/Omega_Genesis_Coder_100k` | Omega Genesis | FULL โญ | | |
| | `open_tool_trace` | `WithinUsAI/OpenToolTrace-X` | Tool traces | 48 | | |
| | `genesis_v11` | `WithinUsAI/Genesis_v1_1_Update...` | Genesis v1.1 | partial | | |
| <br> | |
| ### ๐ง Distilled Category (9 sources) | |
| | Source | Upstream | Distilled From | | |
| |---|---|---| | |
| | `claude_mythos` | `WithinUsAI/claude_mythos_distilled_25k` | Claude | | |
| | `gemini35` | `WithinUsAI/gemini_3.5_flash_distilled_25k` | Gemini 3.5 Flash | | |
| | `fable5_cleaned` | `WithinUsAI/fable_5_distillation_merged_cleaned_25k` | Fable-5 | | |
| | `grok44` | `WithinUsAI/Grok4.4_heavy_max_distill_god_seed_25k` | Grok 4.4 | | |
| | `gemini_pro32` | `WithinUsAI/GeminiPro3.2_max_distill_god_seed_25k` | Gemini Pro 3.2 | | |
| | `gpt55_thinking` | `WithinUsAI/GPT5.5_thinking_max_distill_god_seed_25K` | GPT-5.5 | | |
| | `gpt55_distilled` | `WithinUsAI/GPT_5.5_Distilled` | GPT-5.5 | | |
| | `claude_opus_48_distill` | `11-47/claude_opus_4.8_distill_5k` | Claude Opus 4.8 โญ | | |
| | `claude_opus_48_max_thinking` | `11-47/claude_opus_4.8_max_thinking_5k_v2` | Opus 4.8 Max โญ | | |
| <br> | |
| ### ๐ฌ Science ยท โ๏ธ Applied ยท ๐ Humanities ยท ๐งฎ Math ยท ๐ Instruction ยท ๐ Cybersecurity ยท ๐ Index | |
| <details> | |
| <summary>๐ Click to expand all other categories</summary> | |
| **๐ฌ Science (7 sources):** `science_25k`, `physics_25k`, `chemistry_25k`, `biology_25k`, `medical_25k`, `cs_25k`, `biology_r2med` (R2MED/Biology) | |
| **โ๏ธ Applied (8 sources):** `robotics_25k`, `nano_25k`, `materials_25k`, `earth_climate_25k`, `renewable_energy_25k`, `evolution_25k`, `universe_25k`, `kardashev_25k` | |
| **๐ Humanities (8 sources):** `psychology_25k`, `economics_25k`, `law_25k`, `statistics_25k`, `sports_25k`, `human_25k`, `conscience_25k`, `supernatural_25k` | |
| **๐งฎ Math (2 sources):** `math_25k`, `deepseek_prover_v1` (27,503 Lean proofs) | |
| **๐ Instruction (3 sources):** `alpaca` (52K), `oasst` (32K), `dolly` (15K) | |
| **๐ Cybersecurity (6 sources):** `high_quality_cybersecurity`, `heimdall_v1_1`, `fenrir_v2_1`, `clydeiii_cybersecurity`, `precinct6_cybersecurity`, `savani_cyber_attack` | |
| **๐ Index (2 sources):** `species_25k`, `transport_25k` | |
| </details> | |
| --- | |
| ## ๐ก๏ธ Cybersecurity Deep Dive: Attack & Defense | |
| ### โ๏ธ Why This Matters | |
| Modern AI systems are increasingly deployed in security-critical environmentsโyet most open-source training data ignores real-world adversarial scenarios. **The Open Distillation Codex** includes a dedicated `cybersecurity` category designed to equip models with: | |
| - **Attack Awareness**: Recognize and generate realistic attack patterns, exploits, penetration testing commands, and social engineering dialogues. | |
| - **Defense Proficiency**: Learn to propose defensive measures, detect anomalies, and articulate incident response protocols. | |
| - **Exploit Understanding**: Analyze and explain software vulnerabilities, craft proof-of-concept code (for educational purposes), and understand exploit chains. | |
| - **Red/Blue Team Simulation**: Engage in multi-turn conversations mimicking red team attack planning and blue team defense coordination. | |
| - **Threat Intelligence**: Summarize, classify, and reason about cyber threat reports, CVEs, and IOCs (Indicators of Compromise). | |
| This makes the dataset a powerful foundation for building **cybersecurity-aware LLMs**, **security co-pilots**, and **automated vulnerability assessment tools**. | |
| ### ๐ Whatโs Inside the Cybersecurity Category? | |
| | Source | Description | Data Format | Key Themes | | |
| |:---|:---|:---|:---| | |
| | `high_quality_cybersecurity` | Manually curated high-quality instructionโresponse pairs covering attack techniques, defense, and policy | JSONL (shards) | MITRE ATT&CK, OWASP, incident response | | |
| | `heimdall_v1_1` | ~78 MB of security conversations, including red/blue team dialogues and threat analysis | JSONL | Multi-turn chat, tool usage | | |
| | `fenrir_v2_1` | 411 MB, 2.1M+ entries โ massive corpus of cybersecurity Q&A, exploit descriptions, and code snippets | JSONL | Exploit code, CVEs, vulnerability research | | |
| | `clydeiii_cybersecurity` | 20 MB yearly security corpus, aggregated from public reports and advisories | JSONL | Year-in-review, trends, threat landscape | | |
| | `precinct6_cybersecurity` | 2.1 GB graph-based dataset with network signals, attack graphs, and reference materials | JSONL (graph+signals+ref) | Network attacks, lateral movement, detection | | |
| | `savani_cyber_attack` | 17 MB CSV of labeled cyber attack incidents with detailed features | CSV | Attack classification, feature analysis | | |
| ### ๐งช Attack & Exploit Examples | |
| Here are a few representative samples (sanitized) from the dataset: | |
| **Example 1 โ SQL Injection Exploit** | |
| ```json | |
| { | |
| "source": "fenrir_v2_1", | |
| "instruction": "Explain how an attacker could exploit a vulnerable login form using SQL injection. Provide a sample malicious input.", | |
| "response": "An attacker could bypass authentication by injecting SQL logic. For a query like:\n`SELECT * FROM users WHERE username='$user' AND password='$pass'`,\nthe attacker might enter `admin' --` as the username and anything as password. The resulting query becomes:\n`SELECT * FROM users WHERE username='admin' --' AND password='...'`\nThe `--` comments out the password check, granting access as admin. More advanced attacks use UNION to extract data." | |
| } | |
| ``` | |
| **Example 2 โ Red Team Command Sequence** | |
| ```json | |
| { | |
| "source": "heimdall_v1_1", | |
| "instruction": "As a red team operator, list the steps you would take after gaining initial access to a Windows machine using a phishing payload.", | |
| "response": "1. Establish persistence: `schtasks /create /tn \"Updater\" /tr \"C:\\malware.exe\" /sc hourly`\n2. Enumerate local users: `net user /domain`\n3. Dump credentials: `mimikatz sekurlsa::logonpasswords`\n4. Lateral movement: use PsExec to move to file server.\n5. Exfiltrate data: compress sensitive files and upload via DNS tunneling.\nAlways maintain C2 communication over HTTPS to blend in." | |
| } | |
| ``` | |
| **Example 3 โ Defense Playbook (Blue Team)** | |
| ```json | |
| { | |
| "source": "high_quality_cybersecurity", | |
| "instruction": "An IDS alert flags unusual outbound DNS traffic from a workstation. What's the immediate response?", | |
| "response": "1. Isolate the workstation from the network (VLAN change or port shutdown).\n2. Capture volatile memory and network logs for forensics.\n3. Check DNS queries: if long, random-looking subdomains, suspect DNS tunneling.\n4. Scan for malware with updated signatures.\n5. Review firewall logs for similar traffic from other hosts.\n6. If confirmed, initiate incident response playbook for data exfiltration." | |
| } | |
| ``` | |
| ### ๐ How to Train a Cybersecurity-Focused LLM | |
| ```python | |
| from datasets import load_dataset | |
| REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset" | |
| # Load only cybersecurity category | |
| cyber_ds = load_dataset(REPO, split="train", | |
| data_files="data/cybersecurity/**/*.jsonl", | |
| streaming=True) | |
| # Or load specific sources | |
| fenrir = load_dataset(REPO, split="train", | |
| data_files="data/cybersecurity/fenrir_v2_1/*.jsonl") | |
| # Format for SFT | |
| def format_security_sample(example): | |
| return { | |
| "text": f"### Security Task:\n{example['instruction']}\n\n### Expert Response:\n{example['response']}" | |
| } | |
| cyber_ds = cyber_ds.map(format_security_sample) | |
| # Now train with your favourite framework (transformers, axolotl, etc.) | |
| ``` | |
| **Curriculum Idea**: | |
| 1. Start with `high_quality_cybersecurity` and `heimdall_v1_1` for foundational attack/defense conversations. | |
| 2. Introduce `fenrir_v2_1` for exploit code and vulnerability deep dives. | |
| 3. Use `precinct6_cybersecurity` for network-level attack graph understanding. | |
| ### ๐ก๏ธ Ethical & Responsible Use | |
| - **For Defensive Purposes Only**: This data is intended to strengthen AI for defense, threat detection, and security education. Do not use it to generate active attack code without proper authorization. | |
| - **No Zero-Day Exploits**: The dataset contains only already-public vulnerabilities and techniques. It does not include zero-day or weaponized exploits. | |
| - **Responsible Disclosure**: If you fine-tune a model with this data, we recommend adding a safety preamble warning that generated security content must be used legally and ethically. | |
| - **Dual-Use Awareness**: While we believe open access improves collective security, we acknowledge the dual-use nature. Users are expected to follow applicable laws and guidelines. | |
| > โ ๏ธ **Disclaimer**: This dataset includes descriptions of attack techniques for educational purposes. The maintainers are not responsible for misuse. | |
| ### ๐ Future Additions | |
| - Integration with CTF (Capture The Flag) challenge walkthroughs. | |
| - More blue team procedures and SOAR playbooks. | |
| - Anonymized real-world incident response logs (with permission). | |
| --- | |
| ## ๐ ๏ธ How to Use & Train | |
| ### 1๏ธโฃ Load Categorized JSONL Data | |
| ```python | |
| from datasets import load_dataset | |
| REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset" | |
| # โ Load a single category โ | |
| ds = load_dataset(REPO, split="train", data_files="data/coding/*/*.jsonl", streaming=True) | |
| # โ Load a specific source โ | |
| ds = load_dataset(REPO, split="train", data_files="data/coding/vibe_instruct_v2/*.jsonl", streaming=True) | |
| # โ Load everything (18M+ samples) โ | |
| ds = load_dataset(REPO, split="train", streaming=True) | |
| for sample in ds: | |
| print(sample["source"], sample["instruction"][:80]) | |
| ``` | |
| <br> | |
| ### 2๏ธโฃ Stream the 64 GB `archives/` GitHub Repositories | |
| ```python | |
| from huggingface_hub import hf_hub_download | |
| import tarfile | |
| REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset" | |
| # โ Option A: Download & extract ONE repository โ | |
| hf_hub_download( | |
| repo_id=REPO, | |
| repo_type="dataset", | |
| filename="archives/0x101__lakewatch.tar.gz", | |
| local_dir="./repos", | |
| ) | |
| with tarfile.open("./repos/archives/0x101__lakewatch.tar.gz", "r:gz") as tar: | |
| tar.extractall("./extracted/0x101__lakewatch") | |
| # โ Option B: Stream files WITHOUT full extraction โ | |
| def stream_repo_files(archive_name, max_files=100): | |
| """Stream file contents from tar.gz without extracting to disk.""" | |
| local_path = hf_hub_download(repo_id=REPO, repo_type="dataset", filename=archive_name) | |
| with tarfile.open(local_path, "r:gz") as tar: | |
| count = 0 | |
| for member in tar: | |
| if member.isfile() and count < max_files: | |
| f = tar.extractfile(member) | |
| if f: | |
| yield { | |
| "path": member.name, | |
| "content": f.read().decode("utf-8", errors="ignore")[:4000], | |
| } | |
| count += 1 | |
| import os | |
| os.remove(local_path) # Clean up | |
| # Stream files from a specific repo | |
| for file_data in stream_repo_files("archives/0x101__lakewatch.tar.gz"): | |
| print(f"๐ {file_data['path']}: {file_data['content'][:100]}...") | |
| # โ Option C: Use pre-extracted JSONL shards (475K samples) โ | |
| code_ds = load_dataset( | |
| REPO, split="train", | |
| data_files="data/coding/fable5_repos_full/*.jsonl", | |
| streaming=True | |
| ) | |
| # Each sample: instruction = "<repo>/<file>", response = "<content>" | |
| ``` | |
| <br> | |
| ### 3๏ธโฃ SFT Training Script (Hugging Face Trainer) | |
| ```python | |
| import torch | |
| from datasets import load_dataset | |
| from transformers import ( | |
| AutoTokenizer, | |
| AutoModelForCausalLM, | |
| TrainingArguments, | |
| Trainer, | |
| DataCollatorForLanguageModeling, | |
| ) | |
| # โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ | |
| # CONFIGURATION | |
| # โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ | |
| MODEL_NAME = "meta-llama/Llama-3.1-8B" | |
| DATASET_REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset" | |
| OUTPUT_DIR = "./sft-output" | |
| MAX_SEQ_LEN = 2048 | |
| # โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ | |
| # LOAD MODEL & TOKENIZER | |
| # โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ | |
| tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME) | |
| tokenizer.pad_token = tokenizer.eos_token | |
| model = AutoModelForCausalLM.from_pretrained( | |
| MODEL_NAME, | |
| torch_dtype=torch.bfloat16, | |
| device_map="auto", | |
| attn_implementation="flash_attention_2", | |
| ) | |
| # โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ | |
| # LOAD & FORMAT DATASET | |
| # โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ | |
| def format_instruction(sample): | |
| text = f"### Instruction:\n{sample['instruction']}\n\n### Response:\n{sample['response']}" | |
| return {"text": text} | |
| def tokenize(examples): | |
| return tokenizer( | |
| examples["text"], | |
| truncation=True, | |
| max_length=MAX_SEQ_LEN, | |
| padding="max_length", | |
| ) | |
| # Load coding category (use "data/**/*.jsonl" for full 18M+) | |
| train_ds = load_dataset( | |
| DATASET_REPO, | |
| split="train", | |
| data_files="data/coding/*/*.jsonl", | |
| streaming=True, | |
| ) | |
| train_ds = train_ds.map(format_instruction).filter(lambda x: len(x["text"]) > 0) | |
| train_ds = train_ds.map(tokenize, batched=True) | |
| # โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ | |
| # TRAIN | |
| # โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ | |
| training_args = TrainingArguments( | |
| output_dir=OUTPUT_DIR, | |
| num_train_epochs=3, | |
| per_device_train_batch_size=4, | |
| gradient_accumulation_steps=4, | |
| warmup_steps=500, | |
| logging_steps=100, | |
| save_steps=2000, | |
| learning_rate=2e-5, | |
| bf16=True, | |
| gradient_checkpointing=True, | |
| optim="adamw_torch", | |
| ) | |
| trainer = Trainer( | |
| model=model, | |
| args=training_args, | |
| train_dataset=train_ds, | |
| data_collator=DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False), | |
| ) | |
| trainer.train() | |
| trainer.save_model(OUTPUT_DIR) | |
| ``` | |
| <br> | |
| ### 4๏ธโฃ Curriculum Learning Across Categories | |
| ```python | |
| from datasets import load_dataset, interleave_datasets | |
| REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset" | |
| # โ Phase 1: Foundation (math + science) โ | |
| phase1_math = load_dataset(REPO, split="train", data_files="data/math/**/*.jsonl", streaming=True) | |
| phase1_sci = load_dataset(REPO, split="train", data_files="data/science/**/*.jsonl", streaming=True) | |
| phase1 = interleave_datasets([phase1_math, phase1_sci]) | |
| # โ Phase 2: Add coding traces โ | |
| phase2 = load_dataset(REPO, split="train", data_files="data/coding/**/*.jsonl", streaming=True) | |
| # โ Phase 3: Add distilled reasoning + cybersecurity โ | |
| phase3_distilled = load_dataset(REPO, split="train", data_files="data/distilled/**/*.jsonl", streaming=True) | |
| phase3_cyber = load_dataset(REPO, split="train", data_files="data/cybersecurity/**/*.jsonl", streaming=True) | |
| phase3 = interleave_datasets([phase3_distilled, phase3_cyber]) | |
| # Train sequentially | |
| # trainer.train(phase1) # epochs 0-1 | |
| # trainer.train(phase2) # epochs 1-2 | |
| # trainer.train(phase3) # epochs 2-3 | |
| ``` | |
| <br> | |
| ### ๐ Schema Reference | |
| ```json | |
| { | |
| "source": "fable5_2m", | |
| "source_dataset": "Crownelius/Complete-FABLE.5-traces-2M", | |
| "instruction": "<the prompt / question / file path>", | |
| "response": "<the completion / answer / file content>", | |
| "category": "coding" | |
| } | |
| ``` | |
| | Field | Type | Max Length | Description | | |
| |---|---|---|---| | |
| | `source` | string | 200 | Short slug identifying upstream dataset | | |
| | `source_dataset` | string | 200 | Full HF repo id (`org/name`) | | |
| | `instruction` | string | 4,000 | User-side content (prompt/question/file path) | | |
| | `response` | string | 4,000 | Assistant-side content (completion/answer/file content) | | |
| | `category` | string | 50 | One of 8 categories | | |
| --- | |
| ## ๐ Licensing & Limitations | |
| ### ๐ License | |
| The **collection as a whole** is released under the **MIT License**. | |
| Each upstream dataset retains its **original license**. The `source_dataset` field on every row identifies the upstream โ look it up on Hugging Face to determine its specific license. | |
| | License | Applies To | | |
| |---|---| | |
| | `MIT` | Most WithinUsAI datasets, OpenAssistant | | |
| | `Apache-2.0` | DeepSeek, OpenThoughts | | |
| | `CC-BY-4.0` | Dolly, various | | |
| | `CC-BY-SA-3.0` | Databricks Dolly | | |
| | `AGPL-3.0` | Some Fable-5 traces | | |
| ### โ Intended Use Cases (Our Vision) | |
| - Fine-tuning open-source LLMs for instruction following | |
| - Training coding agents and code-completion models | |
| - Reasoning chain distillation research | |
| - Domain-specific adaptation (math, science, cybersecurity) | |
| - Repository-scale context training (using `archives/`) | |
| ### โ Not Recommended For | |
| - Deploying models without safety evaluation | |
| - Generating harmful, biased, or deceptive content | |
| - High-stakes domains (medical, legal, financial) without expert review | |
| - Claiming models "know" facts โ this is distilled output, not ground truth | |
| ### โ ๏ธ Limitations | |
| 1. **Field length cap**: `instruction` and `response` capped at 4,000 characters. For full content, use `archives/`. | |
| 2. **Distillation artifacts**: Samples are model-generated โ may contain hallucinations or biases. | |
| 3. **Partial recovery**: A few upstream datasets (GOD_Coder variants, Genesis_v1.1) had format errors and were partially recovered via raw JSONL parsing. | |
| ### ๐ Citation | |
| ```bibtex | |
| @misc{open_distillation_codex_2026, | |
| title = {The Open Distillation Codex: 18M+ samples + 7090 code repositories from 73 sources with Cybersecurity Attack & Defense}, | |
| author = {Manusagents}, | |
| year = {2026}, | |
| url = {https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset}, | |
| note = {v8.2 - No skip, full. 516 shards + 7090 archives, 73 sources, 8 categories, 76 GB+} | |
| } | |
| ``` | |
| --- | |
| ## ๐ Changelog | |
| | Version | Date | Key Changes | | |
| |---|---|---| | |
| | `v1.0`โ`v5.0` | 2026-07-01 to 05 | Progressive builds: 117K โ 20.7M samples | | |
| | `v6.0` | 2026-07-06 | Category restructuring: `data/<category>/<source>/shard-*.jsonl` | | |
| | `v7.0` | 2026-07-06 | Training scripts + full processing started | | |
| | `v8.0 FINAL` | 2026-07-06 | **ALL sources FULLY processed โ no skipping. Verified 79.13 GB.** | | |
| | `v8.1` | 2026-07-08 | Added 5 external cybersecurity datasets. Total 81.2 GB, 73 sources. | | |
| | `v8.2` | 2026-07-18 | **Final numbers rectified: 18M+ samples, 76 GB+ total. All sources no skip, fully verified. Enhanced cybersecurity deep-dive with attack/defense examples, training scripts, ethical guidelines.** | | |
| --- | |
| <div align="center"> | |
| <br> | |
| ### ๐ The Open Distillation Codex ๐ | |
| **73 sources** ยท **8 categories** ยท **7,090 repositories** ยท **516 shards** ยท **76 GB+** | |
| <br> | |
| *No skip. Full. 18M+ samples. Built one archive at a time. Released under MIT.* | |
| <br> | |
| --- | |
| > *"Two layers. Eight categories. Seventy-three sources. One codex. No skip. Full. Armed with cybersecurity attack and defense."* | |
| <br> | |
| <img src="https://img.shields.io/badge/Built%20with-Streaming%20Pipeline-blue?style=flat-square" alt="Streaming"> | |
| <img src="https://img.shields.io/badge/No-Skipping-brightgreen?style=flat-square" alt="No Skip"> | |
| <img src="https://img.shields.io/badge/Full%20Processing-success?style=flat-square" alt="Full"> | |
| <img src="https://img.shields.io/badge/Format-JSONL-orange?style=flat-square" alt="JSONL"> | |
| <img src="https://img.shields.io/badge/HuggingFace-Dataset-yellow?style=flat-square" alt="HF"> | |
| <img src="https://img.shields.io/badge/Cybersecurity-Deep%20Dive-purple?style=flat-square" alt="Cyber"> | |
| <br><br> | |
| **โ The Open Distillation Codex โ** | |
| </div> |
Xet Storage Details
- Size:
- 36.9 kB
- Xet hash:
- 1f1f591a1199da185c990be099214e1d869f6fbf24eeda931baa906a5c617429
ยท
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.