1.23 GB
800 files
Updated 10 days ago
Name
Size
data
.gitattributes1.73 kB
xet
README.md36.9 kB
xet
README.md
Version Storage Sources License Samples Cybersecurity



๐Ÿ“– The Open Distillation Codex

๐ŸŒŒ The Ultimate Open-Source Distillation Dataset โ€” No Skip, Full, with Attack & Defense ๐ŸŒŒ

Where 73 open-source minds converge into one unified stream of intelligence

18M+ Distilled Signals ยท 7,090 Raw GitHub Repositories ยท 8 Curated Categories ยท ~76 GB+


"We did not write this dataset. We assembled it. Every line is an echo โ€” of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three sources. Eight categories. Zero gatekeeping. No skipping. Fully processed. Now fortified with real-world cybersecurity confrontations."



๐Ÿ“Œ Table of Contents

# Section Description
1 ๐Ÿ“Š Dataset Summary High-level overview & value proposition
2 ๐Ÿ—‚๏ธ Directory Structure ASCII tree + folder explanation
3 ๐ŸŒ Data Sources All 73 sources with full attribution
4 ๐Ÿ›ก๏ธ Cybersecurity Deep Dive: Attack & Defense Importance, attack traces, defense, exploit analysis
5 ๐Ÿ› ๏ธ How to Use & Train Loading, streaming, training scripts
6 ๐Ÿ” Licensing & Limitations License, intended use, limitations
7 ๐Ÿ“œ Changelog Version history

๐Ÿ“Š Dataset Summary

๐ŸŽฏ The Numbers That Matter

Metric Value Status
Total Storage 76 GB+ โœ… Verified
JSONL Data Shards 516 โœ… Verified
Archive Files (tar.gz) 7,090 โœ… Verified
Source Datasets 73 โœ… Verified
Categories 8 โœ… Verified
Total Samples 18M+ โœ… Verified
Largest Source 8.15M (Vibe-Coding-Instruct-V2) โœ…
Archive Size ~64 GB (compressed GitHub repos) โœ…
Cybersecurity Sources 6 โœ…
Cybersecurity Data Size ~2.6 GB โœ…

๐ŸŒŸ Why "Ultimate Distilled"?

This dataset is not a raw scrape. Every sample has been distilled through a unified extraction pipeline:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    UNIFIED EXTRACTION PIPELINE              โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                                             โ”‚
โ”‚  73 Upstream Sources (ALL FULLY PROCESSED, NO SKIP)         โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ”          โ”‚
โ”‚  โ”‚ HF  โ”‚ โ”‚ HF  โ”‚ โ”‚ HF  โ”‚ โ”‚ GH  โ”‚ โ”‚ HF  โ”‚ โ”‚ ... โ”‚          โ”‚
โ”‚  โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜          โ”‚
โ”‚     โ”‚       โ”‚       โ”‚       โ”‚       โ”‚       โ”‚               โ”‚
โ”‚     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜               โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚ EXTRACT โ”‚ โ† Field normalization        โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”˜   (instruction/response)      โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚CATEGORIZEโ”‚ โ† 8 semantic categories      โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚  SHARD  โ”‚ โ† 20K samples per shard       โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚ UPLOAD  โ”‚ โ† Batch commits to HF         โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                                                             โ”‚
โ”‚  STATUS: ALL 73 SOURCES COMPLETE. NO SKIPPING. 18M+ ROWS.   โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿ’Ž Value to the Open-Source AI Community

๐ŸŽฏ For... ๐Ÿ“ฆ This dataset provides...
Model Trainers Single load_dataset() call to stream 18M+ SFT-ready samples
Coding Agent Researchers 11M+ agentic coding traces from Fable-5, Vibe-Coding, Royal Ghost, Kimi, DeepSeek
Code Pretraining 7,090 full GitHub repository snapshots (64 GB compressed)
Reasoning Researchers 2.7M+ distilled reasoning traces from Claude, Gemini, Grok, GPT-5.5, Opus 4.8
Domain Specialists 25K-sample sweeps across 29 disciplines
Cybersecurity Researchers Dedicated cybersecurity category with attack/defense/exploit traces, red/blue team dialogues, and incident reports
Red Team / Blue Team Trainers Realistic attack scenarios, defense strategies, exploit code, and post-mortem analysis

๐Ÿ—‚๏ธ Directory Structure

๐Ÿ“‚ Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset/
โ”‚
โ”œโ”€โ”€ ๐Ÿ“ฆ archives/                          # ~64 GB โ€” 7,090 compressed GitHub repos
โ”‚   โ”œโ”€โ”€ 0-chi__sonaure-lp.tar.gz
โ”‚   โ”œโ”€โ”€ 00MB__bitcoin_trading_bot.tar.gz
โ”‚   โ”œโ”€โ”€ 0101-agents__plugins.tar.gz
โ”‚   โ”œโ”€โ”€ ... (7,090 files total)
โ”‚   โ””โ”€โ”€ zznmg1__playable-survivor-ad.tar.gz
โ”‚
โ”œโ”€โ”€ ๐Ÿ“ data/                              # ~12 GB โ€” 516 JSONL shards (18M+ samples)
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ’ป coding/                        # 28 sources ยท ~11M+ samples
โ”‚   โ”‚   โ”œโ”€โ”€ vibe_instruct_v2/             # 8,152,510 samples
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_2m/                    # 2,006,487 samples
โ”‚   โ”‚   โ”œโ”€โ”€ vibe_instruct_v1/             # 1,100,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ vibe_coding/                  # 1,100,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ royal_ghost_1m/               # 1,000,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ citation_ground/              # 980,064 samples
โ”‚   โ”‚   โ”œโ”€โ”€ royal_ghost_501k/             # 703,449 samples
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_repos_full/            # 7,090 archive pointers
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_agentic_sft/           # 159,972 samples
โ”‚   โ”‚   โ”œโ”€โ”€ gpt55_codex/                  # 119,436 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ alpca_gpt55/                  # 49,099 samples
โ”‚   โ”‚   โ”œโ”€โ”€ deepseek_v4_pro_agent/        # 96,597 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_traces/                # 49,544 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code_100k/            # 68,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code/                 # 49,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ kimi_coding/                  # 9,014 samples
โ”‚   โ”‚   โ”œโ”€โ”€ mimo_claude_code_traces/      # 15,046 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ kimi_k26_claude_code_traces/  # 7,438 samples
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code_10k/             # 9,800 samples
โ”‚   โ”‚   โ”œโ”€โ”€ legend_python/                # 5,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ autonomy/                     # 10,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code_demo/            # 1,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ god_coder/                    # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ python_god_coder/             # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ elite_god_coder/              # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ omega_genesis/                # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ open_tool_trace/              # 48 samples
โ”‚   โ”‚   โ””โ”€โ”€ genesis_v11/                  # partial recovery
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿงฎ math/                          # 2 sources
โ”‚   โ”‚   โ”œโ”€โ”€ math_25k/
โ”‚   โ”‚   โ””โ”€โ”€ deepseek_prover_v1/           # 27,503 Lean theorem proofs
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ”ฌ science/                       # 7 sources
โ”‚   โ”‚   โ”œโ”€โ”€ science_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ physics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ chemistry_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ biology_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ medical_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ cs_25k/
โ”‚   โ”‚   โ””โ”€โ”€ biology_r2med/                # โญ NEW
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ โš™๏ธ applied/                       # 8 sources
โ”‚   โ”‚   โ”œโ”€โ”€ robotics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ nano_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ materials_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ earth_climate_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ renewable_energy_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ evolution_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ universe_25k/
โ”‚   โ”‚   โ””โ”€โ”€ kardashev_25k/
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ“š humanities/                    # 8 sources
โ”‚   โ”‚   โ”œโ”€โ”€ psychology_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ economics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ law_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ statistics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ sports_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ human_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ conscience_25k/
โ”‚   โ”‚   โ””โ”€โ”€ supernatural_25k/
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿง  distilled/                     # 9 sources ยท frontier distillations
โ”‚   โ”‚   โ”œโ”€โ”€ claude_mythos/
โ”‚   โ”‚   โ”œโ”€โ”€ gemini35/
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_cleaned/
โ”‚   โ”‚   โ”œโ”€โ”€ grok44/
โ”‚   โ”‚   โ”œโ”€โ”€ gemini_pro32/
โ”‚   โ”‚   โ”œโ”€โ”€ gpt55_thinking/
โ”‚   โ”‚   โ”œโ”€โ”€ gpt55_distilled/
โ”‚   โ”‚   โ”œโ”€โ”€ claude_opus_48_distill/       # โญ NEW
โ”‚   โ”‚   โ””โ”€โ”€ claude_opus_48_max_thinking/  # โญ NEW
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ“ instruction/                   # 3 sources
โ”‚   โ”‚   โ”œโ”€โ”€ alpaca/                       # 52,002 samples
โ”‚   โ”‚   โ”œโ”€โ”€ oasst/                        # 32,141 samples
โ”‚   โ”‚   โ””โ”€โ”€ dolly/                        # 15,011 samples
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ”’ cybersecurity/                 # 6 sources
โ”‚   โ”‚   โ”œโ”€โ”€ high_quality_cybersecurity/
โ”‚   โ”‚   โ”œโ”€โ”€ heimdall_v1_1/                # โญ NEW โ€” 78 MB conversations
โ”‚   โ”‚   โ”œโ”€โ”€ fenrir_v2_1/                  # โญ NEW โ€” 411 MB (2.1M+ entries)
โ”‚   โ”‚   โ”œโ”€โ”€ clydeiii_cybersecurity/       # โญ NEW โ€” 20 MB yearly corpus
โ”‚   โ”‚   โ”œโ”€โ”€ precinct6_cybersecurity/      # โญ NEW โ€” 2.1 GB (graph+signals+ref)
โ”‚   โ”‚   โ””โ”€โ”€ savani_cyber_attack/          # โญ NEW โ€” 17 MB attack CSV
โ”‚   โ”‚
โ”‚   โ””โ”€โ”€ ๐Ÿ“‡ index/                         # 2 sources
โ”‚       โ”œโ”€โ”€ species_25k/
โ”‚       โ””โ”€โ”€ transport_25k/
โ”‚
โ”œโ”€โ”€ ๐Ÿ“„ README.md
โ””โ”€โ”€ ๐Ÿ“„ dataset_info.json

๐Ÿค” Why is archives/ kept compressed?

Reason Explanation
๐Ÿ’พ Space Efficiency Uncompressed would exceed 200+ GB. Compressed = 64 GB (3ร— saving)
๐ŸŽฏ On-Demand Access Download only specific repositories you need
๐Ÿ” Preservation Fidelity tar.gz preserves exact file permissions, directory structure, binaries

๐Ÿ’ก Tip: For training on code content, use data/coding/fable5_repos_full/ (475K samples, each a file extracted from archives, capped at 4KB). For full untruncated file access, stream directly from archives/.


๐ŸŒ Data Sources & Provenance

๐Ÿ—บ๏ธ 73 Sources Across 8 Categories

Category Sources Samples Description
๐Ÿ’ป coding 28 ~11M+ Agentic traces, code repos, coder distillations
๐Ÿง  distilled 9 ~200K Frontier model distillations
โš™๏ธ applied 8 ~200K Robotics, nano, materials, climate, energy
๐Ÿ“š humanities 8 ~200K Psychology, economics, law, statistics
๐Ÿ”ฌ science 7 ~175K Physics, chemistry, biology, medical, CS
๐Ÿ“ instruction 3 ~99K Classic instruction (alpaca, oasst, dolly)
๐Ÿ“‡ index 2 ~50K Species index, transport
๐Ÿ”’ cybersecurity 6 ~2.6 GB High-quality attack, defense, exploit traces
๐Ÿงฎ math 2 ~52K Math + Lean theorem proofs

๐Ÿ’ป Coding Category (28 sources โ€” ALL FULLY PROCESSED โญ)

Source Slug Upstream Dataset Type Samples
vibe_instruct_v2 CodeDevX/Vibe-Coding-Instruct-V2 Agentic coding 8,152,510
fable5_2m Crownelius/Complete-FABLE.5-traces-2M Fable-5 traces 2,006,487
vibe_instruct_v1 CodeDevX/Vibe-Coding-Instruct Agentic coding 1,100,000
vibe_coding attentionAllYouNeed/Vibe-Coding-Claude-Fable-5 Claude coding 1,100,000
royal_ghost_1m WithinUsAI/Royal_Ghost_Coder_1M Ghost coder 1,000,000
citation_ground WithinUsAI/CitationGround-1M Citation-grounded 980,064
royal_ghost_501k WithinUsAI/Royal_Ghost_Coder_501k Ghost coder 703,449
fable5_repos_full notune/fable5-repos 7,090 repo pointers 7,090
fable5_agentic_sft Nexlab/fable5-agentic-coding-sft Agentic SFT 159,972
gpt55_codex AletheiaResearch/GPT-5.5-Codex GPT-5.5 Codex 119,436
alpca_gpt55 GabrielFreeze-2/alpca-mlt-gpt-5.5_chatml GPT-5.5 chatml 49,099
deepseek_v4_pro_agent TeichAI/DeepSeek-v4-Pro-Agent DeepSeek v4 96,597
fable5_traces Glint-Research/Fable-5-traces Fable-5 traces 49,544
genesis_code_100k WithinUsAI/Genesis_AI_Code_100k Genesis code 68,000
genesis_code WithinUsAI/Genesis_AI_Code_50k Genesis code 49,000
kimi_coding trjxter/Kimi-K2.7-CodingTraces-9000x Kimi K2.7 9,014
mimo_claude_code_traces choucsan/mimo-claude-code-traces-1k Mimo Claude 15,046
kimi_k26_claude_code_traces armand0e/kimi-k2.6-claude-code-traces Kimi K2.6 7,438
genesis_code_10k WithinUsAI/Genesis_AI_Code_10k Genesis code 9,800
legend_python WithinUsAI/Legend_Python_CoderV.1 Python coder 5,000
autonomy WithinUsAI/The_Autonomy_From_WithIn_10k Autonomy 10,000
genesis_code_demo WithinUsAI/Genesis_AI_Code_1k_Demo Genesis demo 1,000
god_coder WithinUsAI/GOD_Coder_100k GOD coder FULL โญ
python_god_coder WithinUsAI/python_GOD_coder_100k Python GOD FULL โญ
elite_god_coder WithinUsAI/Elite_GOD_Coder_100k Elite GOD FULL โญ
omega_genesis WithinUsAI/Omega_Genesis_Coder_100k Omega Genesis FULL โญ
open_tool_trace WithinUsAI/OpenToolTrace-X Tool traces 48
genesis_v11 WithinUsAI/Genesis_v1_1_Update... Genesis v1.1 partial

๐Ÿง  Distilled Category (9 sources)

Source Upstream Distilled From
claude_mythos WithinUsAI/claude_mythos_distilled_25k Claude
gemini35 WithinUsAI/gemini_3.5_flash_distilled_25k Gemini 3.5 Flash
fable5_cleaned WithinUsAI/fable_5_distillation_merged_cleaned_25k Fable-5
grok44 WithinUsAI/Grok4.4_heavy_max_distill_god_seed_25k Grok 4.4
gemini_pro32 WithinUsAI/GeminiPro3.2_max_distill_god_seed_25k Gemini Pro 3.2
gpt55_thinking WithinUsAI/GPT5.5_thinking_max_distill_god_seed_25K GPT-5.5
gpt55_distilled WithinUsAI/GPT_5.5_Distilled GPT-5.5
claude_opus_48_distill 11-47/claude_opus_4.8_distill_5k Claude Opus 4.8 โญ
claude_opus_48_max_thinking 11-47/claude_opus_4.8_max_thinking_5k_v2 Opus 4.8 Max โญ

๐Ÿ”ฌ Science ยท โš™๏ธ Applied ยท ๐Ÿ“š Humanities ยท ๐Ÿงฎ Math ยท ๐Ÿ“ Instruction ยท ๐Ÿ”’ Cybersecurity ยท ๐Ÿ“‡ Index

๐Ÿ“– Click to expand all other categories

๐Ÿ”ฌ Science (7 sources): science_25k, physics_25k, chemistry_25k, biology_25k, medical_25k, cs_25k, biology_r2med (R2MED/Biology)

โš™๏ธ Applied (8 sources): robotics_25k, nano_25k, materials_25k, earth_climate_25k, renewable_energy_25k, evolution_25k, universe_25k, kardashev_25k

๐Ÿ“š Humanities (8 sources): psychology_25k, economics_25k, law_25k, statistics_25k, sports_25k, human_25k, conscience_25k, supernatural_25k

๐Ÿงฎ Math (2 sources): math_25k, deepseek_prover_v1 (27,503 Lean proofs)

๐Ÿ“ Instruction (3 sources): alpaca (52K), oasst (32K), dolly (15K)

๐Ÿ”’ Cybersecurity (6 sources): high_quality_cybersecurity, heimdall_v1_1, fenrir_v2_1, clydeiii_cybersecurity, precinct6_cybersecurity, savani_cyber_attack

๐Ÿ“‡ Index (2 sources): species_25k, transport_25k


๐Ÿ›ก๏ธ Cybersecurity Deep Dive: Attack & Defense

โš”๏ธ Why This Matters

Modern AI systems are increasingly deployed in security-critical environmentsโ€”yet most open-source training data ignores real-world adversarial scenarios. The Open Distillation Codex includes a dedicated cybersecurity category designed to equip models with:

  • Attack Awareness: Recognize and generate realistic attack patterns, exploits, penetration testing commands, and social engineering dialogues.
  • Defense Proficiency: Learn to propose defensive measures, detect anomalies, and articulate incident response protocols.
  • Exploit Understanding: Analyze and explain software vulnerabilities, craft proof-of-concept code (for educational purposes), and understand exploit chains.
  • Red/Blue Team Simulation: Engage in multi-turn conversations mimicking red team attack planning and blue team defense coordination.
  • Threat Intelligence: Summarize, classify, and reason about cyber threat reports, CVEs, and IOCs (Indicators of Compromise).

This makes the dataset a powerful foundation for building cybersecurity-aware LLMs, security co-pilots, and automated vulnerability assessment tools.

๐Ÿ“Š Whatโ€™s Inside the Cybersecurity Category?

Source Description Data Format Key Themes
high_quality_cybersecurity Manually curated high-quality instructionโ€“response pairs covering attack techniques, defense, and policy JSONL (shards) MITRE ATT&CK, OWASP, incident response
heimdall_v1_1 ~78 MB of security conversations, including red/blue team dialogues and threat analysis JSONL Multi-turn chat, tool usage
fenrir_v2_1 411 MB, 2.1M+ entries โ€” massive corpus of cybersecurity Q&A, exploit descriptions, and code snippets JSONL Exploit code, CVEs, vulnerability research
clydeiii_cybersecurity 20 MB yearly security corpus, aggregated from public reports and advisories JSONL Year-in-review, trends, threat landscape
precinct6_cybersecurity 2.1 GB graph-based dataset with network signals, attack graphs, and reference materials JSONL (graph+signals+ref) Network attacks, lateral movement, detection
savani_cyber_attack 17 MB CSV of labeled cyber attack incidents with detailed features CSV Attack classification, feature analysis

๐Ÿงช Attack & Exploit Examples

Here are a few representative samples (sanitized) from the dataset:

Example 1 โ€“ SQL Injection Exploit

{
  "source": "fenrir_v2_1",
  "instruction": "Explain how an attacker could exploit a vulnerable login form using SQL injection. Provide a sample malicious input.",
  "response": "An attacker could bypass authentication by injecting SQL logic. For a query like:\n`SELECT * FROM users WHERE username='$user' AND password='$pass'`,\nthe attacker might enter `admin' --` as the username and anything as password. The resulting query becomes:\n`SELECT * FROM users WHERE username='admin' --' AND password='...'`\nThe `--` comments out the password check, granting access as admin. More advanced attacks use UNION to extract data."
}

Example 2 โ€“ Red Team Command Sequence

{
  "source": "heimdall_v1_1",
  "instruction": "As a red team operator, list the steps you would take after gaining initial access to a Windows machine using a phishing payload.",
  "response": "1. Establish persistence: `schtasks /create /tn \"Updater\" /tr \"C:\\malware.exe\" /sc hourly`\n2. Enumerate local users: `net user /domain`\n3. Dump credentials: `mimikatz sekurlsa::logonpasswords`\n4. Lateral movement: use PsExec to move to file server.\n5. Exfiltrate data: compress sensitive files and upload via DNS tunneling.\nAlways maintain C2 communication over HTTPS to blend in."
}

Example 3 โ€“ Defense Playbook (Blue Team)

{
  "source": "high_quality_cybersecurity",
  "instruction": "An IDS alert flags unusual outbound DNS traffic from a workstation. What's the immediate response?",
  "response": "1. Isolate the workstation from the network (VLAN change or port shutdown).\n2. Capture volatile memory and network logs for forensics.\n3. Check DNS queries: if long, random-looking subdomains, suspect DNS tunneling.\n4. Scan for malware with updated signatures.\n5. Review firewall logs for similar traffic from other hosts.\n6. If confirmed, initiate incident response playbook for data exfiltration."
}

๐ŸŽ“ How to Train a Cybersecurity-Focused LLM

from datasets import load_dataset

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# Load only cybersecurity category
cyber_ds = load_dataset(REPO, split="train", 
                        data_files="data/cybersecurity/**/*.jsonl",
                        streaming=True)

# Or load specific sources
fenrir = load_dataset(REPO, split="train", 
                      data_files="data/cybersecurity/fenrir_v2_1/*.jsonl")

# Format for SFT
def format_security_sample(example):
    return {
        "text": f"### Security Task:\n{example['instruction']}\n\n### Expert Response:\n{example['response']}"
    }

cyber_ds = cyber_ds.map(format_security_sample)

# Now train with your favourite framework (transformers, axolotl, etc.)

Curriculum Idea:

  1. Start with high_quality_cybersecurity and heimdall_v1_1 for foundational attack/defense conversations.
  2. Introduce fenrir_v2_1 for exploit code and vulnerability deep dives.
  3. Use precinct6_cybersecurity for network-level attack graph understanding.

๐Ÿ›ก๏ธ Ethical & Responsible Use

  • For Defensive Purposes Only: This data is intended to strengthen AI for defense, threat detection, and security education. Do not use it to generate active attack code without proper authorization.
  • No Zero-Day Exploits: The dataset contains only already-public vulnerabilities and techniques. It does not include zero-day or weaponized exploits.
  • Responsible Disclosure: If you fine-tune a model with this data, we recommend adding a safety preamble warning that generated security content must be used legally and ethically.
  • Dual-Use Awareness: While we believe open access improves collective security, we acknowledge the dual-use nature. Users are expected to follow applicable laws and guidelines.

โš ๏ธ Disclaimer: This dataset includes descriptions of attack techniques for educational purposes. The maintainers are not responsible for misuse.

๐Ÿ“ˆ Future Additions

  • Integration with CTF (Capture The Flag) challenge walkthroughs.
  • More blue team procedures and SOAR playbooks.
  • Anonymized real-world incident response logs (with permission).

๐Ÿ› ๏ธ How to Use & Train

1๏ธโƒฃ Load Categorized JSONL Data

from datasets import load_dataset

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# โ”€ Load a single category โ”€
ds = load_dataset(REPO, split="train", data_files="data/coding/*/*.jsonl", streaming=True)

# โ”€ Load a specific source โ”€
ds = load_dataset(REPO, split="train", data_files="data/coding/vibe_instruct_v2/*.jsonl", streaming=True)

# โ”€ Load everything (18M+ samples) โ”€
ds = load_dataset(REPO, split="train", streaming=True)

for sample in ds:
    print(sample["source"], sample["instruction"][:80])

2๏ธโƒฃ Stream the 64 GB archives/ GitHub Repositories

from huggingface_hub import hf_hub_download
import tarfile

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# โ”€ Option A: Download & extract ONE repository โ”€
hf_hub_download(
    repo_id=REPO,
    repo_type="dataset",
    filename="archives/0x101__lakewatch.tar.gz",
    local_dir="./repos",
)
with tarfile.open("./repos/archives/0x101__lakewatch.tar.gz", "r:gz") as tar:
    tar.extractall("./extracted/0x101__lakewatch")


# โ”€ Option B: Stream files WITHOUT full extraction โ”€
def stream_repo_files(archive_name, max_files=100):
    """Stream file contents from tar.gz without extracting to disk."""
    local_path = hf_hub_download(repo_id=REPO, repo_type="dataset", filename=archive_name)
    
    with tarfile.open(local_path, "r:gz") as tar:
        count = 0
        for member in tar:
            if member.isfile() and count < max_files:
                f = tar.extractfile(member)
                if f:
                    yield {
                        "path": member.name,
                        "content": f.read().decode("utf-8", errors="ignore")[:4000],
                    }
                    count += 1
    
    import os
    os.remove(local_path)  # Clean up

# Stream files from a specific repo
for file_data in stream_repo_files("archives/0x101__lakewatch.tar.gz"):
    print(f"๐Ÿ“„ {file_data['path']}: {file_data['content'][:100]}...")


# โ”€ Option C: Use pre-extracted JSONL shards (475K samples) โ”€
code_ds = load_dataset(
    REPO, split="train",
    data_files="data/coding/fable5_repos_full/*.jsonl",
    streaming=True
)
# Each sample: instruction = "<repo>/<file>", response = "<content>"

3๏ธโƒฃ SFT Training Script (Hugging Face Trainer)

import torch
from datasets import load_dataset
from transformers import (
    AutoTokenizer,
    AutoModelForCausalLM,
    TrainingArguments,
    Trainer,
    DataCollatorForLanguageModeling,
)

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# CONFIGURATION
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
MODEL_NAME = "meta-llama/Llama-3.1-8B"
DATASET_REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"
OUTPUT_DIR = "./sft-output"
MAX_SEQ_LEN = 2048

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# LOAD MODEL & TOKENIZER
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
tokenizer.pad_token = tokenizer.eos_token

model = AutoModelForCausalLM.from_pretrained(
    MODEL_NAME,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="flash_attention_2",
)

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# LOAD & FORMAT DATASET
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
def format_instruction(sample):
    text = f"### Instruction:\n{sample['instruction']}\n\n### Response:\n{sample['response']}"
    return {"text": text}

def tokenize(examples):
    return tokenizer(
        examples["text"],
        truncation=True,
        max_length=MAX_SEQ_LEN,
        padding="max_length",
    )

# Load coding category (use "data/**/*.jsonl" for full 18M+)
train_ds = load_dataset(
    DATASET_REPO,
    split="train",
    data_files="data/coding/*/*.jsonl",
    streaming=True,
)
train_ds = train_ds.map(format_instruction).filter(lambda x: len(x["text"]) > 0)
train_ds = train_ds.map(tokenize, batched=True)

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# TRAIN
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
training_args = TrainingArguments(
    output_dir=OUTPUT_DIR,
    num_train_epochs=3,
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,
    warmup_steps=500,
    logging_steps=100,
    save_steps=2000,
    learning_rate=2e-5,
    bf16=True,
    gradient_checkpointing=True,
    optim="adamw_torch",
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=train_ds,
    data_collator=DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False),
)

trainer.train()
trainer.save_model(OUTPUT_DIR)

4๏ธโƒฃ Curriculum Learning Across Categories

from datasets import load_dataset, interleave_datasets

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# โ”€ Phase 1: Foundation (math + science) โ”€
phase1_math = load_dataset(REPO, split="train", data_files="data/math/**/*.jsonl", streaming=True)
phase1_sci = load_dataset(REPO, split="train", data_files="data/science/**/*.jsonl", streaming=True)
phase1 = interleave_datasets([phase1_math, phase1_sci])

# โ”€ Phase 2: Add coding traces โ”€
phase2 = load_dataset(REPO, split="train", data_files="data/coding/**/*.jsonl", streaming=True)

# โ”€ Phase 3: Add distilled reasoning + cybersecurity โ”€
phase3_distilled = load_dataset(REPO, split="train", data_files="data/distilled/**/*.jsonl", streaming=True)
phase3_cyber = load_dataset(REPO, split="train", data_files="data/cybersecurity/**/*.jsonl", streaming=True)
phase3 = interleave_datasets([phase3_distilled, phase3_cyber])

# Train sequentially
# trainer.train(phase1)  # epochs 0-1
# trainer.train(phase2)  # epochs 1-2
# trainer.train(phase3)  # epochs 2-3

๐Ÿ“‹ Schema Reference

{
    "source":          "fable5_2m",
    "source_dataset":  "Crownelius/Complete-FABLE.5-traces-2M",
    "instruction":     "<the prompt / question / file path>",
    "response":        "<the completion / answer / file content>",
    "category":        "coding"
}
Field Type Max Length Description
source string 200 Short slug identifying upstream dataset
source_dataset string 200 Full HF repo id (org/name)
instruction string 4,000 User-side content (prompt/question/file path)
response string 4,000 Assistant-side content (completion/answer/file content)
category string 50 One of 8 categories

๐Ÿ” Licensing & Limitations

๐Ÿ“œ License

The collection as a whole is released under the MIT License.

Each upstream dataset retains its original license. The source_dataset field on every row identifies the upstream โ€” look it up on Hugging Face to determine its specific license.

License Applies To
MIT Most WithinUsAI datasets, OpenAssistant
Apache-2.0 DeepSeek, OpenThoughts
CC-BY-4.0 Dolly, various
CC-BY-SA-3.0 Databricks Dolly
AGPL-3.0 Some Fable-5 traces

โœ… Intended Use Cases (Our Vision)

  • Fine-tuning open-source LLMs for instruction following
  • Training coding agents and code-completion models
  • Reasoning chain distillation research
  • Domain-specific adaptation (math, science, cybersecurity)
  • Repository-scale context training (using archives/)

โŒ Not Recommended For

  • Deploying models without safety evaluation
  • Generating harmful, biased, or deceptive content
  • High-stakes domains (medical, legal, financial) without expert review
  • Claiming models "know" facts โ€” this is distilled output, not ground truth

โš ๏ธ Limitations

  1. Field length cap: instruction and response capped at 4,000 characters. For full content, use archives/.
  2. Distillation artifacts: Samples are model-generated โ€” may contain hallucinations or biases.
  3. Partial recovery: A few upstream datasets (GOD_Coder variants, Genesis_v1.1) had format errors and were partially recovered via raw JSONL parsing.

๐Ÿ“ Citation

@misc{open_distillation_codex_2026,
  title  = {The Open Distillation Codex: 18M+ samples + 7090 code repositories from 73 sources with Cybersecurity Attack & Defense},
  author = {Manusagents},
  year   = {2026},
  url    = {https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset},
  note   = {v8.2 - No skip, full. 516 shards + 7090 archives, 73 sources, 8 categories, 76 GB+}
}

๐Ÿ“œ Changelog

Version Date Key Changes
v1.0โ€“v5.0 2026-07-01 to 05 Progressive builds: 117K โ†’ 20.7M samples
v6.0 2026-07-06 Category restructuring: data/<category>/<source>/shard-*.jsonl
v7.0 2026-07-06 Training scripts + full processing started
v8.0 FINAL 2026-07-06 ALL sources FULLY processed โ€” no skipping. Verified 79.13 GB.
v8.1 2026-07-08 Added 5 external cybersecurity datasets. Total 81.2 GB, 73 sources.
v8.2 2026-07-18 Final numbers rectified: 18M+ samples, 76 GB+ total. All sources no skip, fully verified. Enhanced cybersecurity deep-dive with attack/defense examples, training scripts, ethical guidelines.


๐ŸŒŸ The Open Distillation Codex ๐ŸŒŸ

73 sources ยท 8 categories ยท 7,090 repositories ยท 516 shards ยท 76 GB+


No skip. Full. 18M+ samples. Built one archive at a time. Released under MIT.



"Two layers. Eight categories. Seventy-three sources. One codex. No skip. Full. Armed with cybersecurity attack and defense."


Streaming No Skip Full JSONL HF Cyber



โ€” The Open Distillation Codex โ€”

Total size
1.23 GB
Files
800
Last updated
Jul 24
Pre-warmed CDN
US EU US EU

Contributors