# Data Guide: Datasets, Mixing, and Internal Data Curation ## Pre-Training Data Sources ### Tier 1: Core Datasets (must-have) | Dataset | What | Size | HF Link | License | |---------|------|------|---------|---------| | **FineWeb-Edu** | High-quality educational web text | ~4.3TB | [HuggingFaceFW/fineweb-edu](https://hf.co/datasets/HuggingFaceFW/fineweb-edu) | ODC-By | | **The Stack v2 (dedup)** | Deduplicated source code (600+ languages) | ~67TB raw | [bigcode/the-stack-v2-dedup](https://hf.co/datasets/bigcode/the-stack-v2-dedup) | Various (per-file) | | **Your Internal Code** | Internal Python libraries + apps | varies | Local | Internal | ### Tier 2: Domain-Specific (recommended) | Dataset | What | Size | HF Link | |---------|------|------|---------| | **OpenR1-Math-220k** | Math problems with CoT reasoning | ~2GB | [open-r1/OpenR1-Math-220k](https://hf.co/datasets/open-r1/OpenR1-Math-220k) | | **Proof-Pile-2** | Math/science papers + textbooks | ~53GB | [EleutherAI/proof-pile-2](https://hf.co/datasets/EleutherAI/proof-pile-2) | | **Wikipedia** | Structured knowledge | ~20GB | [wikimedia/wikipedia](https://hf.co/datasets/wikimedia/wikipedia) | | **Slurm/Bash data** (curated) | HPC scripts + documentation | 5-20GB | Self-curated | ### Tier 3: SFT-Quality (for annealing phase) | Dataset | What | Size | HF Link | |---------|------|------|---------| | **Magicoder-Evol-Instruct-110K** | Code instruction pairs | 110K examples | [ise-uiuc/Magicoder-Evol-Instruct-110K](https://hf.co/datasets/ise-uiuc/Magicoder-Evol-Instruct-110K) | | **OpenR1-Math-220k (messages)** | Math instruction w/ reasoning | 220K examples | Same as above | | **Synthetic internal data** | Your library Q&A, docs, examples | 50-200K examples | Self-generated | --- ## Data Mixing Strategy ### Phase 1a: General Pre-Training (70% of compute) ```python data_mix = { "fineweb-edu": 0.30, # General knowledge "the-stack-v2-python": 0.20, # Python code (primary language) "the-stack-v2-other": 0.15, # Other code (bash, JS, C, etc.) "openr1-math": 0.05, # Math reasoning "proof-pile": 0.05, # Math/science "wikipedia": 0.05, # Structured knowledge "github-docs": 0.05, # GitHub issues, READMEs, docs "slurm-bash": 0.05, # HPC/shell data "internal-libraries": 0.10, # YOUR internal code + docs } ``` ### Phase 1b: Code-Heavy Annealing (20% of compute) ```python data_mix_annealing = { "the-stack-v2-python": 0.35, # Increase Python "internal-libraries": 0.20, # Increase internal code "magicoder-evol": 0.10, # High-quality code instructions "openr1-math": 0.10, # Math reasoning "slurm-bash": 0.10, # Increase HPC data "fineweb-edu": 0.15, # Reduced general text } ``` ### Phase 1c: Long-Context Extension (10% of compute) ```python # Same mix as annealing but with long-context examples # Filter for documents > 8K tokens # Progressively increase from 4K → 32K → 131K context ``` ### Why These Ratios? Based on published training recipes: | Model | Code % | Math % | Web % | Other % | |-------|--------|--------|-------|---------| | DeepSeek-V3 | Not published but heavy code emphasis | | | | | StarCoder2 | 85% code + 10% math + 5% web | | | | | Llama 3 | 50% general + 25% code + 25% other | | | | | OLMoE | ~40% web + 30% code + 20% math/sci + 10% other | | | | | **Ours (recommended)** | **~45% code** | **~10% math** | **~30% web** | **~15% other** | The high code percentage (45%) reflects your primary use case. The 10% internal data ensures the model deeply understands your libraries. --- ## Internal Data Curation Pipeline ### Step 1: Crawl Your Repositories ```bash python data_curation/crawl_internal_repos.py \ --repos \ /path/to/your/library-core \ /path/to/your/library-utils \ /path/to/your/app-1 \ /path/to/your/app-2 \ /path/to/your/slurm-scripts \ --output data/raw/internal_raw.jsonl \ --include-tests \ --include-docs \ --repo-level ``` This produces: - **File-level examples:** Individual Python files with metadata headers - **Repo-level examples:** Combined package views showing cross-file dependencies - **Test files:** Unit tests that teach the model expected behavior - **Documentation:** READMEs, docstrings, markdown docs ### Step 2: Generate Synthetic Data with an API Model ```bash export OPENAI_API_KEY="sk-..." python data_curation/generate_synthetic_data.py \ --input data/raw/internal_raw.jsonl \ --output data/raw/internal_synthetic.jsonl \ --api openai \ --model gpt-4o \ --tasks docs qa completion bugfix slurm reasoning \ --num-per-file 5 \ --max-parallel 10 ``` **What this generates for each source file:** 1. **Documentation** — Full API docs, usage guides, tutorials 2. **Q&A Pairs** — "How do I use X?" → answer with code examples 3. **Code Completion** — Function stubs → full implementations 4. **Bug-Fix Pairs** — Buggy code → explanation → fix 5. **Slurm Scripts** — Job scripts using your library 6. **Reasoning** — Step-by-step problem solving with `` tags **Cost estimate:** At ~$0.03/generation (GPT-4o), 500 files × 6 tasks × 5 examples = 15K generations = ~$450 ### Step 3: Quality Filter and Deduplicate ```bash python data_curation/filter_and_dedup.py \ --input data/raw/internal_raw.jsonl data/raw/internal_synthetic.jsonl \ --output data/filtered/internal_filtered.jsonl \ --min-length 100 \ --max-length 100000 \ --min-quality 0.3 \ --dedup-threshold 0.7 ``` ### Step 4: Tokenize for Megatron-LM ```bash python data_curation/tokenize_for_megatron.py \ --input data/filtered/internal_filtered.jsonl \ --output-prefix data/tokenized/internal-libraries \ --tokenizer openai/gpt-oss-120b \ --workers 32 ``` This produces `internal-libraries_text_document.bin` and `.idx` files that Megatron-LM reads directly during training. ### Step 5: Prepare SFT Data For SFT (Phase 2), you need data in ChatML format: ```bash python data_curation/generate_synthetic_data.py \ --input data/raw/internal_raw.jsonl \ --output data/sft/internal_library_qa.jsonl \ --tasks qa reasoning \ --num-per-file 10 # Convert to ChatML messages format: python -c " import json with open('data/sft/internal_library_qa.jsonl') as f, \ open('data/sft/internal_sft_messages.jsonl', 'w') as out: for line in f: item = json.loads(line) text = item['text'] # Parse Q&A into messages if 'Question:' in text and 'Answer:' in text: parts = text.split('Answer:', 1) question = parts[0].replace('Question:', '').strip() answer = parts[1].strip() messages = [ {'role': 'system', 'content': 'You are a helpful coding assistant with deep knowledge of our internal libraries and HPC systems.'}, {'role': 'user', 'content': question}, {'role': 'assistant', 'content': answer} ] out.write(json.dumps({'messages': messages}) + '\n') " ``` --- ## Slurm/Bash Data Curation Since there's no standard Slurm dataset: ### 1. Your own cluster scripts ```bash python data_curation/prepare_slurm_data.py \ --output data/filtered/slurm_bash.jsonl \ --search-dirs ~/slurm-scripts /shared/job-scripts \ --include-manpages \ --include-synthetic ``` ### 2. GitHub mining (optional but valuable) Search GitHub for Slurm scripts: ```bash # Use GitHub API to find .sbatch files # Example search: "SBATCH language:shell filename:*.sbatch" # Collect, filter, and add to training data ``` ### 3. Documentation scraping - Slurm docs: https://slurm.schedmd.com/documentation.html - Convert HTML → Markdown for training ### 4. Synthetic generation (most effective) Use the API model to generate diverse Slurm scripts: ```bash python data_curation/generate_synthetic_data.py \ --input data/raw/internal_raw.jsonl \ --output data/raw/slurm_synthetic.jsonl \ --tasks slurm \ --num-per-file 10 ``` --- ## Data Quality Tips 1. **Upweight high-quality data in later epochs.** During annealing, increase the mix of curated, instruction-quality data and decrease noisy web data. 2. **Repeat internal data.** Your internal code is ~10% of the mix but might be <1% of total tokens. The model sees it many more times than public data. This is intentional and beneficial — it's similar to "curriculum learning" where important data is repeated. 3. **Include test files.** Unit tests are incredibly valuable training data — they show the model the expected input/output behavior of your libraries. The model learns "when I call X with Y, it should return Z." 4. **Include git history selectively.** Commit messages and diffs teach the model about code evolution and debugging patterns. But raw git logs are noisy — filter for meaningful commits. 5. **Decontaminate evaluation data.** Remove any overlap between your training data and evaluation benchmarks (HumanEval, MBPP, etc.) to get honest benchmark scores.