| # Data Guide: Datasets, Mixing, and Internal Data Curation |
|
|
| ## Pre-Training Data Sources |
|
|
| ### Tier 1: Core Datasets (must-have) |
|
|
| | Dataset | What | Size | HF Link | License | |
| |---------|------|------|---------|---------| |
| | **FineWeb-Edu** | High-quality educational web text | ~4.3TB | [HuggingFaceFW/fineweb-edu](https://hf.co/datasets/HuggingFaceFW/fineweb-edu) | ODC-By | |
| | **The Stack v2 (dedup)** | Deduplicated source code (600+ languages) | ~67TB raw | [bigcode/the-stack-v2-dedup](https://hf.co/datasets/bigcode/the-stack-v2-dedup) | Various (per-file) | |
| | **Your Internal Code** | Internal Python libraries + apps | varies | Local | Internal | |
|
|
| ### Tier 2: Domain-Specific (recommended) |
|
|
| | Dataset | What | Size | HF Link | |
| |---------|------|------|---------| |
| | **OpenR1-Math-220k** | Math problems with CoT reasoning | ~2GB | [open-r1/OpenR1-Math-220k](https://hf.co/datasets/open-r1/OpenR1-Math-220k) | |
| | **Proof-Pile-2** | Math/science papers + textbooks | ~53GB | [EleutherAI/proof-pile-2](https://hf.co/datasets/EleutherAI/proof-pile-2) | |
| | **Wikipedia** | Structured knowledge | ~20GB | [wikimedia/wikipedia](https://hf.co/datasets/wikimedia/wikipedia) | |
| | **Slurm/Bash data** (curated) | HPC scripts + documentation | 5-20GB | Self-curated | |
|
|
| ### Tier 3: SFT-Quality (for annealing phase) |
|
|
| | Dataset | What | Size | HF Link | |
| |---------|------|------|---------| |
| | **Magicoder-Evol-Instruct-110K** | Code instruction pairs | 110K examples | [ise-uiuc/Magicoder-Evol-Instruct-110K](https://hf.co/datasets/ise-uiuc/Magicoder-Evol-Instruct-110K) | |
| | **OpenR1-Math-220k (messages)** | Math instruction w/ reasoning | 220K examples | Same as above | |
| | **Synthetic internal data** | Your library Q&A, docs, examples | 50-200K examples | Self-generated | |
|
|
| --- |
|
|
| ## Data Mixing Strategy |
|
|
| ### Phase 1a: General Pre-Training (70% of compute) |
|
|
| ```python |
| data_mix = { |
| "fineweb-edu": 0.30, # General knowledge |
| "the-stack-v2-python": 0.20, # Python code (primary language) |
| "the-stack-v2-other": 0.15, # Other code (bash, JS, C, etc.) |
| "openr1-math": 0.05, # Math reasoning |
| "proof-pile": 0.05, # Math/science |
| "wikipedia": 0.05, # Structured knowledge |
| "github-docs": 0.05, # GitHub issues, READMEs, docs |
| "slurm-bash": 0.05, # HPC/shell data |
| "internal-libraries": 0.10, # YOUR internal code + docs |
| } |
| ``` |
|
|
| ### Phase 1b: Code-Heavy Annealing (20% of compute) |
|
|
| ```python |
| data_mix_annealing = { |
| "the-stack-v2-python": 0.35, # Increase Python |
| "internal-libraries": 0.20, # Increase internal code |
| "magicoder-evol": 0.10, # High-quality code instructions |
| "openr1-math": 0.10, # Math reasoning |
| "slurm-bash": 0.10, # Increase HPC data |
| "fineweb-edu": 0.15, # Reduced general text |
| } |
| ``` |
|
|
| ### Phase 1c: Long-Context Extension (10% of compute) |
|
|
| ```python |
| # Same mix as annealing but with long-context examples |
| # Filter for documents > 8K tokens |
| # Progressively increase from 4K β 32K β 131K context |
| ``` |
|
|
| ### Why These Ratios? |
|
|
| Based on published training recipes: |
|
|
| | Model | Code % | Math % | Web % | Other % | |
| |-------|--------|--------|-------|---------| |
| | DeepSeek-V3 | Not published but heavy code emphasis | | | | |
| | StarCoder2 | 85% code + 10% math + 5% web | | | | |
| | Llama 3 | 50% general + 25% code + 25% other | | | | |
| | OLMoE | ~40% web + 30% code + 20% math/sci + 10% other | | | | |
| | **Ours (recommended)** | **~45% code** | **~10% math** | **~30% web** | **~15% other** | |
|
|
| The high code percentage (45%) reflects your primary use case. The 10% internal data |
| ensures the model deeply understands your libraries. |
|
|
| --- |
|
|
| ## Internal Data Curation Pipeline |
|
|
| ### Step 1: Crawl Your Repositories |
|
|
| ```bash |
| python data_curation/crawl_internal_repos.py \ |
| --repos \ |
| /path/to/your/library-core \ |
| /path/to/your/library-utils \ |
| /path/to/your/app-1 \ |
| /path/to/your/app-2 \ |
| /path/to/your/slurm-scripts \ |
| --output data/raw/internal_raw.jsonl \ |
| --include-tests \ |
| --include-docs \ |
| --repo-level |
| ``` |
|
|
| This produces: |
| - **File-level examples:** Individual Python files with metadata headers |
| - **Repo-level examples:** Combined package views showing cross-file dependencies |
| - **Test files:** Unit tests that teach the model expected behavior |
| - **Documentation:** READMEs, docstrings, markdown docs |
|
|
| ### Step 2: Generate Synthetic Data with an API Model |
|
|
| ```bash |
| export OPENAI_API_KEY="sk-..." |
| python data_curation/generate_synthetic_data.py \ |
| --input data/raw/internal_raw.jsonl \ |
| --output data/raw/internal_synthetic.jsonl \ |
| --api openai \ |
| --model gpt-4o \ |
| --tasks docs qa completion bugfix slurm reasoning \ |
| --num-per-file 5 \ |
| --max-parallel 10 |
| ``` |
|
|
| **What this generates for each source file:** |
|
|
| 1. **Documentation** β Full API docs, usage guides, tutorials |
| 2. **Q&A Pairs** β "How do I use X?" β answer with code examples |
| 3. **Code Completion** β Function stubs β full implementations |
| 4. **Bug-Fix Pairs** β Buggy code β explanation β fix |
| 5. **Slurm Scripts** β Job scripts using your library |
| 6. **Reasoning** β Step-by-step problem solving with `<think>` tags |
|
|
| **Cost estimate:** At ~$0.03/generation (GPT-4o), 500 files Γ 6 tasks Γ 5 examples = 15K generations = ~$450 |
|
|
| ### Step 3: Quality Filter and Deduplicate |
|
|
| ```bash |
| python data_curation/filter_and_dedup.py \ |
| --input data/raw/internal_raw.jsonl data/raw/internal_synthetic.jsonl \ |
| --output data/filtered/internal_filtered.jsonl \ |
| --min-length 100 \ |
| --max-length 100000 \ |
| --min-quality 0.3 \ |
| --dedup-threshold 0.7 |
| ``` |
|
|
| ### Step 4: Tokenize for Megatron-LM |
|
|
| ```bash |
| python data_curation/tokenize_for_megatron.py \ |
| --input data/filtered/internal_filtered.jsonl \ |
| --output-prefix data/tokenized/internal-libraries \ |
| --tokenizer openai/gpt-oss-120b \ |
| --workers 32 |
| ``` |
|
|
| This produces `internal-libraries_text_document.bin` and `.idx` files |
| that Megatron-LM reads directly during training. |
|
|
| ### Step 5: Prepare SFT Data |
|
|
| For SFT (Phase 2), you need data in ChatML format: |
|
|
| ```bash |
| python data_curation/generate_synthetic_data.py \ |
| --input data/raw/internal_raw.jsonl \ |
| --output data/sft/internal_library_qa.jsonl \ |
| --tasks qa reasoning \ |
| --num-per-file 10 |
| |
| # Convert to ChatML messages format: |
| python -c " |
| import json |
| |
| with open('data/sft/internal_library_qa.jsonl') as f, \ |
| open('data/sft/internal_sft_messages.jsonl', 'w') as out: |
| for line in f: |
| item = json.loads(line) |
| text = item['text'] |
| |
| # Parse Q&A into messages |
| if 'Question:' in text and 'Answer:' in text: |
| parts = text.split('Answer:', 1) |
| question = parts[0].replace('Question:', '').strip() |
| answer = parts[1].strip() |
| messages = [ |
| {'role': 'system', 'content': 'You are a helpful coding assistant with deep knowledge of our internal libraries and HPC systems.'}, |
| {'role': 'user', 'content': question}, |
| {'role': 'assistant', 'content': answer} |
| ] |
| out.write(json.dumps({'messages': messages}) + '\n') |
| " |
| ``` |
|
|
| --- |
|
|
| ## Slurm/Bash Data Curation |
|
|
| Since there's no standard Slurm dataset: |
|
|
| ### 1. Your own cluster scripts |
| ```bash |
| python data_curation/prepare_slurm_data.py \ |
| --output data/filtered/slurm_bash.jsonl \ |
| --search-dirs ~/slurm-scripts /shared/job-scripts \ |
| --include-manpages \ |
| --include-synthetic |
| ``` |
|
|
| ### 2. GitHub mining (optional but valuable) |
| Search GitHub for Slurm scripts: |
| ```bash |
| # Use GitHub API to find .sbatch files |
| # Example search: "SBATCH language:shell filename:*.sbatch" |
| # Collect, filter, and add to training data |
| ``` |
|
|
| ### 3. Documentation scraping |
| - Slurm docs: https://slurm.schedmd.com/documentation.html |
| - Convert HTML β Markdown for training |
|
|
| ### 4. Synthetic generation (most effective) |
| Use the API model to generate diverse Slurm scripts: |
| ```bash |
| python data_curation/generate_synthetic_data.py \ |
| --input data/raw/internal_raw.jsonl \ |
| --output data/raw/slurm_synthetic.jsonl \ |
| --tasks slurm \ |
| --num-per-file 10 |
| ``` |
|
|
| --- |
|
|
| ## Data Quality Tips |
|
|
| 1. **Upweight high-quality data in later epochs.** During annealing, increase the mix of |
| curated, instruction-quality data and decrease noisy web data. |
|
|
| 2. **Repeat internal data.** Your internal code is ~10% of the mix but might be <1% of total |
| tokens. The model sees it many more times than public data. This is intentional and beneficial β |
| it's similar to "curriculum learning" where important data is repeated. |
|
|
| 3. **Include test files.** Unit tests are incredibly valuable training data β they show the model |
| the expected input/output behavior of your libraries. The model learns "when I call X with Y, |
| it should return Z." |
|
|
| 4. **Include git history selectively.** Commit messages and diffs teach the model about code |
| evolution and debugging patterns. But raw git logs are noisy β filter for meaningful commits. |
|
|
| 5. **Decontaminate evaluation data.** Remove any overlap between your training data and |
| evaluation benchmarks (HumanEval, MBPP, etc.) to get honest benchmark scores. |
|
|