File size: 5,775 Bytes
994182c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 | # Infosec Qwen Training
This directory is the executable handoff for `docs/PROJECT_GOAL.md`.
Current target:
- Base model: `Qwen/Qwen3.6-27B`
- Primary benchmark: CyberGym
- First training run: Stage 1 LoRA SFT, with a smaller QLoRA pilot first
- GPU provider: Vast.ai through the local `vastai` CLI
- Reasoning policy: keep Qwen thinking mode on and require `<think>...</think>` in assistant training turns
## Files
- `recipes/stage1_qwen36_27b_lora_sft.md` - runbook for Stage 0 smoke tests, the QLoRA pilot, and Stage 1 LoRA SFT.
- `recipes/pretraining_cybergym_baseline.md` - mandatory base-model CyberGym benchmark before training.
- `benchmarks/cybergym_smoke_tasks.txt` - official 10-task subset for the pre-training smoke benchmark.
- `configs/baseline_cybergym.yaml` - CyberGym baseline config and artifact contract.
- `configs/pilot_qlora.yaml` - small, cheap pipeline validation config.
- `configs/stage1_lora_sft.yaml` - main Stage 1 LoRA SFT config.
- `configs/data_mix.yaml` - dataset groups, caps, and formatting rules derived from the project goal.
- `requirements-cu126.txt` - remote Python environment requirements for the PyTorch CUDA 12.6 image.
- `recipes/data_preparation.md` - manifest-driven download → adapt → synthesize → split workflow.
- `configs/datasets.yaml` - per-source download + adapter manifest (HF ids, splits, license, auth, think status), schemas verified against the live datasets-server.
- `scripts/hf_download.py` - probe dataset schemas (stdlib) and download splits to raw JSONL (`datasets`).
- `scripts/sft_adapters.py` - per-source schema adapters (detection, vuln/fix pairs, conversations, MCQ, DPO-to-SFT, etc.).
- `scripts/build_sft_dataset.py` - orchestrator: adapt → route `<think>` (ready vs synthesize) → dedup → source caps → data card.
- `scripts/synthesize_think.py` - rejection-sampling `<think>` backfill against a teacher endpoint (with `--mock` offline mode).
- `scripts/normalize_sft_jsonl.py` - ad-hoc converter for a single common-schema JSONL export.
- `scripts/split_jsonl.py` - deterministic train/validation split helper.
- `scripts/decontaminate.py` - removes train rows overlapping eval sets (13-gram + fuzzy, optional embedding); emits the decontam report gate.
- `scripts/build_eval_sets.py` - builds held-out vuln-detection + knowledge-MCQ scoring sets from reserved splits.
- `scripts/eval_endpoint.py` - scores any OpenAI-compatible endpoint on those sets (base baseline + per-epoch), comparable reports.
- `scripts/run_base_baseline.sh` - one command on the GPU host: build eval sets, serve base vLLM, score the base model.
- `tests/test_data_pipeline.py` - offline end-to-end test of the data + eval + decontam layer (54 checks, no GPU/heavy deps).
- `scripts/vastai_search_h100.sh` - safe offer search command. This does not rent hardware.
- `scripts/vastai_create_stage1.sh` - gated instance creation command. This rents hardware only when `CONFIRM_CREATE=1`.
- `scripts/vastai_onstart.sh` - Vast.ai startup/bootstrap script.
- `scripts/phase0_qwen_smoke.py` - model/config/tokenizer load check before spending on training.
- `scripts/check_training_gates.py` - refuses to start training when baseline/data gates are missing.
- `scripts/cybergym_verify_agent_logs.py` - verifies OpenHands CyberGym runs into JSONL.
- `scripts/summarize_cybergym_results.py` - turns verifier JSONL into a markdown score report.
- `scripts/run_sft.py` - Transformers/PEFT LoRA or QLoRA SFT runner.
- `scripts/merge_lora.py` - merge selected adapter into a standalone checkpoint.
- `PROGRESS.md` - current status and next commands.
## Quick Start
Search for a pilot GPU:
```bash
NUM_GPUS=1 STORAGE_GB=650 bash training/scripts/vastai_search_h100.sh
```
Search for the main Stage 1 GPU shape:
```bash
NUM_GPUS=2 STORAGE_GB=900 bash training/scripts/vastai_search_h100.sh
```
Create an instance from a chosen offer id:
```bash
CONFIRM_CREATE=1 bash training/scripts/vastai_create_stage1.sh <offer-id>
```
After SSH into the instance:
```bash
git clone https://github.com/ModelsLab/infosec.git
cd infosec
bash training/scripts/vastai_onstart.sh
python training/scripts/phase0_qwen_smoke.py --model Qwen/Qwen3.6-27B --mode config
```
Before training, run the CyberGym baseline:
```bash
sed -n '1,240p' training/recipes/pretraining_cybergym_baseline.md
```
Check gates before a pilot or Stage 1 run:
```bash
python training/scripts/check_training_gates.py --config training/configs/pilot_qlora.yaml
python training/scripts/check_training_gates.py --config training/configs/stage1_lora_sft.yaml
```
Prepare the SFT data (see `recipes/data_preparation.md` for the full flow):
```bash
python training/scripts/hf_download.py --probe-manifest training/configs/datasets.yaml
python training/scripts/hf_download.py --all --profile pilot
python training/scripts/build_sft_dataset.py --profile pilot --mix-name pilot
python training/tests/test_data_pipeline.py # offline sanity check, no GPU needed
```
Run SFT after the baseline and data prep exist:
```bash
torchrun --nproc_per_node=1 training/scripts/run_sft.py --config training/configs/pilot_qlora.yaml
torchrun --nproc_per_node=2 training/scripts/run_sft.py --config training/configs/stage1_lora_sft.yaml
```
Run the full load gate only after confirming the rented GPU has enough VRAM and disk:
```bash
python training/scripts/phase0_qwen_smoke.py --model Qwen/Qwen3.6-27B --mode load --device-map auto --dtype bfloat16
```
## Safety Boundary
This repo is for authorized security research and education. Training/eval artifacts should stay scoped to public benchmark tasks, public vulnerability datasets, and owned or explicitly authorized environments. Do not use the recipe to run exploit attempts against third-party live systems.
|