# Training Progress Last updated: 2026-06-27 ## Current State - `training/` contains the Stage 0 / Stage 1 recipe layer plus a tested data-engineering layer. - `training/scripts/run_sft.py` and `training/scripts/merge_lora.py` provide runnable LoRA/QLoRA training and merge entrypoints. - `training/scripts/check_training_gates.py` blocks training until the CyberGym baseline and processed JSONL data exist. - **Data engineering (new):** `datasets.yaml` manifest + `hf_download.py` (probe/download) + `sft_adapters.py` (per-source schema adapters) + `build_sft_dataset.py` (adapt → route `` → dedup → data card) + `synthesize_think.py` (rejection-sampling reasoning backfill). All schemas in the manifest were verified against the live HF datasets-server on 2026-06-27. - **Eval + decontam (new):** `decontaminate.py` (13-gram + fuzzy collision gate), `build_eval_sets.py` (held-out vuln-detection + disjoint knowledge-MCQ sets), and `eval_endpoint.py` (reusable scorer for the base baseline and every checkpoint). `check_training_gates.py` now hard-requires the decontamination report. - **Offline test green:** `python training/tests/test_data_pipeline.py` runs adapt → build → synthesize(`--mock`) → re-build → split, plus build-eval-sets → eval(`--mock`) → decontaminate (54 checks) with no GPU/heavy deps. - `normalize_sft_jsonl.py` / `split_jsonl.py` remain as the simple ad-hoc data path. - Local `vastai` CLI is installed and usable. - Hugging Face API confirms `Qwen/Qwen3.6-27B` exists, is public, Apache-2.0, and reports `model_type=qwen3_5`. - Docker Hub lists `pytorch/pytorch:2.12.1-cuda12.6-cudnn9-devel`, which is the default image in the Vast.ai create script. - CyberGym pre-training baseline is now a hard gate before QLoRA or LoRA starts. ## Data findings (2026-06-27) - PrimeVul / DiverseVul are raw `func`+`target` detection tables → handled by the new `detection_func_target` adapter (the generic normalizer rejected them outright). - `cwe` fields arrive as stringified JSON lists; adapters parse them. - MegaVul / PrimeVul / DiverseVul have `test` splits → reserved for held-out vuln-detection eval (`--eval`). - `SecCoderX/Reasoning_Vulnerability_Detection_SFT_Cold_Start` and `Leopo1d/OpenVul_Rejection_Sampling_Reasoning_SFT` return **HTTP 401** (gated) → disabled in the manifest until `huggingface-cli login` + access approval. ## Next Commands Safe offer search: ```bash NUM_GPUS=1 STORAGE_GB=650 bash training/scripts/vastai_search_h100.sh NUM_GPUS=2 STORAGE_GB=900 bash training/scripts/vastai_search_h100.sh ``` Cost-incurring create step, after choosing an offer: ```bash CONFIRM_CREATE=1 bash training/scripts/vastai_create_stage1.sh ``` Remote bootstrap: ```bash git clone https://github.com/ModelsLab/infosec.git cd infosec bash training/scripts/vastai_onstart.sh python training/scripts/phase0_qwen_smoke.py --model Qwen/Qwen3.6-27B --mode config python training/scripts/phase0_qwen_smoke.py --model Qwen/Qwen3.6-27B --mode load --device-map auto --dtype bfloat16 ``` Mandatory benchmark before training: ```bash sed -n '1,240p' training/recipes/pretraining_cybergym_baseline.md ``` Data prep recipe: ```bash sed -n '1,240p' training/recipes/data_preparation.md ``` Gate checks: ```bash python training/scripts/check_training_gates.py --config training/configs/pilot_qlora.yaml python training/scripts/check_training_gates.py --config training/configs/stage1_lora_sft.yaml ``` Training entrypoints after gates pass: ```bash torchrun --nproc_per_node=1 training/scripts/run_sft.py --config training/configs/pilot_qlora.yaml torchrun --nproc_per_node=2 training/scripts/run_sft.py --config training/configs/stage1_lora_sft.yaml python training/scripts/merge_lora.py --config training/configs/stage1_lora_sft.yaml ``` ## Data prep next commands ```bash python training/scripts/hf_download.py --probe-manifest training/configs/datasets.yaml # verify schemas python training/scripts/hf_download.py --all --profile pilot # cheap raw pull python training/scripts/build_sft_dataset.py --profile pilot --mix-name pilot # build + data card python training/scripts/synthesize_think.py --input data/think_synthesis/pilot.to_synthesize.jsonl \ --output data/think_synthesis/pilot.synthesized.jsonl --mock # swap --mock for a real TEACHER_BASE_URL python training/scripts/build_sft_dataset.py --profile pilot --mix-name pilot \ --include-synthesized data/think_synthesis/pilot.synthesized.jsonl python training/scripts/split_jsonl.py --input data/processed/pilot.ready.normalized.jsonl \ --train data/processed/pilot_train.chatml.jsonl --val data/processed/pilot_val.chatml.jsonl --val-ratio 0.02 ``` ## Blockers Before Real Training - DONE: Hugging Face download/export scripts and per-source schema adapters now exist and are tested. - DONE: rejection-sampling `` synthesis driver exists (validated with `--mock`; needs a real teacher endpoint). - DONE: decontamination gate (13-gram + fuzzy; optional embedding) + held-out eval sets + reusable endpoint scorer. - NEXT (needs GPU host): provision Vast.ai, Phase-0 load gate, serve base vLLM, capture base eval + CyberGym baseline, then train. - Gated reasoning sets need `huggingface-cli login` + access before they can be enabled. - Processed SFT JSONL is generated by the pipeline but not committed (gitignored); run the data-prep commands above. - The Qwen3.6 hybrid architecture must pass the Phase 0 full-load test on the rented image before running SFT. - CyberGym baseline data/server is not installed on a GPU/Docker host yet. - Training must not start until the base-model CyberGym smoke and frozen level1 baseline reports exist.