File size: 30,238 Bytes
994182c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 | # Project Goal: CyberGym-Targeted Fine-Tune of Qwen3.6-27B
**Status:** Planning / pre-implementation
**Owner:** Adhik Joshi Β· **Context:** college research project (authorized security research / education)
**Last updated:** 2026-06-27
**Repo:** `infosec` (greenfield)
---
## 0. TL;DR
Fine-tune the open-weight **Qwen3.6-27B** (reasoning + agentic, Apache-2.0) to raise its score on **CyberGym** β UC Berkeley's agentic vulnerability-reproduction benchmark β using a curated mix of **offensive + defensive** cybersecurity data, with **`<think>` reasoning traces**, trained as **LoRA** (then merged).
**Confirmed decisions (2026-06-27):**
1. **CyberGym is the PRIMARY objective** (success rate β), not knowledge benchmarks.
2. Data must be **offensive + defensive**.
3. **LoRA** for the trial; **merge** into the base checkpoint if it works (`merge_and_unload`).
4. **`<think>` is required** β the model must reason through tasks.
**The hard truth this creates (read this):** CyberGym is won by an **agentic loop** β explore a C/C++ repo over ~100 tool-calling steps and write a PoC input file that crashes the *pre-patch* build but not the *patched* one. Dedicated 32B *coding-agent* models still score β€2%; frontier models top out ~18β22%. **Supervised fine-tuning on Q&A β even great offensive data β has a low ceiling here.** To actually move CyberGym we need a staged program:
> **SFT (LoRA) on code-vuln + reasoning + agentic-trajectory data β DPO on secure/insecure code β RL on the CyberGym task itself (verified pass/fail reward).** The RL stage is the real lever; SFT makes the model *coherent and reason-y* enough for RL to work.
This document plans all stages, but **Stage 1 (LoRA SFT)** is the trial. Expect SFT to deliver modest CyberGym gains + large gains on code-vuln/knowledge evals; budget for the RL stage to get a big CyberGym jump.
---
## 1. Objective & Success Criteria
**Primary:** increase CyberGym success rate over the base Qwen3.6-27B, without regressing tool-calling / agentic ability.
| Metric | Tier | Baseline | Target |
|---|---|---|---|
| **CyberGym success rate** (level1 sample, OpenHands + vLLM) | **PRIMARY** | base (expect low single digits) | **> base; aim +2β5+ pts via SFT, more via RL** |
| CyberGym 10-task subset β valid tool-calls emitted | Guardrail | base | **no tool-call regression** |
| Vulnerability-detection accuracy (C/C++) | Secondary | base | **+5β15 pts** |
| Cyber-knowledge MCQ (CyberMetric/SecEval/CTI-Bench) | Secondary | base | **+5β15 pts** |
| IFEval / MMLU-general | Guardrail | base | **β₯ base β 2 pts** (no forgetting) |
**Definition of done (Stage 1):** a merged LoRA checkpoint + eval report (CyberGym + code-vuln + guardrails + loss curves) showing net improvement and intact agentic ability, plus a reproducible pipeline in this repo. **Stretch (Stage 2+):** DPO + RL for a larger CyberGym jump.
---
## 2. Base Model β Qwen3.6-27B (verified)
- **Real & live**, released 2026-04-24, ~5.6M downloads, **Apache-2.0** (clean commercial/research use, not gated).
- **~27.8B params**, weights **55.6 GB bf16**. **Dense hybrid attention:** 64 layers = 48 Gated-DeltaNet (linear) + 16 full-attention GQA layers. Hidden 5120, FFN 17408.
- **Reasoning model** with a ChatML **`<think>` channel ON by default** (`enable_thinking=false` injects an empty think block). **Matches your `<think>` requirement natively.**
- **Multimodal** (vision tower) β **freeze it** for text/code data.
- **Untied embeddings**, vocab **248,320**; **no BOS**, EOS = `<|im_end|>`. Native context **256K**.
- **MTP head** (speculative decoding) β freeze/ignore in SFT.
- **Framework risk (Phase-0 gate):** custom hybrid arch (`model_type` ~`qwen3_5`) needs **very recent `transformers` (β₯4.57.1)** + FLA/`causal-conv1d` kernels. Older Axolotl/Unsloth/PEFT may not support it.
- **Fallback if blocked:** `Qwen3-32B` (dense, standard attention, text-only, well-supported) β also has a thinking mode. Memory math scales Γ32/27.
---
## 3. Datasets
### 3.A The CyberGym-aligned set (NEW β the focus going forward)
Verified via HF API (rows/schema/license). Ranked by relevance to the agentic vuln-reproduction task.
| Tier | Dataset | Rows | Schema | Why it matters for CyberGym | License |
|---|---|---|---|---|---|
| **1 Β· C/C++ vuln code** | `athrv/megavul-vulnerability-detection` | 353,873 | instruction/input/output + cwe_ids, cve_id, severity, cvss, repo, func, is_vulnerable | Real C/C++ functions w/ CVE/CWE labels β core memory-safety understanding | β οΈ none stated |
| **1 Β· C++ vuln code** | `smart-code-analyzer-team/cpp-vulnerability-dataset` | 69,299 | raw_snippet, security_context, data_flow_graph, ast_metadataβ¦ | C++ snippets w/ security context + data-flow β pattern recognition | β οΈ none stated |
| **1 Β· CVE patches** | `DetectVul/CVEFixes` (+ `hitoshura25/cvefixes`) | 5,730 | lines, raw_lines, label, type | Pre/post-patch diff lines β mirrors CyberGym's patch differential | β οΈ none stated |
| **1 Β· SARD/Juliet C** | `lynchorange/SARD_Vulnerability_Detection_C` | (n/a) | C test cases | NIST-style labeled vulnerable C | β οΈ verify (SARD is public-domain) |
| **2 Β· Reasoning `<think>`** | `SecCoderX/...Reasoning_Vulnerability_Detection_SFT_Cold_Start` | 37,113 | instruction, output | Vuln-detection **with reasoning** β feeds the `<think>` channel | β οΈ none stated |
| **2 Β· Reasoning `<think>`** | `Leopo1d/OpenVul_Rejection_Sampling...Reasoning...SFT` (+ GT companion) | 6,276 | key, prompt, completion | **Reject-sampled reasoning traces** β exactly our `<think>` method, pre-built | Apache-2.0 |
| **3 Β· Agentic loop (top)** | `nvidia/Nemotron-Terminal-Corpus` | ~366,000 | conversations (terminal commandβobserve) | **Closest public analogue to CyberGym's run-command-observe loop** | CC-BY-4.0 |
| **3 Β· Agentic loop** | `Kwai-Klear/SWE-smith-mini_swe_agent_plus-trajectories-66k` (bash-only), `SWE-bench/SWE-smith-trajectories` (76k), `nebius/SWE-agent-trajectories` (80k) | 66β80k | messages/trajectory + patches + eval logs | Teach the multi-turn tool loop (Python-centric β use for loop *mechanics*, not domain) | MIT / CC-BY-4.0 |
| **3 Β· Reasoning+tools** | `interstellarninja/tool-use-multiturn-reasoning` | 14,600 | conversations, tools (5β13 turns, interleaved `<think>`) | Callβreasonβobserve rhythm **with `<think>`** | Apache-2.0 |
| **3 Β· RL warm-start** | `nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1` | 50,000 | tool-call traj + `pass_rate` | Verifiable-reward format for Stage-3 RL | CC-BY-4.0 |
| **3 Β· Format blueprint** | `SecureAI-SE/fuzz-agent-trajectories` | **10** | thinkβact(tool)βobserve, offensive fuzzing | Exact offensive-agent format match β **seed/template only, not trainable volume** | CC-BY-4.0 |
| **2 Β· Reasoning `<think>` (cyber)** | `theelderemo/pentesting-explanations` | 5,900 | MCQ + explanation + `<think>` | Offensive pentest reasoning with `<think>` | Apache-2.0 |
| **3 Β· Tool-output skill** | `tegridydev/infosec-tool-output` | 776 | SAST findings β next-action | Teaches the "read security-tool output" sub-skill | Apache-2.0 |
| **4 Β· DPO** | `CyberNative/Code_Vulnerability_Security_DPO` | 4,656 | lang, vulnerability, system, question, chosen, rejected | Ready-made secure-vs-vulnerable **preference pairs** for the DPO stage | Apache-2.0 |
| **5 Β· CVE knowledge** | `AlicanKiraz0/All-CVE-Records-Training-Dataset` | 297,441 | System/User/Assistant | Broad CVE breadth (defensive context/recall) | Apache-2.0 |
| **5 Β· Threat-intel (def.)** | `reloading0101/threat-intelligence-dataset` | ~9,200 | CTI Q&A fact-checked vs MITRE/CISA/EPSS | High-provenance defensive breadth | CC-BY-4.0 |
| **6 Β· RL env** | `AdithyaSK/repo2rlenv-cve-patches` | (n/a) | CVE patches β RL env | CVE patches packaged as RL environments β Stage-3 fuel | Apache-2.0 |
**Dedup flags:** `Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset` is **byte-identical** to `AlicanKiraz0/All-CVE-Records-...` (same 297,441 rows / same size) β keep one. (Same pattern as the Soban/Trendyol mirror from the original list.)
**License flags:** several Tier-1 sets have **no stated license**. For a college research project this is usually defensible (academic use of public, derived-from-open-source data), but **record provenance and prefer the licensed equivalents where possible** (OpenVul Apache-2.0, CVEFixes upstream is CC, SARD is NIST public-domain).
**Gated/non-commercial:** `WhitzardAgent/CyberSecurity-1M` & `-100B` are the best-curated recent cyber corpora (exploitation chains, Sigma/YARA, CoT) **but are gated behind an institutional-email approval (a personal Gmail will be rejected) and licensed academic/non-commercial** β usable for a college project only if access is granted; do not assume availability.
> **β οΈ Structural gap (the key dataset finding):** **No public dataset contains end-to-end CyberGym-style C/C++ "PoC β crash β iterate" agent trajectories.** The only exact-format match (`SecureAI-SE/fuzz-agent-trajectories`) is 10 rows β a blueprint, not training volume. Consequence for the plan: the agentic-loop skill must come from **(a)** transferring loop *mechanics* from the generic terminal/SWE trajectory sets above, **(b)** **synthesizing** our own trajectories (run an agent harness over training-split CyberGym/ARVO tasks, keep successful runs as SFT data), and/or **(c)** **Stage-3 RL**, which learns the loop directly from the environment. This is why RL (not SFT) is the real lever for the primary metric.
### 3.B The original 6 (now supplementary)
Still useful as **defensive breadth + instruction-following**, but secondary to 3.A for the CyberGym goal:
- β
`AlicanKiraz0/Fenrir-v2.1` (99,870, EN, Apache-2.0), `Trendyol/Cybersecurity-IT` (53,202, EN, Apache-2.0), `hcnote/High-Quality` (270,271, ZH+EN, MIT, **has offensive/exploit content + code** β useful here).
- β `jobby32/QA-with-negatives` (17,263) β `questionβpositive` SFT only.
- β `Soban1234/...` (mirror of Trendyol β drop). β `clouditera/security-paper` (raw bilingual book text, **unspecified license** β exclude; optional CPT only after legal review).
### 3.D Platform / language breadth (NodeJS Β· Python Β· PHP Β· Apache Β· Linux Β· web)
CyberGym's *core* is C/C++ memory-safety bugs (OSS-Fuzz/ARVO), and **Apache httpd, nginx, and the Linux kernel are themselves C** β so the C/C++ sets above already cover the "systems software" targets. **NodeJS / Python / PHP** are different vuln classes (injection, deserialization, SSRF, path traversal) β valuable for *breadth* and useful if we add web targets, but **secondary to C/C++ for the CyberGym score**. Honest framing: prioritize the first group for CyberGym; add the second for a well-rounded offensive+defensive model.
| Group | Dataset | Rows | Schema | Coverage | License |
|---|---|---|---|---|---|
| **Multi-language (best for your ask)** | `hitoshura25/crossvul` | 9,313 | cwe_id, language, **vulnerable_code, fixed_code**, source | **PHP/JS/Python/C/C++/Java/Ruby/Goβ¦** vuln+fix pairs (CrossVul) | Apache-2.0 |
| C/C++ (large, real-world) | `bstee615/diversevul` | 330,492 | func, target, cwe, project, commit | C/C++ across many OSS projects (incl. Apache/Linux-type) | β οΈ none |
| C/C++ (highest quality) | `colin/PrimeVul` | 233,941 | func, target, cwe, project, commit | Rigorously de-noised C/C++ vuln detection | MIT |
| Linux/systems (C) | `mjbommar/linux-cve-dossiers` | 2,556 | cve_id, package, scope, completeness⦠| **Linux CVE** dossiers (kernel/distro packages) | CC-BY-4.0 |
| Web / OWASP (knowledge) | `shahrukh95/OWASP-and-NVD-question-answer-dataset` | 22,958 | Question, Answer, split | OWASP + NVD web-vuln Q&A | β οΈ none |
| PHP (offensive/defensive) | `nbuser32/PHP-Webshell-Dataset` | 30,270 | instruction/input/output | **PHP webshell** detection/analysis | MIT |
| Web injection (NodeJS/PHP/Python apps) | `PurpleAILAB/chatML_SQL_injection_dataset` (+ XSS sets) | ~2,000 | conversations (ChatML) | SQLi/XSS exploitation in chat format | β οΈ none |
| Python | `maddyrucos/code_vulnerability_python` | 848 | func, target, project | Small Python vuln set | MIT |
Notes: `nasirhm/reposvul` (repo-level, multi-language β adds whole-repo context like CyberGym) and `Humanlearning/CyberSecurity_OWASP-sft-dataset` looked relevant but their HF viewers returned empty β **verify at source before relying on them.** Several web sets are tiny/no-license β use as flavor, not backbone; record provenance.
### 3.C Corpus shape (planning)
- **Code/vuln + reasoning core (3.A tiers 1β2):** ~470K rows.
- **Agentic trajectories (3.A tier 3):** ~80K+ (heavy per-example token count β these are long).
- **DPO pairs:** ~5K. **CVE/defensive breadth:** ~300K (deduped) + original-6 supplementary.
- After dedup + decontamination, plan for **~400β600K SFT examples**, token count dominated by the long agentic trajectories. Cap any single source at β€30β40% of tokens.
---
## 4. Benchmark β CyberGym (verified) + the eval suite
**CyberGym** (sunblaze-ucb, arXiv 2506.02548, ICLR 2026): given a pre-patch C/C++ codebase + a vuln description, an **agent** must produce a **PoC input file** that triggers a sanitizer crash on the pre-patch build **and not** on the patched build (differential validation). **1,507 tasks / 188 projects**; levels 0β3 (more info β easier); a **10-task curated subset** ships for quick tests. Harness is **agentic + Docker**, uses **LiteLLM** β supports **any OpenAI-compatible endpoint** (our vLLM). Use **binary-only mode (~130 GB)**. Baselines: GPT-5 ~22%, Sonnet-4 ~18%, open 32B coding agents β€2%.
**Self-hosted eval path (confirmed):** serve the merged checkpoint with **vLLM** (`--served-model-name my-qwen`, tool-calling enabled) β run OpenHands with `--model openai/my-qwen --llm.base_url http://HOST:PORT/v1`, `LLM_API_KEY=EMPTY`, `--difficulty level1`.
**Eval suite we'll report:**
- **PRIMARY:** CyberGym β 10-task smoke (tool-call sanity) β **~100β200 task level1 sample** (real number) β full 1,507 only if budget allows.
- **Secondary:** held-out **C/C++ vulnerability-detection** accuracy (from MegaVul/CVEFixes test splits); cyber-knowledge MCQ (CyberMetric, SecEval, CTI-Bench).
- **Guardrails:** IFEval (instruction following), MMLU-general (forgetting).
- **Decontaminate** all training data against every eval set (13-gram + fuzzy + embedding); freeze test splits.
---
## 5. Training Strategy (staged)
### Stage 0 β Infra & go/no-go
Confirm Qwen3.6-27B loads + LoRA-trains + serves under vLLM with tool-calling on current `transformers`. If blocked, fall back to Qwen3-32B. Set up storage (~236 GB CyberGym / ~130 GB binary server), W&B, vLLM.
### Stage 1 β LoRA SFT (the trial; "benchmark after each epoch")
**Goal:** make the model fluent in C/C++ vuln reasoning + the agentic tool-use format, with `<think>` traces.
- **Data:** 3.A tiers 1β3 (code-vuln + reasoning + agentic trajectories) + DPO-source positives + defensive breadth, all normalized to ChatML with **`<think>` reasoning**.
- **`<think>` policy:** keep reasoning **ON**. Use the pre-built reasoning datasets (SecCoderX, OpenVul) directly; for code-vuln sets lacking traces, **synthesize `<think>` via rejection sampling** (generate N reasonings with a strong model, keep those that reach the correct label/patch). Apply the format uniformly β inconsistent `<think>` is the #1 way to damage a Qwen reasoning model.
- **Method:** LoRA **r=64, Ξ±=128, rsLoRA on, all linear targets, dropout 0.05**, freeze vision tower + MTP head.
- **Hyperparameters:** lr **2e-4** cosine, warmup 0.03β0.05, **2β3 epochs**, effective batch ~128, **seq 8192** (agentic trajectories are long), packing on, gradient checkpointing on, **loss-masked on assistant tokens**, paged AdamW-8bit.
- **Pilot first:** a **QLoRA** smoke-test on a subset (1Γ 80GB, ~$20β60) to validate the pipeline before the real LoRA run.
- **Per-epoch eval:** at each epoch checkpoint run the fast eval (vuln-detection slice + IFEval + CyberGym 10-task). Select the checkpoint by **best eval + non-rising val loss**, never lowest train loss.
### Stage 2 β DPO (optional, cheap)
Use `CyberNative/Code_Vulnerability_Security_DPO` (chosen=secure, rejected=vulnerable) to sharpen secure/insecure discrimination. LoRA-DPO, lr ~5e-6, 1 epoch. Small, fast, low risk.
### Stage 3 β RL on the CyberGym task β "the real lever" (detailed)
**Why RL fits CyberGym perfectly.** CyberGym hands us a *verifiable, machine-checkable, hard-to-game* reward: the submitted PoC either crashes the pre-patch build **and not** the patched build, or it doesn't. That is the ideal setup for **RLVR (RL with Verifiable Rewards)** β the same paradigm behind reasoning models (e.g. DeepSeek-R1). No reward model to train; **the environment is the judge.** The catch: this is *agentic, long-horizon, sparse-reward* RL (a success needs ~100 correct tool-calling steps, reward only at the end), which is much harder than single-turn math RLVR. So we climb a ladder from cheap to powerful.
**The ladder (do them in order):**
1. **Rejection-sampling Fine-Tuning / Expert Iteration (STaR / RFT) β START HERE, highest ROI.**
- Run the SFT model on **training-split** tasks, many high-temperature rollouts each.
- Keep only trajectories that produced a **validated PoC (reward = 1)**.
- SFT on those winning trajectories (with their `<think>` reasoning).
- Repeat β as the model improves it solves more tasks β generates more training data β improves further.
- This is "RL" via filtered behavior cloning: **no PPO/GRPO infrastructure, far cheaper and more stable**, and it directly closes the "no public C/C++ agent-trajectory data" gap by *manufacturing our own*. For a college budget this alone is likely the biggest single win.
2. **GRPO (Group Relative Policy Optimization) β the on-policy push.**
- Modern policy-gradient for verifiable rewards: sample a *group* of rollouts per task, advantage = reward normalized within the group, **no value/critic model** (cheaper than PPO). Best applied *after* RFT has warmed the model up.
- Alternatives: PPO (needs a critic, heavier), RLOO, or offline trajectory-DPO (chosen = solved run, rejected = failed run) as a lightweight stepping stone.
3. **Curriculum + reward shaping (to fight sparsity).**
- **Curriculum:** start at **level3** (easiest β includes the patch + crash stack trace), then level2 β level1 β level0; or start from CyberGym's **5 "solvable" tasks**. Get signal flowing before tackling hard tasks.
- **Shaped reward (dense, but outcome-dominant):**
- + locate the vulnerable function/file,
- + produce *any* sanitizer crash on the pre-patch build (even the wrong bug),
- **+ full reward for the differential success** (crash pre-patch, clean post-patch β the true objective),
- + small format reward (valid tool-call syntax, `<think>` present, valid PoC file written),
- β small per-step/length penalty for efficiency.
- Keep the differential-success term dominant so the model can't farm the shaping terms.
4. **Anti-reward-hacking.** The differential check already blocks the cheapest exploit (triggering an unrelated crash). Also: keep eval tasks **completely out** of RL training; watch for the agent crashing/escaping the harness itself; cap the tool/step budget.
**Environment & data.**
- The **CyberGym Docker harness *is* the RL environment** β rollouts POST candidate PoCs to the validation server β binary reward. Use **binary-only mode (~130 GB)** for rollout speed.
- **Mandatory train/eval split:** RL-train on a held-in subset of the 1,507 tasks (+ external verified tasks: ARVO / OSS-Fuzz reproducible bugs, `AdithyaSK/repo2rlenv-cve-patches`), evaluate on a **disjoint** held-out subset. Never train on eval tasks. More distinct environments β better generalization.
**Frameworks.**
- **verl** (ByteDance) β leading scalable RLVR/GRPO framework; supports multi-turn agentic rollouts + tool calling; pairs with vLLM. Top pick for the full RL run.
- **OpenRLHF** β PPO/GRPO/RLOO at scale. **SkyRL / RAGEN / Agent-R1 / verifiers** β purpose-built for *multi-turn agentic* RL (the hard part is rollouts that interleave model β environment). **TRL `GRPOTrainer`** β simplest, good for RFT / single-turn proxies.
- The **RFT phase needs no RL framework** β just vLLM (rollouts) + the CyberGym harness (scoring) + a normal SFT trainer.
**Compute reality (why this is the costly phase).** Agentic RL is **rollout-dominated**: each episode = up to ~100 LLM calls + Docker execution, and GRPO needs a *group* (8β16) of episodes per task per step β thousands of ~100-step rollouts. You need a **vLLM inference cluster** (generation) + **many parallel Docker workers** (environment) + **training GPUs** (updates); **throughput of inference/env, not the gradient step, is the bottleneck.** Budget **$1β5K+** and it can balloon β which is exactly why RFT (cheap, off-policy) should capture most of the gain first.
**Recommended RL path (college budget):** SFT (Stage 1) β **RFT/Expert-Iteration** on the training split (prioritize this) β *if budget remains* **GRPO with curriculum + shaped reward via verl** β evaluate on the held-out split throughout.
**Realistic expectation.** Frontier models sit at ~18β22% on CyberGym; dedicated 32B coding agents score β€2% *without* this program. A 27B model after SFT + RFT (+ optional GRPO) can credibly aim for **mid-single digits β low double digits at level1** β a strong, publishable result for an open 27B in a college research setting.
### LoRA β merge
After Stage 1 (and DPO), `peft.merge_and_unload()` folds adapters into base weights β standalone checkpoint, zero inference overhead. **bf16 LoRA merges cleanly; QLoRA merge is lossy** (merge into dequantized bf16, then re-quantize for serving if needed). Re-benchmark the merged model to confirm parity before continuing to RL.
---
## 6. GPU, Time & Cost
~470K SFT rows incl. long agentic trajectories β plan **~600Mβ1B tokens/epoch** (trajectories dominate). 8Γ H100 β $16β24/hr; 1β2Γ H100 β $2β3/GPU-hr.
| Stage | Method | Hardware | Wallclock | Compute $ |
|---|---|---|---|---|
| Pilot | QLoRA, subset | 1Γ H100 80GB | ~6β12 hr | ~$20β60 |
| **1. LoRA SFT** (2β3 ep) | LoRA r=64, bf16 base | **2Γ H100 80GB** (or 8Γ FSDP) | ~2β4 days on 2Γ (~Β½ day on 8Γ) | **~$200β500** |
| 2. DPO | LoRA-DPO | 1β2Γ H100 | ~6β12 hr | ~$30β80 |
| 3. RL (CyberGym reward) | GRPO + Docker rollouts | 4β8Γ H100 + Docker host | daysβweeks | **$1β5K+** (rollout-dominated) |
| CyberGym eval | vLLM + OpenHands | 1Γ H100 + Docker | ~1β2 days/sample | ~$50β200 |
**VRAM:** LoRA bf16 base 55.6 GB + adapters + long-seq activations β **1Γ 80GB (tight, seqβ€4096) or 2Γ 80GB (comfortable, seq 8192)**. Full FT (if ever) ~450β520 GB β 8Γ 80GB ZeRO-3.
**Budget:** Stage 1 trial **~$300β700** all-in (data prep + pilot + run + eval iterations). Add **$1β5K+** only if you commit to Stage-3 RL. Storage ~370 GB.
---
## 7. Loss Curve & Monitoring
- **SFT loss:** starts ~1.2β1.6 (post-warmup), converges ~0.6β0.9. Below ~0.4 = memorization. Agentic-trajectory tokens may sit a bit higher.
- **Healthy:** sharp early drop β smooth decline β flatten; **val tracks train** with small stable gap.
- **Overfitting:** val plateaus/rises while train falls; **sawtooth at epoch boundaries** = memorizing β stop at 2β3 epochs.
- **Track (W&B):** train/val loss, grad norm, LR, tokens/s, **per-epoch eval** (vuln-detection + IFEval + CyberGym-10). For RL: reward curve, KL to ref, pass-rate on held-in tasks.
- **Selection rule:** best **eval score + non-rising val loss**, never lowest train loss (guards against agentic/tool-call regression).
---
## 8. Risks & Mitigations
| Risk | Likelihood | Mitigation |
|---|---|---|
| SFT barely moves CyberGym (it's agentic/RL-shaped) | **High** | Set expectations; Stage-1 win = vuln-detection/knowledge + agentic coherence; Stage-3 RL for the CyberGym jump |
| Framework can't train the hybrid arch | **High** | Phase-0 go/no-go; fall back to Qwen3-32B |
| SFT regresses tool-calling/reasoning | Medium | LoRA (low forgetting); uniform `<think>`; blend agentic trajectories; per-epoch guardrail evals |
| Missing-license Tier-1 data | Medium | Record provenance; prefer licensed equivalents; academic-use framing; legal review before any release |
| Benchmark contamination | Medium | 13-gram + fuzzy + embedding decontam; frozen test splits |
| RL harness cost/complexity | Medium | Gate Stage 3 on Stage-1 success; start with a small task subset |
| `<think>` inconsistency damages reasoning | Medium | Single uniform policy; use pre-built reasoning sets + rejection-sampled traces |
---
## 9. Open Decisions (most now resolved)
β
Primary = CyberGym Β· β
offensive+defensive Β· β
LoRAβmerge Β· β
`<think>` required.
Remaining:
1. **Commit to Stage-3 RL** now (changes budget to $1β5K+), or decide after Stage 1? *(Recommend: decide after Stage 1.)*
2. **Base model:** lock Qwen3.6-27B, pre-authorize Qwen3-32B fallback if Phase-0 blocks? *(Recommend: yes.)*
3. **Reasoning-trace synthesis:** budget for rejection-sampling `<think>` traces over code-vuln sets (needs a strong teacher model + inference $)? *(Recommend: yes β it's core to the `<think>` goal.)*
---
## 10. Phased Milestones
| Phase | Output | Gate |
|---|---|---|
| 0 Infra & go/no-go | Qwen3.6-27B trains+serves (or fallback) | Architecture trainable? |
| 1 Baseline | Base scores on CyberGym + code-vuln + guardrails | Reproducible; must finish before QLoRA/LoRA training |
| 2 Data engineering | Normalized+deduped+decontam mix w/ `<think>` traces + data card | Decontam verified |
| 3 Pilot (QLoRA) | Loss curve + sane gens + 1 eval cycle | No tool-call regression |
| 4 **LoRA SFT + per-epoch eval** | Checkpoints + curves + eval table | Hit Stage-1 criteria |
| 5 Merge + benchmark | Merged checkpoint, CyberGym level1 sample vs base | Net improvement, no regression |
| 6 (Opt) DPO | Sharpened secure/insecure model | Cheap win |
| 7 (Opt) RL on CyberGym | v2 with real CyberGym jump | Cost/benefit justified |
---
## 11. Proposed Repo Structure
```
infosec/
βββ docs/PROJECT_GOAL.md
βββ data/{download,processed,decontam,think_synthesis}/ # fetch, normalize, decontaminate, reasoning-trace synthesis
βββ training/{recipes,configs,scripts,tests}/ # Vast.ai + Qwen recipe, data manifest/adapters, offline tests
βββ eval/{cybergym,vuln_detection,knowledge,judge}/
βββ serve/ # vLLM OpenAI-compatible + tool-calling
βββ reports/ # baseline, per-epoch, merged, RL
```
## 12. Current Repo Artifacts
The implementation layer lives under `training/`:
- `training/README.md` β entrypoint and quick-start commands.
- `training/recipes/pretraining_cybergym_baseline.md` β mandatory base-model CyberGym benchmark before training.
- `training/recipes/stage1_qwen36_27b_lora_sft.md` β Stage 0 load gate, QLoRA pilot, and main LoRA SFT recipe.
- `training/configs/*.yaml` β pilot, Stage 1, and data-mix contracts.
- `training/scripts/vastai_*.sh` β Vast.ai search/create/bootstrap scripts.
- `training/scripts/phase0_qwen_smoke.py` β config/tokenizer/full-load smoke gate for `Qwen/Qwen3.6-27B`.
- `training/scripts/run_sft.py` and `training/scripts/merge_lora.py` β runnable PEFT training and merge entrypoints.
**Data-engineering layer (implemented + offline-tested, 2026-06-27):**
- `training/configs/datasets.yaml` β per-source download + adapter manifest; **all schemas verified against the live HF datasets-server.**
- `training/scripts/hf_download.py` β schema probe (stdlib) + split download (`datasets`), with gated-source detection.
- `training/scripts/sft_adapters.py` β per-source schema adapters; notably `detection_func_target` for PrimeVul/DiverseVul (raw `func`+`target`, which the generic normalizer silently dropped) and `vuln_fix_pair` for CrossVul.
- `training/scripts/build_sft_dataset.py` β adapt β route `<think>` (ready vs synthesize) β dedup β source caps β data card.
- `training/scripts/synthesize_think.py` β rejection-sampling `<think>` backfill (STaR-style) against a teacher endpoint, with an offline `--mock`.
- `training/scripts/normalize_sft_jsonl.py` / `split_jsonl.py` β ad-hoc converter + deterministic split.
- `training/scripts/decontaminate.py` β 13-gram + fuzzy (optional embedding) contamination gate; emits the decontam report the Stage-1 gate now requires.
- `training/scripts/build_eval_sets.py` β held-out C/C++ vuln-detection set (from reserved `test` splits) + disjoint knowledge-MCQ set.
- `training/scripts/eval_endpoint.py` β reusable scorer for any OpenAI-compatible endpoint; same command for the base baseline and every checkpoint (vuln-detection accuracy/precision/recall + MCQ accuracy).
- `training/recipes/data_preparation.md` β the full manifest-driven workflow.
- `training/tests/test_data_pipeline.py` β offline end-to-end test (**54 checks**: adapters β build β synthesize β re-build β split β eval-sets β endpoint-eval β decontaminate).
**Data findings to fold into Β§3:** (a) MegaVul/PrimeVul/DiverseVul ship `test` splits β reserve for the held-out vuln-detection eval, never train on them; (b) the two Tier-2 reasoning sets `SecCoderX/...Cold_Start` and `Leopo1d/OpenVul...` return **HTTP 401 (gated)** anonymously β they need `huggingface-cli login` + access approval before use, so the `<think>`-core relies more on synthesized traces until then. **Still missing:** decontamination against frozen eval splits, and a real teacher endpoint for synthesis.
## Appendix β Source Pointers
- Model: huggingface.co/Qwen/Qwen3.6-27B Β· Benchmark: github.com/sunblaze-ucb/cybergym, arXiv 2506.02548, HF sunblaze-ucb/cybergym
- Datasets: see Β§3 (HF ids). Methods: QLoRA (2305.14314), "LoRA Learns Less, Forgets Less" (2405.09673), GRPO/DPO refs.
|