Project Goal: CyberGym-Targeted Fine-Tune of Qwen3.6-27B
Status: Planning / pre-implementation
Owner: Adhik Joshi Β· Context: college research project (authorized security research / education)
Last updated: 2026-06-27
Repo: infosec (greenfield)
0. TL;DR
Fine-tune the open-weight Qwen3.6-27B (reasoning + agentic, Apache-2.0) to raise its score on CyberGym β UC Berkeley's agentic vulnerability-reproduction benchmark β using a curated mix of offensive + defensive cybersecurity data, with <think> reasoning traces, trained as LoRA (then merged).
Confirmed decisions (2026-06-27):
- CyberGym is the PRIMARY objective (success rate β), not knowledge benchmarks.
- Data must be offensive + defensive.
- LoRA for the trial; merge into the base checkpoint if it works (
merge_and_unload). <think>is required β the model must reason through tasks.
The hard truth this creates (read this): CyberGym is won by an agentic loop β explore a C/C++ repo over ~100 tool-calling steps and write a PoC input file that crashes the pre-patch build but not the patched one. Dedicated 32B coding-agent models still score β€2%; frontier models top out ~18β22%. Supervised fine-tuning on Q&A β even great offensive data β has a low ceiling here. To actually move CyberGym we need a staged program:
SFT (LoRA) on code-vuln + reasoning + agentic-trajectory data β DPO on secure/insecure code β RL on the CyberGym task itself (verified pass/fail reward). The RL stage is the real lever; SFT makes the model coherent and reason-y enough for RL to work.
This document plans all stages, but Stage 1 (LoRA SFT) is the trial. Expect SFT to deliver modest CyberGym gains + large gains on code-vuln/knowledge evals; budget for the RL stage to get a big CyberGym jump.
1. Objective & Success Criteria
Primary: increase CyberGym success rate over the base Qwen3.6-27B, without regressing tool-calling / agentic ability.
| Metric | Tier | Baseline | Target |
|---|---|---|---|
| CyberGym success rate (level1 sample, OpenHands + vLLM) | PRIMARY | base (expect low single digits) | > base; aim +2β5+ pts via SFT, more via RL |
| CyberGym 10-task subset β valid tool-calls emitted | Guardrail | base | no tool-call regression |
| Vulnerability-detection accuracy (C/C++) | Secondary | base | +5β15 pts |
| Cyber-knowledge MCQ (CyberMetric/SecEval/CTI-Bench) | Secondary | base | +5β15 pts |
| IFEval / MMLU-general | Guardrail | base | β₯ base β 2 pts (no forgetting) |
Definition of done (Stage 1): a merged LoRA checkpoint + eval report (CyberGym + code-vuln + guardrails + loss curves) showing net improvement and intact agentic ability, plus a reproducible pipeline in this repo. Stretch (Stage 2+): DPO + RL for a larger CyberGym jump.
2. Base Model β Qwen3.6-27B (verified)
- Real & live, released 2026-04-24, ~5.6M downloads, Apache-2.0 (clean commercial/research use, not gated).
- ~27.8B params, weights 55.6 GB bf16. Dense hybrid attention: 64 layers = 48 Gated-DeltaNet (linear) + 16 full-attention GQA layers. Hidden 5120, FFN 17408.
- Reasoning model with a ChatML
<think>channel ON by default (enable_thinking=falseinjects an empty think block). Matches your<think>requirement natively. - Multimodal (vision tower) β freeze it for text/code data.
- Untied embeddings, vocab 248,320; no BOS, EOS =
<|im_end|>. Native context 256K. - MTP head (speculative decoding) β freeze/ignore in SFT.
- Framework risk (Phase-0 gate): custom hybrid arch (
model_type~`qwen3_5) needs **very recenttransformers(β₯4.57.1)** + FLA/causal-conv1d` kernels. Older Axolotl/Unsloth/PEFT may not support it. - Fallback if blocked:
Qwen3-32B(dense, standard attention, text-only, well-supported) β also has a thinking mode. Memory math scales Γ32/27.
3. Datasets
3.A The CyberGym-aligned set (NEW β the focus going forward)
Verified via HF API (rows/schema/license). Ranked by relevance to the agentic vuln-reproduction task.
| Tier | Dataset | Rows | Schema | Why it matters for CyberGym | License |
|---|---|---|---|---|---|
| 1 Β· C/C++ vuln code | athrv/megavul-vulnerability-detection |
353,873 | instruction/input/output + cwe_ids, cve_id, severity, cvss, repo, func, is_vulnerable | Real C/C++ functions w/ CVE/CWE labels β core memory-safety understanding | β οΈ none stated |
| 1 Β· C++ vuln code | smart-code-analyzer-team/cpp-vulnerability-dataset |
69,299 | raw_snippet, security_context, data_flow_graph, ast_metadataβ¦ | C++ snippets w/ security context + data-flow β pattern recognition | β οΈ none stated |
| 1 Β· CVE patches | DetectVul/CVEFixes (+ hitoshura25/cvefixes) |
5,730 | lines, raw_lines, label, type | Pre/post-patch diff lines β mirrors CyberGym's patch differential | β οΈ none stated |
| 1 Β· SARD/Juliet C | lynchorange/SARD_Vulnerability_Detection_C |
(n/a) | C test cases | NIST-style labeled vulnerable C | β οΈ verify (SARD is public-domain) |
2 Β· Reasoning <think> |
SecCoderX/...Reasoning_Vulnerability_Detection_SFT_Cold_Start |
37,113 | instruction, output | Vuln-detection with reasoning β feeds the <think> channel |
β οΈ none stated |
2 Β· Reasoning <think> |
Leopo1d/OpenVul_Rejection_Sampling...Reasoning...SFT (+ GT companion) |
6,276 | key, prompt, completion | Reject-sampled reasoning traces β exactly our <think> method, pre-built |
Apache-2.0 |
| 3 Β· Agentic loop (top) | nvidia/Nemotron-Terminal-Corpus |
~366,000 | conversations (terminal commandβobserve) | Closest public analogue to CyberGym's run-command-observe loop | CC-BY-4.0 |
| 3 Β· Agentic loop | Kwai-Klear/SWE-smith-mini_swe_agent_plus-trajectories-66k (bash-only), SWE-bench/SWE-smith-trajectories (76k), nebius/SWE-agent-trajectories (80k) |
66β80k | messages/trajectory + patches + eval logs | Teach the multi-turn tool loop (Python-centric β use for loop mechanics, not domain) | MIT / CC-BY-4.0 |
| 3 Β· Reasoning+tools | interstellarninja/tool-use-multiturn-reasoning |
14,600 | conversations, tools (5β13 turns, interleaved <think>) |
Callβreasonβobserve rhythm with <think> |
Apache-2.0 |
| 3 Β· RL warm-start | nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1 |
50,000 | tool-call traj + pass_rate |
Verifiable-reward format for Stage-3 RL | CC-BY-4.0 |
| 3 Β· Format blueprint | SecureAI-SE/fuzz-agent-trajectories |
10 | thinkβact(tool)βobserve, offensive fuzzing | Exact offensive-agent format match β seed/template only, not trainable volume | CC-BY-4.0 |
2 Β· Reasoning <think> (cyber) |
theelderemo/pentesting-explanations |
5,900 | MCQ + explanation + <think> |
Offensive pentest reasoning with <think> |
Apache-2.0 |
| 3 Β· Tool-output skill | tegridydev/infosec-tool-output |
776 | SAST findings β next-action | Teaches the "read security-tool output" sub-skill | Apache-2.0 |
| 4 Β· DPO | CyberNative/Code_Vulnerability_Security_DPO |
4,656 | lang, vulnerability, system, question, chosen, rejected | Ready-made secure-vs-vulnerable preference pairs for the DPO stage | Apache-2.0 |
| 5 Β· CVE knowledge | AlicanKiraz0/All-CVE-Records-Training-Dataset |
297,441 | System/User/Assistant | Broad CVE breadth (defensive context/recall) | Apache-2.0 |
| 5 Β· Threat-intel (def.) | reloading0101/threat-intelligence-dataset |
~9,200 | CTI Q&A fact-checked vs MITRE/CISA/EPSS | High-provenance defensive breadth | CC-BY-4.0 |
| 6 Β· RL env | AdithyaSK/repo2rlenv-cve-patches |
(n/a) | CVE patches β RL env | CVE patches packaged as RL environments β Stage-3 fuel | Apache-2.0 |
Dedup flags: Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset is byte-identical to AlicanKiraz0/All-CVE-Records-... (same 297,441 rows / same size) β keep one. (Same pattern as the Soban/Trendyol mirror from the original list.)
License flags: several Tier-1 sets have no stated license. For a college research project this is usually defensible (academic use of public, derived-from-open-source data), but record provenance and prefer the licensed equivalents where possible (OpenVul Apache-2.0, CVEFixes upstream is CC, SARD is NIST public-domain).
Gated/non-commercial: WhitzardAgent/CyberSecurity-1M & -100B are the best-curated recent cyber corpora (exploitation chains, Sigma/YARA, CoT) but are gated behind an institutional-email approval (a personal Gmail will be rejected) and licensed academic/non-commercial β usable for a college project only if access is granted; do not assume availability.
β οΈ Structural gap (the key dataset finding): No public dataset contains end-to-end CyberGym-style C/C++ "PoC β crash β iterate" agent trajectories. The only exact-format match (
SecureAI-SE/fuzz-agent-trajectories) is 10 rows β a blueprint, not training volume. Consequence for the plan: the agentic-loop skill must come from (a) transferring loop mechanics from the generic terminal/SWE trajectory sets above, (b) synthesizing our own trajectories (run an agent harness over training-split CyberGym/ARVO tasks, keep successful runs as SFT data), and/or (c) Stage-3 RL, which learns the loop directly from the environment. This is why RL (not SFT) is the real lever for the primary metric.
3.B The original 6 (now supplementary)
Still useful as defensive breadth + instruction-following, but secondary to 3.A for the CyberGym goal:
- β
AlicanKiraz0/Fenrir-v2.1(99,870, EN, Apache-2.0),Trendyol/Cybersecurity-IT(53,202, EN, Apache-2.0),hcnote/High-Quality(270,271, ZH+EN, MIT, has offensive/exploit content + code β useful here). - β
jobby32/QA-with-negatives(17,263) βquestionβpositiveSFT only. - β
Soban1234/...(mirror of Trendyol β drop). βclouditera/security-paper(raw bilingual book text, unspecified license β exclude; optional CPT only after legal review).
3.D Platform / language breadth (NodeJS Β· Python Β· PHP Β· Apache Β· Linux Β· web)
CyberGym's core is C/C++ memory-safety bugs (OSS-Fuzz/ARVO), and Apache httpd, nginx, and the Linux kernel are themselves C β so the C/C++ sets above already cover the "systems software" targets. NodeJS / Python / PHP are different vuln classes (injection, deserialization, SSRF, path traversal) β valuable for breadth and useful if we add web targets, but secondary to C/C++ for the CyberGym score. Honest framing: prioritize the first group for CyberGym; add the second for a well-rounded offensive+defensive model.
| Group | Dataset | Rows | Schema | Coverage | License |
|---|---|---|---|---|---|
| Multi-language (best for your ask) | hitoshura25/crossvul |
9,313 | cwe_id, language, vulnerable_code, fixed_code, source | PHP/JS/Python/C/C++/Java/Ruby/Go⦠vuln+fix pairs (CrossVul) | Apache-2.0 |
| C/C++ (large, real-world) | bstee615/diversevul |
330,492 | func, target, cwe, project, commit | C/C++ across many OSS projects (incl. Apache/Linux-type) | β οΈ none |
| C/C++ (highest quality) | colin/PrimeVul |
233,941 | func, target, cwe, project, commit | Rigorously de-noised C/C++ vuln detection | MIT |
| Linux/systems (C) | mjbommar/linux-cve-dossiers |
2,556 | cve_id, package, scope, completeness⦠| Linux CVE dossiers (kernel/distro packages) | CC-BY-4.0 |
| Web / OWASP (knowledge) | shahrukh95/OWASP-and-NVD-question-answer-dataset |
22,958 | Question, Answer, split | OWASP + NVD web-vuln Q&A | β οΈ none |
| PHP (offensive/defensive) | nbuser32/PHP-Webshell-Dataset |
30,270 | instruction/input/output | PHP webshell detection/analysis | MIT |
| Web injection (NodeJS/PHP/Python apps) | PurpleAILAB/chatML_SQL_injection_dataset (+ XSS sets) |
~2,000 | conversations (ChatML) | SQLi/XSS exploitation in chat format | β οΈ none |
| Python | maddyrucos/code_vulnerability_python |
848 | func, target, project | Small Python vuln set | MIT |
Notes: nasirhm/reposvul (repo-level, multi-language β adds whole-repo context like CyberGym) and Humanlearning/CyberSecurity_OWASP-sft-dataset looked relevant but their HF viewers returned empty β verify at source before relying on them. Several web sets are tiny/no-license β use as flavor, not backbone; record provenance.
3.C Corpus shape (planning)
- Code/vuln + reasoning core (3.A tiers 1β2): ~470K rows.
- Agentic trajectories (3.A tier 3): ~80K+ (heavy per-example token count β these are long).
- DPO pairs: ~5K. CVE/defensive breadth: ~300K (deduped) + original-6 supplementary.
- After dedup + decontamination, plan for ~400β600K SFT examples, token count dominated by the long agentic trajectories. Cap any single source at β€30β40% of tokens.
4. Benchmark β CyberGym (verified) + the eval suite
CyberGym (sunblaze-ucb, arXiv 2506.02548, ICLR 2026): given a pre-patch C/C++ codebase + a vuln description, an agent must produce a PoC input file that triggers a sanitizer crash on the pre-patch build and not on the patched build (differential validation). 1,507 tasks / 188 projects; levels 0β3 (more info β easier); a 10-task curated subset ships for quick tests. Harness is agentic + Docker, uses LiteLLM β supports any OpenAI-compatible endpoint (our vLLM). Use binary-only mode (~130 GB). Baselines: GPT-5 ~22%, Sonnet-4 ~18%, open 32B coding agents β€2%.
Self-hosted eval path (confirmed): serve the merged checkpoint with vLLM (--served-model-name my-qwen, tool-calling enabled) β run OpenHands with --model openai/my-qwen --llm.base_url http://HOST:PORT/v1, LLM_API_KEY=EMPTY, --difficulty level1.
Eval suite we'll report:
- PRIMARY: CyberGym β 10-task smoke (tool-call sanity) β ~100β200 task level1 sample (real number) β full 1,507 only if budget allows.
- Secondary: held-out C/C++ vulnerability-detection accuracy (from MegaVul/CVEFixes test splits); cyber-knowledge MCQ (CyberMetric, SecEval, CTI-Bench).
- Guardrails: IFEval (instruction following), MMLU-general (forgetting).
- Decontaminate all training data against every eval set (13-gram + fuzzy + embedding); freeze test splits.
5. Training Strategy (staged)
Stage 0 β Infra & go/no-go
Confirm Qwen3.6-27B loads + LoRA-trains + serves under vLLM with tool-calling on current transformers. If blocked, fall back to Qwen3-32B. Set up storage (~236 GB CyberGym / ~130 GB binary server), W&B, vLLM.
Stage 1 β LoRA SFT (the trial; "benchmark after each epoch")
Goal: make the model fluent in C/C++ vuln reasoning + the agentic tool-use format, with <think> traces.
- Data: 3.A tiers 1β3 (code-vuln + reasoning + agentic trajectories) + DPO-source positives + defensive breadth, all normalized to ChatML with
<think>reasoning. <think>policy: keep reasoning ON. Use the pre-built reasoning datasets (SecCoderX, OpenVul) directly; for code-vuln sets lacking traces, synthesize<think>via rejection sampling (generate N reasonings with a strong model, keep those that reach the correct label/patch). Apply the format uniformly β inconsistent<think>is the #1 way to damage a Qwen reasoning model.- Method: LoRA r=64, Ξ±=128, rsLoRA on, all linear targets, dropout 0.05, freeze vision tower + MTP head.
- Hyperparameters: lr 2e-4 cosine, warmup 0.03β0.05, 2β3 epochs, effective batch ~128, seq 8192 (agentic trajectories are long), packing on, gradient checkpointing on, loss-masked on assistant tokens, paged AdamW-8bit.
- Pilot first: a QLoRA smoke-test on a subset (1Γ 80GB, ~$20β60) to validate the pipeline before the real LoRA run.
- Per-epoch eval: at each epoch checkpoint run the fast eval (vuln-detection slice + IFEval + CyberGym 10-task). Select the checkpoint by best eval + non-rising val loss, never lowest train loss.
Stage 2 β DPO (optional, cheap)
Use CyberNative/Code_Vulnerability_Security_DPO (chosen=secure, rejected=vulnerable) to sharpen secure/insecure discrimination. LoRA-DPO, lr ~5e-6, 1 epoch. Small, fast, low risk.
Stage 3 β RL on the CyberGym task β "the real lever" (detailed)
Why RL fits CyberGym perfectly. CyberGym hands us a verifiable, machine-checkable, hard-to-game reward: the submitted PoC either crashes the pre-patch build and not the patched build, or it doesn't. That is the ideal setup for RLVR (RL with Verifiable Rewards) β the same paradigm behind reasoning models (e.g. DeepSeek-R1). No reward model to train; the environment is the judge. The catch: this is agentic, long-horizon, sparse-reward RL (a success needs ~100 correct tool-calling steps, reward only at the end), which is much harder than single-turn math RLVR. So we climb a ladder from cheap to powerful.
The ladder (do them in order):
Rejection-sampling Fine-Tuning / Expert Iteration (STaR / RFT) β START HERE, highest ROI.
- Run the SFT model on training-split tasks, many high-temperature rollouts each.
- Keep only trajectories that produced a validated PoC (reward = 1).
- SFT on those winning trajectories (with their
<think>reasoning). - Repeat β as the model improves it solves more tasks β generates more training data β improves further.
- This is "RL" via filtered behavior cloning: no PPO/GRPO infrastructure, far cheaper and more stable, and it directly closes the "no public C/C++ agent-trajectory data" gap by manufacturing our own. For a college budget this alone is likely the biggest single win.
GRPO (Group Relative Policy Optimization) β the on-policy push.
- Modern policy-gradient for verifiable rewards: sample a group of rollouts per task, advantage = reward normalized within the group, no value/critic model (cheaper than PPO). Best applied after RFT has warmed the model up.
- Alternatives: PPO (needs a critic, heavier), RLOO, or offline trajectory-DPO (chosen = solved run, rejected = failed run) as a lightweight stepping stone.
Curriculum + reward shaping (to fight sparsity).
- Curriculum: start at level3 (easiest β includes the patch + crash stack trace), then level2 β level1 β level0; or start from CyberGym's 5 "solvable" tasks. Get signal flowing before tackling hard tasks.
- Shaped reward (dense, but outcome-dominant):
- locate the vulnerable function/file,
- produce any sanitizer crash on the pre-patch build (even the wrong bug),
- + full reward for the differential success (crash pre-patch, clean post-patch β the true objective),
- small format reward (valid tool-call syntax,
<think>present, valid PoC file written),
- small format reward (valid tool-call syntax,
- β small per-step/length penalty for efficiency.
- Keep the differential-success term dominant so the model can't farm the shaping terms.
Anti-reward-hacking. The differential check already blocks the cheapest exploit (triggering an unrelated crash). Also: keep eval tasks completely out of RL training; watch for the agent crashing/escaping the harness itself; cap the tool/step budget.
Environment & data.
- The CyberGym Docker harness is the RL environment β rollouts POST candidate PoCs to the validation server β binary reward. Use binary-only mode (~130 GB) for rollout speed.
- Mandatory train/eval split: RL-train on a held-in subset of the 1,507 tasks (+ external verified tasks: ARVO / OSS-Fuzz reproducible bugs,
AdithyaSK/repo2rlenv-cve-patches), evaluate on a disjoint held-out subset. Never train on eval tasks. More distinct environments β better generalization.
Frameworks.
- verl (ByteDance) β leading scalable RLVR/GRPO framework; supports multi-turn agentic rollouts + tool calling; pairs with vLLM. Top pick for the full RL run.
- OpenRLHF β PPO/GRPO/RLOO at scale. SkyRL / RAGEN / Agent-R1 / verifiers β purpose-built for multi-turn agentic RL (the hard part is rollouts that interleave model β environment). TRL
GRPOTrainerβ simplest, good for RFT / single-turn proxies. - The RFT phase needs no RL framework β just vLLM (rollouts) + the CyberGym harness (scoring) + a normal SFT trainer.
Compute reality (why this is the costly phase). Agentic RL is rollout-dominated: each episode = up to ~100 LLM calls + Docker execution, and GRPO needs a group (8β16) of episodes per task per step β thousands of ~100-step rollouts. You need a vLLM inference cluster (generation) + many parallel Docker workers (environment) + training GPUs (updates); throughput of inference/env, not the gradient step, is the bottleneck. Budget $1β5K+ and it can balloon β which is exactly why RFT (cheap, off-policy) should capture most of the gain first.
Recommended RL path (college budget): SFT (Stage 1) β RFT/Expert-Iteration on the training split (prioritize this) β if budget remains GRPO with curriculum + shaped reward via verl β evaluate on the held-out split throughout.
Realistic expectation. Frontier models sit at ~18β22% on CyberGym; dedicated 32B coding agents score β€2% without this program. A 27B model after SFT + RFT (+ optional GRPO) can credibly aim for mid-single digits β low double digits at level1 β a strong, publishable result for an open 27B in a college research setting.
LoRA β merge
After Stage 1 (and DPO), peft.merge_and_unload() folds adapters into base weights β standalone checkpoint, zero inference overhead. bf16 LoRA merges cleanly; QLoRA merge is lossy (merge into dequantized bf16, then re-quantize for serving if needed). Re-benchmark the merged model to confirm parity before continuing to RL.
6. GPU, Time & Cost
470K SFT rows incl. long agentic trajectories β plan **600Mβ1B tokens/epoch** (trajectories dominate). 8Γ H100 β $16β24/hr; 1β2Γ H100 β $2β3/GPU-hr.
| Stage | Method | Hardware | Wallclock | Compute $ |
|---|---|---|---|---|
| Pilot | QLoRA, subset | 1Γ H100 80GB | ~6β12 hr | ~$20β60 |
| 1. LoRA SFT (2β3 ep) | LoRA r=64, bf16 base | 2Γ H100 80GB (or 8Γ FSDP) | ~$200β500 | |
| 2. DPO | LoRA-DPO | 1β2Γ H100 | ~6β12 hr | ~$30β80 |
| 3. RL (CyberGym reward) | GRPO + Docker rollouts | 4β8Γ H100 + Docker host | daysβweeks | $1β5K+ (rollout-dominated) |
| CyberGym eval | vLLM + OpenHands | 1Γ H100 + Docker | ~1β2 days/sample | ~$50β200 |
VRAM: LoRA bf16 base 55.6 GB + adapters + long-seq activations β 1Γ 80GB (tight, seqβ€4096) or 2Γ 80GB (comfortable, seq 8192). Full FT (if ever) ~450β520 GB β 8Γ 80GB ZeRO-3.
Budget: Stage 1 trial ~$300β700 all-in (data prep + pilot + run + eval iterations). Add $1β5K+ only if you commit to Stage-3 RL. Storage ~370 GB.
7. Loss Curve & Monitoring
- SFT loss: starts ~1.2β1.6 (post-warmup), converges ~0.6β0.9. Below ~0.4 = memorization. Agentic-trajectory tokens may sit a bit higher.
- Healthy: sharp early drop β smooth decline β flatten; val tracks train with small stable gap.
- Overfitting: val plateaus/rises while train falls; sawtooth at epoch boundaries = memorizing β stop at 2β3 epochs.
- Track (W&B): train/val loss, grad norm, LR, tokens/s, per-epoch eval (vuln-detection + IFEval + CyberGym-10). For RL: reward curve, KL to ref, pass-rate on held-in tasks.
- Selection rule: best eval score + non-rising val loss, never lowest train loss (guards against agentic/tool-call regression).
8. Risks & Mitigations
| Risk | Likelihood | Mitigation |
|---|---|---|
| SFT barely moves CyberGym (it's agentic/RL-shaped) | High | Set expectations; Stage-1 win = vuln-detection/knowledge + agentic coherence; Stage-3 RL for the CyberGym jump |
| Framework can't train the hybrid arch | High | Phase-0 go/no-go; fall back to Qwen3-32B |
| SFT regresses tool-calling/reasoning | Medium | LoRA (low forgetting); uniform <think>; blend agentic trajectories; per-epoch guardrail evals |
| Missing-license Tier-1 data | Medium | Record provenance; prefer licensed equivalents; academic-use framing; legal review before any release |
| Benchmark contamination | Medium | 13-gram + fuzzy + embedding decontam; frozen test splits |
| RL harness cost/complexity | Medium | Gate Stage 3 on Stage-1 success; start with a small task subset |
<think> inconsistency damages reasoning |
Medium | Single uniform policy; use pre-built reasoning sets + rejection-sampled traces |
9. Open Decisions (most now resolved)
β
Primary = CyberGym Β· β
offensive+defensive Β· β
LoRAβmerge Β· β
<think> required.
Remaining:
- Commit to Stage-3 RL now (changes budget to $1β5K+), or decide after Stage 1? (Recommend: decide after Stage 1.)
- Base model: lock Qwen3.6-27B, pre-authorize Qwen3-32B fallback if Phase-0 blocks? (Recommend: yes.)
- Reasoning-trace synthesis: budget for rejection-sampling
<think>traces over code-vuln sets (needs a strong teacher model + inference $)? (Recommend: yes β it's core to the<think>goal.)
10. Phased Milestones
| Phase | Output | Gate |
|---|---|---|
| 0 Infra & go/no-go | Qwen3.6-27B trains+serves (or fallback) | Architecture trainable? |
| 1 Baseline | Base scores on CyberGym + code-vuln + guardrails | Reproducible; must finish before QLoRA/LoRA training |
| 2 Data engineering | Normalized+deduped+decontam mix w/ <think> traces + data card |
Decontam verified |
| 3 Pilot (QLoRA) | Loss curve + sane gens + 1 eval cycle | No tool-call regression |
| 4 LoRA SFT + per-epoch eval | Checkpoints + curves + eval table | Hit Stage-1 criteria |
| 5 Merge + benchmark | Merged checkpoint, CyberGym level1 sample vs base | Net improvement, no regression |
| 6 (Opt) DPO | Sharpened secure/insecure model | Cheap win |
| 7 (Opt) RL on CyberGym | v2 with real CyberGym jump | Cost/benefit justified |
11. Proposed Repo Structure
infosec/
βββ docs/PROJECT_GOAL.md
βββ data/{download,processed,decontam,think_synthesis}/ # fetch, normalize, decontaminate, reasoning-trace synthesis
βββ training/{recipes,configs,scripts,tests}/ # Vast.ai + Qwen recipe, data manifest/adapters, offline tests
βββ eval/{cybergym,vuln_detection,knowledge,judge}/
βββ serve/ # vLLM OpenAI-compatible + tool-calling
βββ reports/ # baseline, per-epoch, merged, RL
12. Current Repo Artifacts
The implementation layer lives under training/:
training/README.mdβ entrypoint and quick-start commands.training/recipes/pretraining_cybergym_baseline.mdβ mandatory base-model CyberGym benchmark before training.training/recipes/stage1_qwen36_27b_lora_sft.mdβ Stage 0 load gate, QLoRA pilot, and main LoRA SFT recipe.training/configs/*.yamlβ pilot, Stage 1, and data-mix contracts.training/scripts/vastai_*.shβ Vast.ai search/create/bootstrap scripts.training/scripts/phase0_qwen_smoke.pyβ config/tokenizer/full-load smoke gate forQwen/Qwen3.6-27B.training/scripts/run_sft.pyandtraining/scripts/merge_lora.pyβ runnable PEFT training and merge entrypoints.
Data-engineering layer (implemented + offline-tested, 2026-06-27):
training/configs/datasets.yamlβ per-source download + adapter manifest; all schemas verified against the live HF datasets-server.training/scripts/hf_download.pyβ schema probe (stdlib) + split download (datasets), with gated-source detection.training/scripts/sft_adapters.pyβ per-source schema adapters; notablydetection_func_targetfor PrimeVul/DiverseVul (rawfunc+target, which the generic normalizer silently dropped) andvuln_fix_pairfor CrossVul.training/scripts/build_sft_dataset.pyβ adapt β route<think>(ready vs synthesize) β dedup β source caps β data card.training/scripts/synthesize_think.pyβ rejection-sampling<think>backfill (STaR-style) against a teacher endpoint, with an offline--mock.training/scripts/normalize_sft_jsonl.py/split_jsonl.pyβ ad-hoc converter + deterministic split.training/scripts/decontaminate.pyβ 13-gram + fuzzy (optional embedding) contamination gate; emits the decontam report the Stage-1 gate now requires.training/scripts/build_eval_sets.pyβ held-out C/C++ vuln-detection set (from reservedtestsplits) + disjoint knowledge-MCQ set.training/scripts/eval_endpoint.pyβ reusable scorer for any OpenAI-compatible endpoint; same command for the base baseline and every checkpoint (vuln-detection accuracy/precision/recall + MCQ accuracy).training/recipes/data_preparation.mdβ the full manifest-driven workflow.training/tests/test_data_pipeline.pyβ offline end-to-end test (54 checks: adapters β build β synthesize β re-build β split β eval-sets β endpoint-eval β decontaminate).
Data findings to fold into Β§3: (a) MegaVul/PrimeVul/DiverseVul ship test splits β reserve for the held-out vuln-detection eval, never train on them; (b) the two Tier-2 reasoning sets SecCoderX/...Cold_Start and Leopo1d/OpenVul... return HTTP 401 (gated) anonymously β they need huggingface-cli login + access approval before use, so the <think>-core relies more on synthesized traces until then. Still missing: decontamination against frozen eval splits, and a real teacher endpoint for synthesis.
Appendix β Source Pointers
- Model: huggingface.co/Qwen/Qwen3.6-27B Β· Benchmark: github.com/sunblaze-ucb/cybergym, arXiv 2506.02548, HF sunblaze-ucb/cybergym
- Datasets: see Β§3 (HF ids). Methods: QLoRA (2305.14314), "LoRA Learns Less, Forgets Less" (2405.09673), GRPO/DPO refs.