AetherPrior's picture
Update best PRM: combined allpairs K=3 (val acc 0.968)
a8ab1e3 verified
|
Raw
History Blame Contribute Delete
1.13 kB
metadata
license: apache-2.0
base_model: Qwen/Qwen3-8B
tags:
  - reward-model
  - prm
  - code-security
  - qwen3
library_name: transformers
pipeline_tag: text-classification

Qwen3-8B Implementation-level PRM (exec-filtered, thinking)

Scalar Bradley–Terry preference model over implementation guidelines (<high_level> + <implementation>) for secure coding.

Training

  • Base: Qwen/Qwen3-8BAutoModelForSequenceClassification (num_labels=1)
  • Prefs: evolved (t+) / preferred (t) / degraded (t−) guidelines, exec-filtered with thinking enabled (keep t+ iff func∧sec; t− iff ¬sec)
  • Independent pairwise BT: t≻t−, t+≻t−, t+≻t
  • Data: combined 5cwe + other CWEs, K=3 paired attempts (attempt-0 + 2 extras), preferences_synth_triple_exec_combined_allpairs_k3 (6933 pairs; 6192/741)
  • Val accuracy (pairwise): 0.968 (741 pairs)

Scoring

Chat messages: user = coding task, assistant = guideline trace. Serve with vLLM pooling/classify and POST /classify.

Citation / project

Internal: prm_secode_eval/recode-style prm_impl_bt_triple_exec_combined_allpairs_k3 best.