AetherPrior's picture
Update best PRM: combined allpairs K=3 (val acc 0.968)
a8ab1e3 verified
|
Raw
History Blame Contribute Delete
1.13 kB
---
license: apache-2.0
base_model: Qwen/Qwen3-8B
tags:
- reward-model
- prm
- code-security
- qwen3
library_name: transformers
pipeline_tag: text-classification
---
# Qwen3-8B Implementation-level PRM (exec-filtered, thinking)
Scalar Bradley–Terry preference model over **implementation guidelines**
(`<high_level>` + `<implementation>`) for secure coding.
## Training
- Base: `Qwen/Qwen3-8B``AutoModelForSequenceClassification` (`num_labels=1`)
- Prefs: evolved (t+) / preferred (t) / degraded (t−) guidelines, exec-filtered
with thinking enabled (keep t+ iff func∧sec; t− iff ¬sec)
- Independent pairwise BT: t≻t−, t+≻t−, t+≻t
- Data: combined 5cwe + other CWEs, K=3 paired attempts (attempt-0 + 2 extras),
`preferences_synth_triple_exec_combined_allpairs_k3` (6933 pairs; 6192/741)
- Val accuracy (pairwise): **0.968** (741 pairs)
## Scoring
Chat messages: user = coding task, assistant = guideline trace.
Serve with vLLM pooling/classify and `POST /classify`.
## Citation / project
Internal: `prm_secode_eval/recode-style` `prm_impl_bt_triple_exec_combined_allpairs_k3` best.