File size: 3,324 Bytes
3e4433b
6ae3ce9
 
3e4433b
6ae3ce9
3e4433b
6ae3ce9
 
 
 
 
 
 
 
 
 
 
 
 
 
3e4433b
6ae3ce9
 
 
 
 
 
 
 
 
 
 
 
3e4433b
 
6ae3ce9
3e4433b
6ae3ce9
 
 
 
3e4433b
6ae3ce9
 
 
3e4433b
6ae3ce9
3e4433b
6ae3ce9
 
 
 
 
 
3e4433b
6ae3ce9
 
3e4433b
6ae3ce9
3e4433b
6ae3ce9
 
 
 
 
3e4433b
6ae3ce9
3e4433b
6ae3ce9
 
 
 
 
 
3e4433b
6ae3ce9
3e4433b
6ae3ce9
 
 
 
 
 
3e4433b
6ae3ce9
3e4433b
6ae3ce9
 
 
3e4433b
6ae3ce9
3e4433b
6ae3ce9
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
---

license: other
pipeline_tag: text-classification
library_name: transformers
base_model: microsoft/deberta-v3-large
tags:
  - text-classification
  - process-reward-model
  - reasoning-verification
  - step-verification
  - crp
  - context-relay-protocol
datasets:
  - trl-lib/prm800k
  - peiyi9979/Math-Shepherd
  - UW-Madison-Lee-Lab/MMLU-Pro-CoT-Train-Labeled
  - RLHFlow/Mistral-PRM-Data
metrics:
  - accuracy
  - auc
model-index:
  - name: crp-prm-deberta-v1
    results:
      - task:
          type: text-classification
          name: Reasoning-step validity (VALID/INVALID)
        dataset:
          type: trl-lib/prm800k
          name: prm800k held-out test steps
        metrics:
          - type: auc
            value: 0.793
            name: ROC AUC (held-out, independent harness)
---


# CRP PRM DeBERTa β€” process reward model (step verifier)

Part of the [Context Relay Protocol (CRP)](https://crprotocol.io) ML-first
governance layer. Scores whether a reasoning step is entailed by / consistent
with its premises β€” `VALID` or `INVALID` β€” inside the CRP Verification Relay
(`crp/vr/prm.py`, SPEC-049). Input format:

```

premises: <problem and prior steps> [SEP] step: <current step>

```

## Verified results (independent harness, 2026-08-01)

| Metric | Value |
|---|---|
| ROC AUC (389 held-out prm800k steps) | **0.793** |
| Curated reasoning cases (mixed domains) | **8/10** |
| VALID recall (threshold β‰₯ 0.15) | **1.000** β€” never false-flags good steps |
| Best operating point (this slice) | t=0.1 β†’ INVALID recall 0.577 / VALID recall 0.890 |

Trainer-side eval (on a training-matched mix): eval_loss 0.133, INVALID F1

0.977. See *Calibration* below before using thresholds.



## Intended use β€” advisory scorer



Wire as an **advisory** step-quality score: it never false-halts good steps

(VALID recall 1.0) and catches blatant bad ones, but subtle math-domain

errors score low (score compression on out-of-training distribution). Hard

INVALID gating should stay with symbolic verifiers / human checkpoints β€”

exactly how `crp/vr/prm.py` consumes it.



## Calibration (important)



The config's exported `prm_threshold` (0.675) was calibrated on a mix that
includes the same synthetic patterns as training, and **does not transfer to

real distributions** β€” on pure prm800k test steps, P(INVALID) compresses
below 0.5 (mean 0.098 on true INVALID vs 0.053 on true VALID; ranking is
real, absolute calibration is distribution-dependent). Recalibrate on your
own traffic before thresholding, or use argmax (0.5).

## Training data

~400k step-level examples (stratified ~50/50): `trl-lib/prm800k` (human
labels), `peiyi9979/Math-Shepherd` (GPT-4 step labels),
`UW-Madison-Lee-Lab/MMLU-Pro-CoT-Train-Labeled` (14 non-math domains),
`RLHFlow/Mistral-PRM-Data`, plus 20k synthetic agentic (devops / compliance /
code) reasoning examples. DeBERTa-v3-large, bf16, 2 epochs, lr 1e-5,
RunPod RTX 4090.

## Limitations

Not a substitute for formal/symbolic verification. Weakest on long, subtle
multi-step math derivations (premises truncation at 384 tokens). English
only. Scores are distribution-sensitive β€” see Calibration.

## License

Elastic License 2.0 β€” see the CRP repository for details.