crp-prm-deberta-v1 / README.md
ConstantinosVidiniotis's picture
Complete model card: verified metrics, training data, usage, limitations
6ae3ce9 verified
|
Raw
History Blame Contribute Delete
3.32 kB
metadata
license: other
pipeline_tag: text-classification
library_name: transformers
base_model: microsoft/deberta-v3-large
tags:
  - text-classification
  - process-reward-model
  - reasoning-verification
  - step-verification
  - crp
  - context-relay-protocol
datasets:
  - trl-lib/prm800k
  - peiyi9979/Math-Shepherd
  - UW-Madison-Lee-Lab/MMLU-Pro-CoT-Train-Labeled
  - RLHFlow/Mistral-PRM-Data
metrics:
  - accuracy
  - auc
model-index:
  - name: crp-prm-deberta-v1
    results:
      - task:
          type: text-classification
          name: Reasoning-step validity (VALID/INVALID)
        dataset:
          type: trl-lib/prm800k
          name: prm800k held-out test steps
        metrics:
          - type: auc
            value: 0.793
            name: ROC AUC (held-out, independent harness)

CRP PRM DeBERTa β€” process reward model (step verifier)

Part of the Context Relay Protocol (CRP) ML-first governance layer. Scores whether a reasoning step is entailed by / consistent with its premises β€” VALID or INVALID β€” inside the CRP Verification Relay (crp/vr/prm.py, SPEC-049). Input format:

premises: <problem and prior steps> [SEP] step: <current step>

Verified results (independent harness, 2026-08-01)

Metric Value
ROC AUC (389 held-out prm800k steps) 0.793
Curated reasoning cases (mixed domains) 8/10
VALID recall (threshold β‰₯ 0.15) 1.000 β€” never false-flags good steps
Best operating point (this slice) t=0.1 β†’ INVALID recall 0.577 / VALID recall 0.890

Trainer-side eval (on a training-matched mix): eval_loss 0.133, INVALID F1 0.977. See Calibration below before using thresholds.

Intended use β€” advisory scorer

Wire as an advisory step-quality score: it never false-halts good steps (VALID recall 1.0) and catches blatant bad ones, but subtle math-domain errors score low (score compression on out-of-training distribution). Hard INVALID gating should stay with symbolic verifiers / human checkpoints β€” exactly how crp/vr/prm.py consumes it.

Calibration (important)

The config's exported prm_threshold (0.675) was calibrated on a mix that includes the same synthetic patterns as training, and does not transfer to real distributions β€” on pure prm800k test steps, P(INVALID) compresses below 0.5 (mean 0.098 on true INVALID vs 0.053 on true VALID; ranking is real, absolute calibration is distribution-dependent). Recalibrate on your own traffic before thresholding, or use argmax (0.5).

Training data

~400k step-level examples (stratified ~50/50): trl-lib/prm800k (human labels), peiyi9979/Math-Shepherd (GPT-4 step labels), UW-Madison-Lee-Lab/MMLU-Pro-CoT-Train-Labeled (14 non-math domains), RLHFlow/Mistral-PRM-Data, plus 20k synthetic agentic (devops / compliance / code) reasoning examples. DeBERTa-v3-large, bf16, 2 epochs, lr 1e-5, RunPod RTX 4090.

Limitations

Not a substitute for formal/symbolic verification. Weakest on long, subtle multi-step math derivations (premises truncation at 384 tokens). English only. Scores are distribution-sensitive β€” see Calibration.

License

Elastic License 2.0 β€” see the CRP repository for details.