Asilarkness's picture
Add RL v2 step20 model card
22e01b2 verified
|
Raw
History Blame Contribute Delete
1.3 kB
metadata
library_name: pytorch
tags:
  - causal-lm
  - reinforcement-learning
  - human-feedback
  - testgeniy
  - custom-architecture
language:
  - en
  - ru
base_model: Asilarkness/testgeniy

TestGeniy RL v2 - HelpSteer2/OASST real-RL step 20

Experimental post-training checkpoint for the custom TestGeniy architecture. This is a separate candidate folder; it does not replace the primary dialogue_sft_v6 checkpoint.

Training

  • Base: testgeniy_v6_clean/checkpoints/dialogue_sft.pt
  • Algorithm: online group-relative policy gradient with a KL anchor to v6
  • Human reward: reward head trained on human HelpSteer2 ratings and validated on OASST preferences
  • Verifier rewards: real GSM8K, MATH, ARC-Challenge and FOLIO labels
  • Prompts: real-only; no synthetic prompt dataset
  • MTP: disabled
  • Updates: 20
  • Group size: 4
  • Learning rate: 5e-8

Evaluation

Same fixed 100-example suite and evaluator used for the v6 comparison:

Checkpoint GSM8K MATH-500 ARC FOLIO Composite
v6 Dialogue SFT 24 7 26 29 21.50
RL v2 step 20 24 8 26 29 21.75

Human holdouts: OASST validation 55% for both; HelpSteer2-derived 500-pair holdout 41.6% for both.

This checkpoint is an experimental candidate. Validate before replacing the primary model.