--- library_name: pytorch tags: - causal-lm - reinforcement-learning - human-feedback - testgeniy - custom-architecture language: - en - ru base_model: Asilarkness/testgeniy --- # TestGeniy RL v2 - HelpSteer2/OASST real-RL step 20 Experimental post-training checkpoint for the custom TestGeniy architecture. This is a separate candidate folder; it does not replace the primary dialogue_sft_v6 checkpoint. ## Training - Base: testgeniy_v6_clean/checkpoints/dialogue_sft.pt - Algorithm: online group-relative policy gradient with a KL anchor to v6 - Human reward: reward head trained on human HelpSteer2 ratings and validated on OASST preferences - Verifier rewards: real GSM8K, MATH, ARC-Challenge and FOLIO labels - Prompts: real-only; no synthetic prompt dataset - MTP: disabled - Updates: 20 - Group size: 4 - Learning rate: 5e-8 ## Evaluation Same fixed 100-example suite and evaluator used for the v6 comparison: | Checkpoint | GSM8K | MATH-500 | ARC | FOLIO | Composite | |---|---:|---:|---:|---:|---:| | v6 Dialogue SFT | 24 | 7 | 26 | 29 | 21.50 | | RL v2 step 20 | 24 | 8 | 26 | 29 | 21.75 | Human holdouts: OASST validation 55% for both; HelpSteer2-derived 500-pair holdout 41.6% for both. This checkpoint is an experimental candidate. Validate before replacing the primary model.