Trained Verifier Models
Surrogate code verifiers across three model sizes trained using multiple different algorithms as described in the Aletheia paper
Text Generation • 2B • Updated • 72Note Our flagship verifiers, trained using a GRPO-style algorithm with a 16k generation limit
Aletheia-Bench/GRPO-Think-7B-16k
Text Generation • 8B • Updated • 78Note Our flagship verifiers, trained using a GRPO-style algorithm with a 16k generation limit
Aletheia-Bench/GRPO-Think-14B-16k
Text Generation • 15B • Updated • 76Note Our flagship verifiers, trained using a GRPO-style algorithm with a 16k generation limit
Aletheia-Bench/GRPO-Think-1.5B-4k
Text Generation • 2B • Updated • 60Note A variant of the flagship verifiers trained with a 4k generation limit
Aletheia-Bench/GRPO-Think-7B-4k
Text Generation • 8B • Updated • 73Note A variant of the flagship verifiers trained with a 4k generation limit
Aletheia-Bench/GRPO-Think-14B-4k
Text Generation • 15B • Updated • 69Note A variant of the flagship verifiers trained with a 4k generation limit
Aletheia-Bench/GRPO-Think-1.5B-8k
Text Generation • 2B • Updated • 60Note A variant of our flagship verifiers, trained with a 8k generation limit
Aletheia-Bench/GRPO-Think-7B-8k
Text Generation • 8B • Updated • 69Note A variant of our flagship verifiers, trained with a 8k generation limit
Aletheia-Bench/GRPO-Think-14B-8k
Text Generation • 15B • Updated • 78 • 1Note A variant of our flagship verifiers, trained with a 8k generation limit
Aletheia-Bench/GRPO-Instruct-1.5B
Text Generation • 2B • Updated • 87Note A GRPO training ablation that does not generate thinking traces
Aletheia-Bench/GRPO-Instruct-7B
Text Generation • 8B • Updated • 86Note A GRPO training ablation that does not generate thinking traces
Aletheia-Bench/GRPO-Instruct-14B
Text Generation • 15B • Updated • 95Note A GRPO training ablation that does not generate thinking traces
Aletheia-Bench/DPO-Think-1.5B
Text Generation • 2B • Updated • 78Note A verifier trained completely offline using DPO
Aletheia-Bench/DPO-Think-7B
Text Generation • 8B • Updated • 74Note A verifier trained completely offline using DPO
Aletheia-Bench/DPO-Think-14B
Text Generation • 15B • Updated • 158 • 2Note A verifier trained completely offline using DPO
Aletheia-Bench/BatchOnline-GRPO-1.5B
Text Generation • 2B • Updated • 72Note A variant of our flagship verifiers, where the generation policy is synced every 4 gradient updates
Aletheia-Bench/BatchOnline-GRPO-7B
Text Generation • 8B • Updated • 74 • 1Note A variant of our flagship verifiers, where the generation policy is synced every 4 gradient updates
Aletheia-Bench/BatchOnline-GRPO-14B
Text Generation • 15B • Updated • 76 • 1Note A variant of our flagship verifiers, where the generation policy is synced every 4 gradient updates
Aletheia-Bench/RAFT-1.5B
Text Generation • 2B • Updated • 71Note A verifier trained using on-policy rejection sampling on only positive samples
Aletheia-Bench/RAFT-7B
Text Generation • 8B • Updated • 71Note A verifier trained using on-policy rejection sampling on only positive samples
Aletheia-Bench/RAFT-14B
Text Generation • 15B • Updated • 73Note A verifier trained using on-policy rejection sampling on only positive samples
-
Aletheia: What Makes RLVR For Code Verifiers Tick?
Paper • 2601.12186 • Published • 1