wetsoledrysoul commited on
Commit
dc1c7a1
·
verified ·
1 Parent(s): 3f0cda2

Update model/dataset card: arXiv link, updated abstract, unified documentation

Browse files
Files changed (1) hide show
  1. README.md +61 -176
README.md CHANGED
@@ -1,199 +1,84 @@
1
  ---
2
  library_name: transformers
3
- tags: []
 
 
 
 
 
 
 
 
 
 
4
  ---
5
 
6
- # Model Card for Model ID
 
 
 
 
7
 
8
- <!-- Provide a quick summary of what the model is/does. -->
9
 
 
10
 
 
11
 
12
- ## Model Details
13
 
14
- ### Model Description
15
 
16
- <!-- Provide a longer summary of what this model is. -->
17
 
18
- This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
19
 
20
- - **Developed by:** [More Information Needed]
21
- - **Funded by [optional]:** [More Information Needed]
22
- - **Shared by [optional]:** [More Information Needed]
23
- - **Model type:** [More Information Needed]
24
- - **Language(s) (NLP):** [More Information Needed]
25
- - **License:** [More Information Needed]
26
- - **Finetuned from model [optional]:** [More Information Needed]
27
 
28
- ### Model Sources [optional]
29
 
30
- <!-- Provide the basic links for the model. -->
31
 
32
- - **Repository:** [More Information Needed]
33
- - **Paper [optional]:** [More Information Needed]
34
- - **Demo [optional]:** [More Information Needed]
35
 
36
- ## Uses
 
 
 
 
 
 
 
 
37
 
38
- <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
39
 
40
- ### Direct Use
 
 
41
 
42
- <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
43
 
44
- [More Information Needed]
45
 
46
- ### Downstream Use [optional]
 
 
 
 
 
 
 
 
 
 
47
 
48
- <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
49
 
50
- [More Information Needed]
51
-
52
- ### Out-of-Scope Use
53
-
54
- <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
55
-
56
- [More Information Needed]
57
-
58
- ## Bias, Risks, and Limitations
59
-
60
- <!-- This section is meant to convey both technical and sociotechnical limitations. -->
61
-
62
- [More Information Needed]
63
-
64
- ### Recommendations
65
-
66
- <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
67
-
68
- Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
69
-
70
- ## How to Get Started with the Model
71
-
72
- Use the code below to get started with the model.
73
-
74
- [More Information Needed]
75
-
76
- ## Training Details
77
-
78
- ### Training Data
79
-
80
- <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
81
-
82
- [More Information Needed]
83
-
84
- ### Training Procedure
85
-
86
- <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
87
-
88
- #### Preprocessing [optional]
89
-
90
- [More Information Needed]
91
-
92
-
93
- #### Training Hyperparameters
94
-
95
- - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
96
-
97
- #### Speeds, Sizes, Times [optional]
98
-
99
- <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
100
-
101
- [More Information Needed]
102
-
103
- ## Evaluation
104
-
105
- <!-- This section describes the evaluation protocols and provides the results. -->
106
-
107
- ### Testing Data, Factors & Metrics
108
-
109
- #### Testing Data
110
-
111
- <!-- This should link to a Dataset Card if possible. -->
112
-
113
- [More Information Needed]
114
-
115
- #### Factors
116
-
117
- <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
118
-
119
- [More Information Needed]
120
-
121
- #### Metrics
122
-
123
- <!-- These are the evaluation metrics being used, ideally with a description of why. -->
124
-
125
- [More Information Needed]
126
-
127
- ### Results
128
-
129
- [More Information Needed]
130
-
131
- #### Summary
132
-
133
-
134
-
135
- ## Model Examination [optional]
136
-
137
- <!-- Relevant interpretability work for the model goes here -->
138
-
139
- [More Information Needed]
140
-
141
- ## Environmental Impact
142
-
143
- <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
144
-
145
- Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
146
-
147
- - **Hardware Type:** [More Information Needed]
148
- - **Hours used:** [More Information Needed]
149
- - **Cloud Provider:** [More Information Needed]
150
- - **Compute Region:** [More Information Needed]
151
- - **Carbon Emitted:** [More Information Needed]
152
-
153
- ## Technical Specifications [optional]
154
-
155
- ### Model Architecture and Objective
156
-
157
- [More Information Needed]
158
-
159
- ### Compute Infrastructure
160
-
161
- [More Information Needed]
162
-
163
- #### Hardware
164
-
165
- [More Information Needed]
166
-
167
- #### Software
168
-
169
- [More Information Needed]
170
-
171
- ## Citation [optional]
172
-
173
- <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
174
-
175
- **BibTeX:**
176
-
177
- [More Information Needed]
178
-
179
- **APA:**
180
-
181
- [More Information Needed]
182
-
183
- ## Glossary [optional]
184
-
185
- <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
186
-
187
- [More Information Needed]
188
-
189
- ## More Information [optional]
190
-
191
- [More Information Needed]
192
-
193
- ## Model Card Authors [optional]
194
-
195
- [More Information Needed]
196
-
197
- ## Model Card Contact
198
-
199
- [More Information Needed]
 
1
  ---
2
  library_name: transformers
3
+ license: cc-by-nc-sa-4.0
4
+ pipeline_tag: text-generation
5
+ base_model:
6
+ - Qwen/Qwen2.5-7B-Instruct
7
+ datasets:
8
+ - Aletheia-Bench/Aletheia-Train
9
+ tags:
10
+ - code
11
+ - code-verification
12
+ - rlvr
13
+ - grpo
14
  ---
15
 
16
+ <font size=3><div align='center'>
17
+ [[**📖 Paper**](https://arxiv.org/pdf/2601.12186)]
18
+ [[**💻 Code**](https://github.com/insait-institute/aletheia)]
19
+ [[**🤗 Models & Datasets**](https://huggingface.co/Aletheia-Bench)]
20
+ </div></font>
21
 
22
+ # Aletheia: What Makes RLVR For Code Verifiers Tick?
23
 
24
+ Multi-domain thinking verifiers trained via Reinforcement Learning with Verifiable Rewards (RLVR) are a cornerstone of modern post-training. However, their adoption in code generation has lagged behind that of execution feedback due to the prohibitive costs of the full RLVR pipeline. In this work, we ablate three primary choices along the performance-cost trade-off in RLVR: intermediate thinking traces, learning from negative samples, and on-policy training. We introduce **Aletheia**, a controlled, execution-grounded testbed to facilitate a contamination-free analysis of code verifier training recipes across disparate model sizes and covariate shifts across two common verifier application scenarios. Our analysis reveals that the optimal training recipe is scale-dependent: on-policy learning is the primary performance driver for small verifiers, whereas the thinking budget becomes the most vital factor at larger scales. While leveraging negative samples has a consistent impact on top-1 selection accuracy across sizes, their contribution to ranking reconstruction increases monotonically with scale and plays a key role in stabilizing training at large sizes. Our Pareto optimality analysis demonstrates that eliminating on-policy training at larger model scales yields a verifier that performs comparably to the full RLVR recipe. Furthermore, we find that eschewing thinking traces serves as a compute-efficient strategy at lower budgets, offering a strong trade-off between training cost and verifier accuracy. Ultimately, our work provides the empirical foundation necessary to efficiently deploy robust code verifiers, thereby enabling their wider adoption in post-training pipelines for large code generation models.
25
 
26
+ ## 🤖 This Model
27
 
28
+ **`GRPO-Instruct-7B`** is a **GRPO-Instruct** verifier: trained with RLVR (on-policy, with negative samples) but **without intermediate thinking traces**, directly emitting a verdict. It is fine-tuned from [`Qwen/Qwen2.5-7B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) (7B scale) on [`Aletheia-Bench/Aletheia-Train`](https://huggingface.co/datasets/Aletheia-Bench/Aletheia-Train).
29
 
30
+ Given a competitive-programming problem and a set of candidate solutions, the verifier judges and ranks the candidates. See the [GitHub repository](https://github.com/insait-institute/aletheia) for the exact prompt format, plus training and evaluation scripts.
31
 
32
+ ## 📦 Model Zoo
33
 
34
+ Fine-tuned code verifiers at 1.5B, 7B, and 14B scales using several algorithms:
35
 
36
+ | Algorithm | Thinking | Negatives | Online | Description |
37
+ | :--- | :---: | :---: | :---: | :--- |
38
+ | [**GRPO-Think**](https://huggingface.co/Aletheia-Bench/GRPO-Think-7B-16k) | | ✅ | ✅ | Standard GRPO-style approach to training verifiers. |
39
+ | [**GRPO-Instruct**](https://huggingface.co/Aletheia-Bench/GRPO-Instruct-7B) | | ✅ | ✅ | RLVR training without intermediate thinking traces. |
40
+ | [**RAFT**](https://huggingface.co/Aletheia-Bench/RAFT-7B) | | ❌ | ✅ | On-policy rejection sampling fine-tuning using only positive reasoning samples. |
41
+ | [**DPO-Think**](https://huggingface.co/Aletheia-Bench/DPO-Think-7B) | | ✅ | ❌ | Offline preference optimization using pre-collected thinking traces. |
42
+ | [**BatchOnline-GRPO**](https://huggingface.co/Aletheia-Bench/BatchOnline-GRPO-7B) | | | ⚠️ | Semi-online training where the generator policy is synced every 4 steps. |
43
 
44
+ The `-4k` / `-8k` / `-16k` suffix on GRPO-Think checkpoints denotes the reasoning-token budget (maximum completion length) used during training.
45
 
46
+ ## 🎁 Datasets
47
 
48
+ The Aletheia dataset collection includes:
 
 
49
 
50
+ * [**Aletheia-Train**](https://huggingface.co/datasets/Aletheia-Bench/Aletheia-Train): 50,000 training instances, each pairing a competitive-programming problem with 2–5 candidate solutions — exactly one of which is correct (execution-verified) — generated by a pool of weak and strong policy models across Python, C++, and Java.
51
+ * [**Aletheia-Train-Questions**](https://huggingface.co/datasets/Aletheia-Bench/Aletheia-Train-Questions): The 3,574 unique problem statements underlying Aletheia-Train, useful for on-policy sampling.
52
+ * [**Aletheia-DPO**](https://huggingface.co/datasets/Aletheia-Bench/Aletheia-DPO): A companion dataset to Aletheia-Train containing "chosen" and "rejected" verification responses for each instance. The chosen response identifies the correct code snippet, while the rejected response does not.
53
+ * [**Aletheia-Mixed**](https://huggingface.co/datasets/Aletheia-Bench/Aletheia-Mixed): A variant of Aletheia-Train in which ~25% of instances carry adversarial modifications that exploit common LLM biases (authority references, self-declared correctness, misleading comments, etc.).
54
+ * [**Aletheia-Mixed-DPO**](https://huggingface.co/datasets/Aletheia-Bench/Aletheia-Mixed-DPO): The DPO-style companion to Aletheia-Mixed.
55
+ * [**Aletheia-Heldout**](https://huggingface.co/datasets/Aletheia-Bench/Aletheia-Heldout): A completely in-distribution test set.
56
+ * [**Aletheia-Strong**](https://huggingface.co/datasets/Aletheia-Bench/Aletheia-Strong): An OOD test set where the candidates are generated by stronger models.
57
+ * [**Aletheia-Hard**](https://huggingface.co/datasets/Aletheia-Bench/Aletheia-Hard): An OOD test set where the comparison between candidates is more difficult.
58
+ * [**Aletheia-Adv**](https://huggingface.co/datasets/Aletheia-Bench/Aletheia-Adv): An OOD test set where the candidates are adversarially modified to exploit common LLM biases.
59
 
60
+ ## 💡 Intended Uses
61
 
62
+ * **RLHF / RLAIF**: plug-and-play reward function for code generation policy optimization.
63
+ * **Automated evaluation**: LLM-as-a-judge for a variety of code-related tasks.
64
+ * **Research**: study the effects of thinking traces, on-policy learning, and negative samples in training successful code verifiers.
65
 
66
+ ## 📚 Citation
67
 
68
+ If you find this work useful, please cite our paper:
69
 
70
+ ```bibtex
71
+ @misc{venkatkrishna2026aletheiamakesrlvrcode,
72
+ title={Aletheia: What Makes RLVR For Code Verifiers Tick?},
73
+ author={Vatsal Venkatkrishna and Indraneil Paul and Iryna Gurevych},
74
+ year={2026},
75
+ eprint={2601.12186},
76
+ archivePrefix={arXiv},
77
+ primaryClass={cs.SE},
78
+ url={https://arxiv.org/abs/2601.12186},
79
+ }
80
+ ```
81
 
82
+ ## 📄 License
83
 
84
+ This work is licensed under [CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/).