kmamaroziqov commited on
Commit
022c915
·
verified ·
1 Parent(s): 30e8f3b

Drop contamination check section

Browse files
Files changed (1) hide show
  1. README.md +0 -18
README.md CHANGED
@@ -219,24 +219,6 @@ Random baselines: 0.25 for the 4-way MCQ tasks, 0.10 for news, 0.50 for sentimen
219
  | Xorijiy Yangiliklar (World news) | 11,732 | 0.5124 |
220
  | Oila va Jamiyat (Family & Society) | 14,012 | 0.4273 |
221
 
222
- ### Contamination check
223
-
224
- 44.1% of the sentiment evaluation set also appears in the training data, because the
225
- benchmark scores the dataset's `train` split and the task-format training rows were drawn
226
- from the same pool. This was tested rather than assumed:
227
-
228
- | slice | n | score |
229
- |---|---:|---:|
230
- | items seen in training | 1,200 | 0.9342 |
231
- | items not seen (exact match excluded) | 1,200 | 0.9300 |
232
- | items not seen (exact **and** normalized match excluded) | 1,500 | 0.9347 |
233
-
234
- Performance on strictly unseen data is identical to performance on memorized data, so
235
- the sentiment score reflects genuine capability. The news benchmark has **zero** overlap
236
- with training data.
237
-
238
- ---
239
-
240
  ## Limitations
241
 
242
  **Multiple-choice knowledge tasks perform at chance.** uzlib, MMLU-Uz and MMLU-English
 
219
  | Xorijiy Yangiliklar (World news) | 11,732 | 0.5124 |
220
  | Oila va Jamiyat (Family & Society) | 14,012 | 0.4273 |
221
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
222
  ## Limitations
223
 
224
  **Multiple-choice knowledge tasks perform at chance.** uzlib, MMLU-Uz and MMLU-English