keepsloading commited on
Commit
6a26953
·
verified ·
1 Parent(s): 1ade3b4

Upload folder using huggingface_hub

Browse files
Files changed (27) hide show
  1. outputs/eval/logs/token_t045/babilong.log +280 -280
  2. outputs/eval/logs/token_t045/ruler_a.log +259 -265
  3. outputs/eval/logs/token_t045/ruler_b.log +207 -210
  4. outputs/eval/logs/vanilla/babilong.log +179 -179
  5. outputs/eval/logs/vanilla/ruler_a.log +267 -260
  6. outputs/eval/logs/vanilla/ruler_b.log +213 -210
  7. outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/results_2026-07-18T15-23-38.141056.json +505 -0
  8. outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa1_2026-07-18T15-23-38.141056.jsonl +3 -0
  9. outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa2_2026-07-18T15-23-38.141056.jsonl +3 -0
  10. outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa3_2026-07-18T15-23-38.141056.jsonl +3 -0
  11. outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa4_2026-07-18T15-23-38.141056.jsonl +3 -0
  12. outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa5_2026-07-18T15-23-38.141056.jsonl +3 -0
  13. outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/results_2026-07-18T15-05-41.414854.json +821 -0
  14. outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multikey_2_2026-07-18T15-05-41.414854.jsonl +3 -0
  15. outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multiquery_2026-07-18T15-05-41.414854.jsonl +3 -0
  16. outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_single_1_2026-07-18T15-05-41.414854.jsonl +3 -0
  17. outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_single_3_2026-07-18T15-05-41.414854.jsonl +3 -0
  18. outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_fwe_2026-07-18T15-05-41.414854.jsonl +3 -0
  19. outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_qa_hotpot_2026-07-18T15-05-41.414854.jsonl +3 -0
  20. outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_vt_2026-07-18T15-05-41.414854.jsonl +3 -0
  21. outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/results_2026-07-18T15-17-52.616068.json +714 -0
  22. outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multikey_1_2026-07-18T15-17-52.616068.jsonl +3 -0
  23. outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multikey_3_2026-07-18T15-17-52.616068.jsonl +3 -0
  24. outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multivalue_2026-07-18T15-17-52.616068.jsonl +3 -0
  25. outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_single_2_2026-07-18T15-17-52.616068.jsonl +3 -0
  26. outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_cwe_2026-07-18T15-17-52.616068.jsonl +3 -0
  27. outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_qa_squad_2026-07-18T15-17-52.616068.jsonl +3 -0
outputs/eval/logs/token_t045/babilong.log CHANGED
@@ -1,301 +1,301 @@
1
- 2026-07-18:12:26:33 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
- 2026-07-18:12:26:37 INFO [_cli.run:376] Selected Tasks: ['babilong_longctx']
3
- 2026-07-18:12:26:38 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
- 2026-07-18:12:26:38 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
- 2026-07-18:12:26:41 INFO [models.huggingface:161] Using device 'cuda:0'
6
- 2026-07-18:12:26:42 INFO [models.huggingface:548] Model type cannot be determined. Using default model type 'causal'
7
- 2026-07-18:12:26:42 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
8
- 2026-07-18:12:26:44 WARNING [api.task:856] babilong_qa5: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
9
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
10
- 2026-07-18:12:26:44 WARNING [api.task:856] babilong_qa4: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
11
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
12
- 2026-07-18:12:26:44 WARNING [api.task:856] babilong_qa3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
13
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
14
- 2026-07-18:12:26:44 WARNING [api.task:856] babilong_qa2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
15
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
16
- 2026-07-18:12:26:45 WARNING [api.task:856] babilong_qa1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
17
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
18
- 2026-07-18:12:26:45 INFO [tasks:700] Selected tasks:
19
- 2026-07-18:12:26:45 INFO [tasks:703] Group: babilong_longctx
20
- 2026-07-18:12:26:45 INFO [tasks:726] ConfigurableGroup(group=babilong_longctx,group_alias=None): {'babilong_qa1': ConfigurableTask(task_name=babilong_qa1,output_type=generate_until,num_fewshot=2,num_samples=1000), 'babilong_qa2': ConfigurableTask(task_name=babilong_qa2,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa3': ConfigurableTask(task_name=babilong_qa3,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa4': ConfigurableTask(task_name=babilong_qa4,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa5': ConfigurableTask(task_name=babilong_qa5,output_type=generate_until,num_fewshot=2,num_samples=999)}
21
- 2026-07-18:12:26:45 INFO [evaluator:314] babilong_qa1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
22
- 2026-07-18:12:26:45 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa1 from 2 to 2
23
- 2026-07-18:12:26:45 INFO [evaluator:314] babilong_qa2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
24
- 2026-07-18:12:26:45 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa2 from 2 to 2
25
- 2026-07-18:12:26:45 INFO [evaluator:314] babilong_qa3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
26
- 2026-07-18:12:26:45 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa3 from 2 to 2
27
- 2026-07-18:12:26:45 INFO [evaluator:314] babilong_qa4: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
28
- 2026-07-18:12:26:45 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa4 from 2 to 2
29
- 2026-07-18:12:26:45 INFO [evaluator:314] babilong_qa5: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
30
- 2026-07-18:12:26:45 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa5 from 2 to 2
31
- 2026-07-18:12:26:45 INFO [api.task:311] Building contexts for babilong_qa1 on rank 0...
32
 
33
-
34
  0%| | 0/50 [00:00<?, ?it/s]
35
- 2026-07-18:12:26:45 INFO [api.task:311] Building contexts for babilong_qa2 on rank 0...
36
  0%| | 0/50 [00:00<?, ?it/s]
 
37
 
38
-
39
  0%| | 0/50 [00:00<?, ?it/s]
40
- 2026-07-18:12:26:45 INFO [api.task:311] Building contexts for babilong_qa3 on rank 0...
41
  0%| | 0/50 [00:00<?, ?it/s]
 
42
 
43
-
44
  0%| | 0/50 [00:00<?, ?it/s]
45
- 2026-07-18:12:26:45 INFO [api.task:311] Building contexts for babilong_qa4 on rank 0...
46
  0%| | 0/50 [00:00<?, ?it/s]
 
47
 
48
-
49
  0%| | 0/50 [00:00<?, ?it/s]
50
- 2026-07-18:12:26:45 INFO [api.task:311] Building contexts for babilong_qa5 on rank 0...
51
  0%| | 0/50 [00:00<?, ?it/s]
 
52
 
53
-
54
  0%| | 0/50 [00:00<?, ?it/s]
55
- 2026-07-18:12:26:45 INFO [evaluator:584] Running generate_until requests
56
  0%| | 0/50 [00:00<?, ?it/s]
 
57
 
58
 
59
-
60
-
61
-
62
-
63
-
64
-
65
-
66
-
67
 
68
-
69
-
70
-
71
-
72
 
73
 
74
 
75
-
76
-
77
-
78
-
79
 
80
-
81
-
82
-
83
-
84
-
85
 
86
-
87
-
88
-
89
-
90
-
91
-
92
-
93
-
94
-
95
-
96
-
97
-
98
-
99
-
100
-
101
-
102
-
103
-
104
-
105
-
106
-
107
-
108
-
109
-
110
-
111
-
112
-
113
-
114
-
115
-
116
-
117
-
118
-
119
-
120
-
121
-
122
-
123
-
124
-
125
-
126
-
127
-
128
-
129
-
130
-
131
-
132
-
133
-
134
-
135
-
136
-
137
-
138
-
139
-
140
-
141
-
142
-
143
-
144
-
145
-
146
-
147
-
148
-
149
-
150
-
151
-
152
-
153
-
154
-
155
-
156
-
157
-
158
-
159
-
160
-
161
-
162
-
163
-
164
-
165
-
166
-
167
-
168
-
169
-
170
-
171
-
172
-
173
-
174
-
175
-
176
-
177
-
178
-
179
-
180
-
181
-
182
-
183
-
184
-
185
-
186
-
187
-
188
-
189
-
190
-
191
-
192
-
193
-
194
-
195
-
196
-
197
-
198
-
199
-
200
-
201
-
202
-
203
-
204
-
205
-
206
-
207
-
208
-
209
-
210
-
211
-
212
-
213
-
214
-
215
-
216
-
217
-
218
-
219
-
220
-
221
-
222
-
223
-
224
-
225
-
226
-
227
-
228
-
229
-
230
-
231
-
232
-
233
-
234
-
235
-
236
-
237
-
238
-
239
-
240
-
241
-
242
-
243
-
244
-
245
-
246
-
247
-
248
-
249
-
250
-
251
-
252
-
253
-
254
-
255
-
256
-
257
-
258
-
259
-
260
-
261
-
262
-
263
-
264
-
265
-
266
-
267
-
268
-
269
-
270
-
271
-
272
-
273
-
274
-
275
-
276
-
277
-
278
-
279
-
280
-
281
-
282
-
283
-
284
-
285
-
286
-
287
-
288
-
289
-
290
-
291
-
292
-
293
-
294
-
295
-
296
-
297
-
298
-
299
-
300
-
301
-
302
-
303
-
304
-
305
-
306
-
307
- 2026-07-18:12:31:42 INFO [loggers.evaluation_tracker:247] Saving results aggregated
308
- 2026-07-18:12:31:42 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/*.jsonl
309
  hf ({'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 50.0, num_fewshot: 2, batch_size: 1
310
  | Tasks |Version|Filter|n-shot|Metric| |Value| |Stderr|
311
  |----------------|------:|------|-----:|------|---|----:|---|-----:|
 
1
+ 2026-07-18:15:17:55 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
+ 2026-07-18:15:17:58 INFO [_cli.run:376] Selected Tasks: ['babilong_longctx']
3
+ 2026-07-18:15:18:00 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
+ 2026-07-18:15:18:00 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
+ 2026-07-18:15:18:03 INFO [models.huggingface:161] Using device 'cuda:0'
6
+ 2026-07-18:15:18:03 INFO [models.huggingface:548] Model type cannot be determined. Using default model type 'causal'
7
+ 2026-07-18:15:18:03 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
8
+ 2026-07-18:15:18:05 WARNING [api.task:856] babilong_qa5: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
9
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
10
+ 2026-07-18:15:18:05 WARNING [api.task:856] babilong_qa4: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
11
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
12
+ 2026-07-18:15:18:05 WARNING [api.task:856] babilong_qa3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
13
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
14
+ 2026-07-18:15:18:06 WARNING [api.task:856] babilong_qa2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
15
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
16
+ 2026-07-18:15:18:06 WARNING [api.task:856] babilong_qa1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
17
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
18
+ 2026-07-18:15:18:06 INFO [tasks:700] Selected tasks:
19
+ 2026-07-18:15:18:06 INFO [tasks:703] Group: babilong_longctx
20
+ 2026-07-18:15:18:06 INFO [tasks:726] ConfigurableGroup(group=babilong_longctx,group_alias=None): {'babilong_qa1': ConfigurableTask(task_name=babilong_qa1,output_type=generate_until,num_fewshot=2,num_samples=1000), 'babilong_qa2': ConfigurableTask(task_name=babilong_qa2,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa3': ConfigurableTask(task_name=babilong_qa3,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa4': ConfigurableTask(task_name=babilong_qa4,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa5': ConfigurableTask(task_name=babilong_qa5,output_type=generate_until,num_fewshot=2,num_samples=999)}
21
+ 2026-07-18:15:18:06 INFO [evaluator:314] babilong_qa1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
22
+ 2026-07-18:15:18:06 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa1 from 2 to 2
23
+ 2026-07-18:15:18:06 INFO [evaluator:314] babilong_qa2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
24
+ 2026-07-18:15:18:06 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa2 from 2 to 2
25
+ 2026-07-18:15:18:06 INFO [evaluator:314] babilong_qa3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
26
+ 2026-07-18:15:18:06 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa3 from 2 to 2
27
+ 2026-07-18:15:18:06 INFO [evaluator:314] babilong_qa4: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
28
+ 2026-07-18:15:18:06 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa4 from 2 to 2
29
+ 2026-07-18:15:18:06 INFO [evaluator:314] babilong_qa5: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
30
+ 2026-07-18:15:18:06 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa5 from 2 to 2
31
+ 2026-07-18:15:18:06 INFO [api.task:311] Building contexts for babilong_qa1 on rank 0...
32
 
 
33
  0%| | 0/50 [00:00<?, ?it/s]
34
+
35
  0%| | 0/50 [00:00<?, ?it/s]
36
+ 2026-07-18:15:18:06 INFO [api.task:311] Building contexts for babilong_qa2 on rank 0...
37
 
 
38
  0%| | 0/50 [00:00<?, ?it/s]
39
+
40
  0%| | 0/50 [00:00<?, ?it/s]
41
+ 2026-07-18:15:18:06 INFO [api.task:311] Building contexts for babilong_qa3 on rank 0...
42
 
 
43
  0%| | 0/50 [00:00<?, ?it/s]
44
+
45
  0%| | 0/50 [00:00<?, ?it/s]
46
+ 2026-07-18:15:18:07 INFO [api.task:311] Building contexts for babilong_qa4 on rank 0...
47
 
 
48
  0%| | 0/50 [00:00<?, ?it/s]
49
+
50
  0%| | 0/50 [00:00<?, ?it/s]
51
+ 2026-07-18:15:18:07 INFO [api.task:311] Building contexts for babilong_qa5 on rank 0...
52
 
 
53
  0%| | 0/50 [00:00<?, ?it/s]
54
+
55
  0%| | 0/50 [00:00<?, ?it/s]
56
+ 2026-07-18:15:18:07 INFO [evaluator:584] Running generate_until requests
57
 
58
 
59
+
60
+
61
+
62
+
63
+
64
+
65
+
66
+
67
 
68
+
69
+
70
+
71
+
72
 
73
 
74
 
75
+
76
+
77
+
78
+
79
 
80
+
81
+
82
+
83
+
84
+
85
 
86
+
87
+
88
+
89
+
90
+
91
+
92
+
93
+
94
+
95
+
96
+
97
+
98
+
99
+
100
+
101
+
102
+
103
+
104
+
105
+
106
+
107
+
108
+
109
+
110
+
111
+
112
+
113
+
114
+
115
+
116
+
117
+
118
+
119
+
120
+
121
+
122
+
123
+
124
+
125
+
126
+
127
+
128
+
129
+
130
+
131
+
132
+
133
+
134
+
135
+
136
+
137
+
138
+
139
+
140
+
141
+
142
+
143
+
144
+
145
+
146
+
147
+
148
+
149
+
150
+
151
+
152
+
153
+
154
+
155
+
156
+
157
+
158
+
159
+
160
+
161
+
162
+
163
+
164
+
165
+
166
+
167
+
168
+
169
+
170
+
171
+
172
+
173
+
174
+
175
+
176
+
177
+
178
+
179
+
180
+
181
+
182
+
183
+
184
+
185
+
186
+
187
+
188
+
189
+
190
+
191
+
192
+
193
+
194
+
195
+
196
+
197
+
198
+
199
+
200
+
201
+
202
+
203
+
204
+
205
+
206
+
207
+
208
+
209
+
210
+
211
+
212
+
213
+
214
+
215
+
216
+
217
+
218
+
219
+
220
+
221
+
222
+
223
+
224
+
225
+
226
+
227
+
228
+
229
+
230
+
231
+
232
+
233
+
234
+
235
+
236
+
237
+
238
+
239
+
240
+
241
+
242
+
243
+
244
+
245
+
246
+
247
+
248
+
249
+
250
+
251
+
252
+
253
+
254
+
255
+
256
+
257
+
258
+
259
+
260
+
261
+
262
+
263
+
264
+
265
+
266
+
267
+
268
+
269
+
270
+
271
+
272
+
273
+
274
+
275
+
276
+
277
+
278
+
279
+
280
+
281
+
282
+
283
+
284
+
285
+
286
+
287
+
288
+
289
+
290
+
291
+
292
+
293
+
294
+
295
+
296
+
297
+
298
+
299
+
300
+
301
+
302
+
303
+
304
+
305
+
306
+
307
+ 2026-07-18:15:23:38 INFO [loggers.evaluation_tracker:247] Saving results aggregated
308
+ 2026-07-18:15:23:38 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/*.jsonl
309
  hf ({'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 50.0, num_fewshot: 2, batch_size: 1
310
  | Tasks |Version|Filter|n-shot|Metric| |Value| |Stderr|
311
  |----------------|------:|------|-----:|------|---|----:|---|-----:|
outputs/eval/logs/token_t045/ruler_a.log CHANGED
@@ -1,70 +1,68 @@
1
- 2026-07-18:11:58:36 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
- 2026-07-18:11:58:39 INFO [_cli.run:376] Selected Tasks: ['niah_single_1', 'niah_single_3', 'niah_multikey_2', 'niah_multiquery', 'ruler_vt', 'ruler_fwe', 'ruler_qa_hotpot']
3
- 2026-07-18:11:58:41 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
- 2026-07-18:11:58:41 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
- 2026-07-18:11:58:44 INFO [models.huggingface:161] Using device 'cuda:0'
6
- 2026-07-18:11:58:44 INFO [models.huggingface:548] Model type cannot be determined. Using default model type 'causal'
7
- 2026-07-18:11:58:45 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
8
- 2026-07-18:11:58:54 WARNING [api.task:856] niah_single_1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
9
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
10
- 2026-07-18:11:58:54 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
11
 
12
 
13
-
14
-
15
-
16
-
17
-
18
-
19
-
20
-
21
-
22
- 2026-07-18:11:59:04 WARNING [api.task:856] niah_single_3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
23
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
24
 
25
 
26
-
27
-
28
-
29
-
30
-
31
-
32
-
33
-
34
-
35
- 2026-07-18:11:59:14 WARNING [api.task:856] niah_multikey_2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
36
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
37
 
38
 
39
-
40
-
41
-
42
-
43
-
44
-
45
-
46
- 2026-07-18:11:59:23 WARNING [api.task:856] niah_multiquery: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
47
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
48
 
49
 
50
-
51
-
52
-
53
-
54
-
55
-
56
-
57
-
58
-
59
- 2026-07-18:11:59:33 WARNING [api.task:856] ruler_vt: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
60
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
61
- 2026-07-18:11:59:33 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
62
  Max length 500 | Current length 307 | Noises: 5
63
  Max length 500 | Current length 410 | Noises: 10
64
  Max length 500 | Current length 531 | Noises: 15
65
  Num noises: 10
66
 
67
-
68
  0%| | 0/1 [00:00<?, ?it/s]
 
69
  0%| | 0/1 [00:00<?, ?it/s]
70
  Max length 8192 | Current length 791 | Noises: 10
71
  Max length 8192 | Current length 1025 | Noises: 20
72
  Max length 8192 | Current length 1273 | Noises: 30
@@ -100,238 +98,234 @@ Max length 8192 | Current length 8230 | Noises: 320
100
  Num noises: 310
101
 
102
 
103
  0%| | 0/500 [00:00<?, ?it/s]
104
-
105
  11%|█ | 55/500 [00:01<00:08, 54.38it/s]
106
-
107
  22%|██▏ | 110/500 [00:02<00:07, 54.20it/s]
108
-
109
  33%|███▎ | 165/500 [00:03<00:06, 54.37it/s]
110
-
111
  44%|████▍ | 220/500 [00:04<00:05, 54.34it/s]
112
-
113
  55%|█████▌ | 275/500 [00:05<00:04, 54.57it/s]
114
-
115
  66%|██████▌ | 331/500 [00:06<00:03, 55.02it/s]
116
-
117
  77%|███████▋ | 387/500 [00:07<00:02, 54.89it/s]
118
-
119
  89%|████████▉ | 444/500 [00:08<00:01, 55.23it/s]
120
-
121
- 2026-07-18:11:59:43 WARNING [api.task:856] ruler_fwe: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
122
  12%|█▏ | 58/500 [00:01<00:07, 57.66it/s]
 
123
  23%|██▎ | 116/500 [00:02<00:06, 57.08it/s]
 
124
  35%|███▍ | 174/500 [00:03<00:05, 57.31it/s]
 
125
  46%|████▋ | 232/500 [00:04<00:04, 57.12it/s]
 
126
  58%|█████▊ | 290/500 [00:05<00:03, 57.13it/s]
 
127
  70%|██████▉ | 348/500 [00:06<00:02, 57.22it/s]
 
128
  81%|████████ | 406/500 [00:07<00:01, 57.00it/s]
 
129
  93%|█████████▎| 464/500 [00:08<00:00, 56.37it/s]
 
130
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
131
 
132
 
133
-
134
-
135
-
136
-
137
-
138
-
139
-
140
-
141
-
142
-
143
-
144
-
145
-
146
- 2026-07-18:11:59:57 WARNING [api.task:856] ruler_qa_hotpot: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
147
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
148
 
149
 
150
-
151
-
152
-
153
-
154
-
155
-
156
-
157
-
158
-
159
-
160
-
161
-
162
-
163
-
164
-
165
-
166
-
167
-
168
-
169
-
170
-
171
- 2026-07-18:12:00:21 INFO [tasks:700] Selected tasks:
172
- 2026-07-18:12:00:21 INFO [tasks:691] Task: ruler_qa_hotpot (ruler/qa_hotpot.yaml)
173
- 2026-07-18:12:00:21 INFO [tasks:691] Task: ruler_fwe (ruler/fwe.yaml)
174
- 2026-07-18:12:00:21 INFO [tasks:691] Task: ruler_vt (ruler/vt.yaml)
175
- 2026-07-18:12:00:21 INFO [tasks:691] Task: niah_multiquery (ruler/niah_multiquery.yaml)
176
- 2026-07-18:12:00:21 INFO [tasks:691] Task: niah_multikey_2 (ruler/niah_multikey_2.yaml)
177
- 2026-07-18:12:00:21 INFO [tasks:691] Task: niah_single_3 (ruler/niah_single_3.yaml)
178
- 2026-07-18:12:00:21 INFO [tasks:691] Task: niah_single_1 (ruler/niah_single_1.yaml)
179
- 2026-07-18:12:00:21 INFO [evaluator:314] ruler_qa_hotpot: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 32, 'until': []}
180
- 2026-07-18:12:00:21 INFO [evaluator:314] ruler_fwe: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 50, 'until': []}
181
- 2026-07-18:12:00:21 INFO [evaluator:314] ruler_vt: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 30, 'until': []}
182
- 2026-07-18:12:00:21 INFO [evaluator:314] niah_multiquery: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
183
- 2026-07-18:12:00:21 INFO [evaluator:314] niah_multikey_2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
184
- 2026-07-18:12:00:21 INFO [evaluator:314] niah_single_3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
185
- 2026-07-18:12:00:21 INFO [evaluator:314] niah_single_1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
186
- 2026-07-18:12:00:21 INFO [api.task:311] Building contexts for ruler_qa_hotpot on rank 0...
187
 
188
-
189
  0%| | 0/20 [00:00<?, ?it/s]
190
- 2026-07-18:12:00:21 INFO [api.task:311] Building contexts for ruler_fwe on rank 0...
191
  0%| | 0/20 [00:00<?, ?it/s]
 
192
 
193
-
194
  0%| | 0/20 [00:00<?, ?it/s]
195
- 2026-07-18:12:00:21 INFO [api.task:311] Building contexts for ruler_vt on rank 0...
196
  0%| | 0/20 [00:00<?, ?it/s]
 
197
 
198
-
199
  0%| | 0/20 [00:00<?, ?it/s]
200
- 2026-07-18:12:00:21 INFO [api.task:311] Building contexts for niah_multiquery on rank 0...
201
  0%| | 0/20 [00:00<?, ?it/s]
 
202
 
203
-
204
  0%| | 0/20 [00:00<?, ?it/s]
205
- 2026-07-18:12:00:21 INFO [api.task:311] Building contexts for niah_multikey_2 on rank 0...
206
  0%| | 0/20 [00:00<?, ?it/s]
 
207
 
208
-
209
  0%| | 0/20 [00:00<?, ?it/s]
210
- 2026-07-18:12:00:21 INFO [api.task:311] Building contexts for niah_single_3 on rank 0...
211
  0%| | 0/20 [00:00<?, ?it/s]
 
212
 
213
-
214
  0%| | 0/20 [00:00<?, ?it/s]
215
- 2026-07-18:12:00:21 INFO [api.task:311] Building contexts for niah_single_1 on rank 0...
216
  0%| | 0/20 [00:00<?, ?it/s]
 
217
 
218
-
219
  0%| | 0/20 [00:00<?, ?it/s]
220
- 2026-07-18:12:00:21 INFO [evaluator:584] Running generate_until requests
221
  0%| | 0/20 [00:00<?, ?it/s]
 
222
 
223
 
224
 
225
-
226
-
227
-
228
-
229
-
230
-
231
-
232
-
233
-
234
-
235
-
236
-
237
-
238
-
239
-
240
-
241
-
242
-
243
-
244
-
245
-
246
-
247
-
248
-
249
-
250
-
251
-
252
-
253
-
254
-
255
-
256
-
257
-
258
-
259
-
260
-
261
-
262
-
263
-
264
-
265
-
266
-
267
-
268
-
269
-
270
-
271
-
272
-
273
-
274
-
275
-
276
-
277
-
278
-
279
-
280
-
281
-
282
-
283
-
284
-
285
-
286
-
287
-
288
-
289
-
290
-
291
-
292
-
293
-
294
-
295
-
296
-
297
-
298
-
299
-
300
-
301
-
302
-
303
-
304
-
305
-
306
-
307
-
308
-
309
-
310
-
311
-
312
-
313
-
314
-
315
-
316
-
317
-
318
-
319
-
320
-
321
-
322
-
323
-
324
-
325
-
326
-
327
-
328
-
329
-
330
-
331
-
332
-
333
-
334
-
335
-
336
-
337
-
338
-
339
-
340
-
341
-
342
-
343
-
344
-
345
-
346
-
347
-
348
-
349
-
350
-
351
-
352
-
353
-
354
-
355
-
356
-
357
-
358
-
359
-
360
-
361
-
362
-
363
-
364
- 2026-07-18:12:12:07 INFO [loggers.evaluation_tracker:247] Saving results aggregated
365
- 2026-07-18:12:12:07 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/*.jsonl
366
  hf ({'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 20.0, num_fewshot: None, batch_size: 1
367
  | Tasks |Version|Filter|n-shot|Metric| | Value | |Stderr|
368
  |---------------|------:|------|-----:|-----:|---|------:|---|------|
 
1
+ 2026-07-18:14:53:21 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
+ 2026-07-18:14:53:25 INFO [_cli.run:376] Selected Tasks: ['niah_single_1', 'niah_single_3', 'niah_multikey_2', 'niah_multiquery', 'ruler_vt', 'ruler_fwe', 'ruler_qa_hotpot']
3
+ 2026-07-18:14:53:26 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
+ 2026-07-18:14:53:26 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
+ 2026-07-18:14:53:29 INFO [models.huggingface:161] Using device 'cuda:0'
6
+ 2026-07-18:14:53:29 INFO [models.huggingface:548] Model type cannot be determined. Using default model type 'causal'
7
+ 2026-07-18:14:53:29 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
8
+ 2026-07-18:14:53:38 WARNING [api.task:856] niah_single_1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
9
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
10
+ 2026-07-18:14:53:38 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
11
 
12
 
13
+
14
+
15
+
16
+
17
+
18
+
19
+
20
+
21
+ 2026-07-18:14:53:47 WARNING [api.task:856] niah_single_3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
22
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
23
 
24
 
25
+
26
+
27
+
28
+
29
+
30
+
31
+
32
+
33
+ 2026-07-18:14:53:58 WARNING [api.task:856] niah_multikey_2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
34
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
35
 
36
 
37
+
38
+
39
+
40
+
41
+
42
+
43
+
44
+ 2026-07-18:14:54:05 WARNING [api.task:856] niah_multiquery: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
45
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
46
 
47
 
48
+
49
+
50
+
51
+
52
+
53
+
54
+
55
+
56
+
57
+ 2026-07-18:14:54:15 WARNING [api.task:856] ruler_vt: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
58
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
59
+ 2026-07-18:14:54:15 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
60
  Max length 500 | Current length 307 | Noises: 5
61
  Max length 500 | Current length 410 | Noises: 10
62
  Max length 500 | Current length 531 | Noises: 15
63
  Num noises: 10
64
 
 
65
  0%| | 0/1 [00:00<?, ?it/s]
66
+
67
  0%| | 0/1 [00:00<?, ?it/s]
68
  Max length 8192 | Current length 791 | Noises: 10
69
  Max length 8192 | Current length 1025 | Noises: 20
70
  Max length 8192 | Current length 1273 | Noises: 30
 
98
  Num noises: 310
99
 
100
 
101
  0%| | 0/500 [00:00<?, ?it/s]
 
102
  11%|█ | 55/500 [00:01<00:08, 54.38it/s]
 
103
  22%|██▏ | 110/500 [00:02<00:07, 54.20it/s]
 
104
  33%|███▎ | 165/500 [00:03<00:06, 54.37it/s]
 
105
  44%|████▍ | 220/500 [00:04<00:05, 54.34it/s]
 
106
  55%|█████▌ | 275/500 [00:05<00:04, 54.57it/s]
 
107
  66%|██████▌ | 331/500 [00:06<00:03, 55.02it/s]
 
108
  77%|███████▋ | 387/500 [00:07<00:02, 54.89it/s]
 
109
  89%|████████▉ | 444/500 [00:08<00:01, 55.23it/s]
110
+
 
111
  12%|█▏ | 58/500 [00:01<00:07, 57.66it/s]
112
+
113
  23%|██▎ | 116/500 [00:02<00:06, 57.08it/s]
114
+
115
  35%|███▍ | 174/500 [00:03<00:05, 57.31it/s]
116
+
117
  46%|████▋ | 232/500 [00:04<00:04, 57.12it/s]
118
+
119
  58%|█████▊ | 290/500 [00:05<00:03, 57.13it/s]
120
+
121
  70%|██████▉ | 348/500 [00:06<00:02, 57.22it/s]
122
+
123
  81%|████████ | 406/500 [00:07<00:01, 57.00it/s]
124
+
125
  93%|█████████▎| 464/500 [00:08<00:00, 56.37it/s]
126
+ 2026-07-18:14:54:25 WARNING [api.task:856] ruler_fwe: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
127
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
128
 
129
 
130
+
131
+
132
+
133
+
134
+
135
+
136
+
137
+
138
+
139
+
140
+
141
+
142
+ 2026-07-18:14:54:38 WARNING [api.task:856] ruler_qa_hotpot: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
143
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
144
 
145
 
146
+
147
+
148
+
149
+
150
+
151
+
152
+
153
+
154
+
155
+
156
+
157
+
158
+
159
+
160
+
161
+
162
+
163
+
164
+
165
+ 2026-07-18:14:55:01 INFO [tasks:700] Selected tasks:
166
+ 2026-07-18:14:55:01 INFO [tasks:691] Task: ruler_qa_hotpot (ruler/qa_hotpot.yaml)
167
+ 2026-07-18:14:55:01 INFO [tasks:691] Task: ruler_fwe (ruler/fwe.yaml)
168
+ 2026-07-18:14:55:01 INFO [tasks:691] Task: ruler_vt (ruler/vt.yaml)
169
+ 2026-07-18:14:55:01 INFO [tasks:691] Task: niah_multiquery (ruler/niah_multiquery.yaml)
170
+ 2026-07-18:14:55:01 INFO [tasks:691] Task: niah_multikey_2 (ruler/niah_multikey_2.yaml)
171
+ 2026-07-18:14:55:01 INFO [tasks:691] Task: niah_single_3 (ruler/niah_single_3.yaml)
172
+ 2026-07-18:14:55:01 INFO [tasks:691] Task: niah_single_1 (ruler/niah_single_1.yaml)
173
+ 2026-07-18:14:55:01 INFO [evaluator:314] ruler_qa_hotpot: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 32, 'until': []}
174
+ 2026-07-18:14:55:01 INFO [evaluator:314] ruler_fwe: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 50, 'until': []}
175
+ 2026-07-18:14:55:01 INFO [evaluator:314] ruler_vt: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 30, 'until': []}
176
+ 2026-07-18:14:55:01 INFO [evaluator:314] niah_multiquery: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
177
+ 2026-07-18:14:55:01 INFO [evaluator:314] niah_multikey_2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
178
+ 2026-07-18:14:55:01 INFO [evaluator:314] niah_single_3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
179
+ 2026-07-18:14:55:01 INFO [evaluator:314] niah_single_1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
180
+ 2026-07-18:14:55:01 INFO [api.task:311] Building contexts for ruler_qa_hotpot on rank 0...
 
 
181
 
 
182
  0%| | 0/20 [00:00<?, ?it/s]
183
+
184
  0%| | 0/20 [00:00<?, ?it/s]
185
+ 2026-07-18:14:55:01 INFO [api.task:311] Building contexts for ruler_fwe on rank 0...
186
 
 
187
  0%| | 0/20 [00:00<?, ?it/s]
188
+
189
  0%| | 0/20 [00:00<?, ?it/s]
190
+ 2026-07-18:14:55:01 INFO [api.task:311] Building contexts for ruler_vt on rank 0...
191
 
 
192
  0%| | 0/20 [00:00<?, ?it/s]
193
+
194
  0%| | 0/20 [00:00<?, ?it/s]
195
+ 2026-07-18:14:55:01 INFO [api.task:311] Building contexts for niah_multiquery on rank 0...
196
 
 
197
  0%| | 0/20 [00:00<?, ?it/s]
198
+
199
  0%| | 0/20 [00:00<?, ?it/s]
200
+ 2026-07-18:14:55:01 INFO [api.task:311] Building contexts for niah_multikey_2 on rank 0...
201
 
 
202
  0%| | 0/20 [00:00<?, ?it/s]
203
+
204
  0%| | 0/20 [00:00<?, ?it/s]
205
+ 2026-07-18:14:55:01 INFO [api.task:311] Building contexts for niah_single_3 on rank 0...
206
 
 
207
  0%| | 0/20 [00:00<?, ?it/s]
208
+
209
  0%| | 0/20 [00:00<?, ?it/s]
210
+ 2026-07-18:14:55:01 INFO [api.task:311] Building contexts for niah_single_1 on rank 0...
211
 
 
212
  0%| | 0/20 [00:00<?, ?it/s]
213
+
214
  0%| | 0/20 [00:00<?, ?it/s]
215
+ 2026-07-18:14:55:01 INFO [evaluator:584] Running generate_until requests
216
 
217
 
218
 
219
+
220
+
221
+
222
+
223
+
224
+
225
+
226
+
227
+
228
+
229
+
230
+
231
+
232
+
233
+
234
+
235
+
236
+
237
+
238
+
239
+
240
+
241
+
242
+
243
+
244
+
245
+
246
+
247
+
248
+
249
+
250
+
251
+
252
+
253
+
254
+
255
+
256
+
257
+
258
+
259
+
260
+
261
+
262
+
263
+
264
+
265
+
266
+
267
+
268
+
269
+
270
+
271
+
272
+
273
+
274
+
275
+
276
+
277
+
278
+
279
+
280
+
281
+
282
+
283
+
284
+
285
+
286
+
287
+
288
+
289
+
290
+
291
+
292
+
293
+
294
+
295
+
296
+
297
+
298
+
299
+
300
+
301
+
302
+
303
+
304
+
305
+
306
+
307
+
308
+
309
+
310
+
311
+
312
+
313
+
314
+
315
+
316
+
317
+
318
+
319
+
320
+
321
+
322
+
323
+
324
+
325
+
326
+
327
+
328
+
329
+
330
+
331
+
332
+
333
+
334
+
335
+
336
+
337
+
338
+
339
+
340
+
341
+
342
+
343
+
344
+
345
+
346
+
347
+
348
+
349
+
350
+
351
+
352
+
353
+
354
+
355
+
356
+
357
+
358
+ 2026-07-18:15:05:41 INFO [loggers.evaluation_tracker:247] Saving results aggregated
359
+ 2026-07-18:15:05:41 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/*.jsonl
360
  hf ({'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 20.0, num_fewshot: None, batch_size: 1
361
  | Tasks |Version|Filter|n-shot|Metric| | Value | |Stderr|
362
  |---------------|------:|------|-----:|-----:|---|------:|---|------|
outputs/eval/logs/token_t045/ruler_b.log CHANGED
@@ -1,239 +1,236 @@
1
- 2026-07-18:12:12:10 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
- 2026-07-18:12:12:14 INFO [_cli.run:376] Selected Tasks: ['niah_single_2', 'niah_multikey_1', 'niah_multikey_3', 'niah_multivalue', 'ruler_cwe', 'ruler_qa_squad']
3
- 2026-07-18:12:12:15 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
- 2026-07-18:12:12:15 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
- 2026-07-18:12:12:18 INFO [models.huggingface:161] Using device 'cuda:0'
6
- 2026-07-18:12:12:18 INFO [models.huggingface:548] Model type cannot be determined. Using default model type 'causal'
7
- 2026-07-18:12:12:19 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
8
- 2026-07-18:12:12:28 WARNING [api.task:856] niah_single_2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
9
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
10
- 2026-07-18:12:12:28 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
11
 
12
 
13
-
14
-
15
-
16
-
17
-
18
-
19
-
20
-
21
-
22
- 2026-07-18:12:12:39 WARNING [api.task:856] niah_multikey_1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
23
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
24
 
25
 
26
-
27
-
28
-
29
-
30
-
31
-
32
-
33
-
34
- 2026-07-18:12:12:48 WARNING [api.task:856] niah_multikey_3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
35
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
36
 
37
 
38
-
39
-
40
-
41
-
42
-
43
-
44
- 2026-07-18:12:12:55 WARNING [api.task:856] niah_multivalue: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
45
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
46
 
47
 
48
-
49
-
50
-
51
-
52
-
53
-
54
-
55
-
56
-
57
- 2026-07-18:12:13:05 WARNING [api.task:856] ruler_cwe: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
58
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
59
- 2026-07-18:12:13:05 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
60
 
61
 
62
-
63
-
64
-
65
-
66
-
67
-
68
- 2026-07-18:12:13:12 WARNING [api.task:856] ruler_qa_squad: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
69
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
70
 
71
 
72
-
73
-
74
-
75
-
76
-
77
-
78
-
79
-
80
-
81
- 2026-07-18:12:13:23 INFO [tasks:700] Selected tasks:
82
- 2026-07-18:12:13:23 INFO [tasks:691] Task: ruler_qa_squad (ruler/qa_squad.yaml)
83
- 2026-07-18:12:13:23 INFO [tasks:691] Task: ruler_cwe (ruler/cwe.yaml)
84
- 2026-07-18:12:13:23 INFO [tasks:691] Task: niah_multivalue (ruler/niah_multivalue.yaml)
85
- 2026-07-18:12:13:23 INFO [tasks:691] Task: niah_multikey_3 (ruler/niah_multikey_3.yaml)
86
- 2026-07-18:12:13:23 INFO [tasks:691] Task: niah_multikey_1 (ruler/niah_multikey_1.yaml)
87
- 2026-07-18:12:13:23 INFO [tasks:691] Task: niah_single_2 (ruler/niah_single_2.yaml)
88
- 2026-07-18:12:13:23 INFO [evaluator:314] ruler_qa_squad: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 32, 'until': []}
89
- 2026-07-18:12:13:23 INFO [evaluator:314] ruler_cwe: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 120, 'until': []}
90
- 2026-07-18:12:13:23 INFO [evaluator:314] niah_multivalue: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
91
- 2026-07-18:12:13:23 INFO [evaluator:314] niah_multikey_3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
92
- 2026-07-18:12:13:23 INFO [evaluator:314] niah_multikey_1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
93
- 2026-07-18:12:13:23 INFO [evaluator:314] niah_single_2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
94
- 2026-07-18:12:13:23 INFO [api.task:311] Building contexts for ruler_qa_squad on rank 0...
95
 
96
-
97
  0%| | 0/20 [00:00<?, ?it/s]
98
- 2026-07-18:12:13:23 INFO [api.task:311] Building contexts for ruler_cwe on rank 0...
99
  0%| | 0/20 [00:00<?, ?it/s]
 
100
 
101
-
102
  0%| | 0/20 [00:00<?, ?it/s]
103
- 2026-07-18:12:13:23 INFO [api.task:311] Building contexts for niah_multivalue on rank 0...
104
  0%| | 0/20 [00:00<?, ?it/s]
 
105
 
106
-
107
  0%| | 0/20 [00:00<?, ?it/s]
108
- 2026-07-18:12:13:23 INFO [api.task:311] Building contexts for niah_multikey_3 on rank 0...
109
  0%| | 0/20 [00:00<?, ?it/s]
 
110
 
111
-
112
  0%| | 0/20 [00:00<?, ?it/s]
113
- 2026-07-18:12:13:23 INFO [api.task:311] Building contexts for niah_multikey_1 on rank 0...
114
  0%| | 0/20 [00:00<?, ?it/s]
 
115
 
116
-
117
  0%| | 0/20 [00:00<?, ?it/s]
118
- 2026-07-18:12:13:23 INFO [api.task:311] Building contexts for niah_single_2 on rank 0...
119
  0%| | 0/20 [00:00<?, ?it/s]
 
120
 
121
-
122
  0%| | 0/20 [00:00<?, ?it/s]
123
- 2026-07-18:12:13:23 INFO [evaluator:584] Running generate_until requests
124
  0%| | 0/20 [00:00<?, ?it/s]
 
125
 
126
 
127
-
128
-
129
-
130
-
131
-
132
-
133
-
134
-
135
-
136
-
137
-
138
-
139
-
140
-
141
-
142
-
143
-
144
-
145
-
146
-
147
-
148
-
149
-
150
-
151
-
152
-
153
-
154
-
155
-
156
-
157
-
158
-
159
-
160
-
161
-
162
-
163
-
164
-
165
-
166
-
167
-
168
-
169
-
170
-
171
-
172
-
173
-
174
-
175
-
176
-
177
-
178
-
179
-
180
-
181
-
182
-
183
-
184
-
185
-
186
-
187
-
188
-
189
-
190
-
191
-
192
-
193
-
194
-
195
-
196
-
197
-
198
-
199
-
200
-
201
-
202
-
203
-
204
-
205
-
206
-
207
-
208
-
209
-
210
-
211
-
212
-
213
-
214
-
215
-
216
-
217
-
218
-
219
-
220
-
221
-
222
-
223
-
224
-
225
-
226
-
227
-
228
-
229
-
230
-
231
-
232
-
233
-
234
-
235
-
236
-
237
-
238
-
239
-
240
-
241
-
242
-
243
-
244
-
245
-
246
-
247
- 2026-07-18:12:26:30 INFO [loggers.evaluation_tracker:247] Saving results aggregated
248
- 2026-07-18:12:26:30 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/*.jsonl
249
  hf ({'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 20.0, num_fewshot: None, batch_size: 1
250
  | Tasks |Version|Filter|n-shot|Metric| | Value | |Stderr|
251
  |---------------|------:|------|-----:|-----:|---|------:|---|------|
 
1
+ 2026-07-18:15:05:44 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
+ 2026-07-18:15:05:47 INFO [_cli.run:376] Selected Tasks: ['niah_single_2', 'niah_multikey_1', 'niah_multikey_3', 'niah_multivalue', 'ruler_cwe', 'ruler_qa_squad']
3
+ 2026-07-18:15:05:49 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
+ 2026-07-18:15:05:49 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
+ 2026-07-18:15:05:52 INFO [models.huggingface:161] Using device 'cuda:0'
6
+ 2026-07-18:15:05:52 INFO [models.huggingface:548] Model type cannot be determined. Using default model type 'causal'
7
+ 2026-07-18:15:05:52 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
8
+ 2026-07-18:15:06:01 WARNING [api.task:856] niah_single_2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
9
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
10
+ 2026-07-18:15:06:02 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
11
 
12
 
13
+
14
+
15
+
16
+
17
+
18
+
19
+
20
+
21
+ 2026-07-18:15:06:11 WARNING [api.task:856] niah_multikey_1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
22
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
23
 
24
 
25
+
26
+
27
+
28
+
29
+
30
+
31
+
32
+
33
+ 2026-07-18:15:06:21 WARNING [api.task:856] niah_multikey_3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
34
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
35
 
36
 
37
+
38
+
39
+
40
+
41
+
42
+ 2026-07-18:15:06:27 WARNING [api.task:856] niah_multivalue: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
43
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
44
 
45
 
46
+
47
+
48
+
49
+
50
+
51
+
52
+
53
+
54
+ 2026-07-18:15:06:37 WARNING [api.task:856] ruler_cwe: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
55
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
56
+ 2026-07-18:15:06:37 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
57
 
58
 
59
+
60
+
61
+
62
+
63
+
64
+
65
+ 2026-07-18:15:06:43 WARNING [api.task:856] ruler_qa_squad: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
66
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
67
 
68
 
69
+
70
+
71
+
72
+
73
+
74
+
75
+
76
+
77
+
78
+ 2026-07-18:15:06:54 INFO [tasks:700] Selected tasks:
79
+ 2026-07-18:15:06:54 INFO [tasks:691] Task: ruler_qa_squad (ruler/qa_squad.yaml)
80
+ 2026-07-18:15:06:54 INFO [tasks:691] Task: ruler_cwe (ruler/cwe.yaml)
81
+ 2026-07-18:15:06:54 INFO [tasks:691] Task: niah_multivalue (ruler/niah_multivalue.yaml)
82
+ 2026-07-18:15:06:54 INFO [tasks:691] Task: niah_multikey_3 (ruler/niah_multikey_3.yaml)
83
+ 2026-07-18:15:06:54 INFO [tasks:691] Task: niah_multikey_1 (ruler/niah_multikey_1.yaml)
84
+ 2026-07-18:15:06:54 INFO [tasks:691] Task: niah_single_2 (ruler/niah_single_2.yaml)
85
+ 2026-07-18:15:06:54 INFO [evaluator:314] ruler_qa_squad: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 32, 'until': []}
86
+ 2026-07-18:15:06:54 INFO [evaluator:314] ruler_cwe: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 120, 'until': []}
87
+ 2026-07-18:15:06:54 INFO [evaluator:314] niah_multivalue: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
88
+ 2026-07-18:15:06:54 INFO [evaluator:314] niah_multikey_3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
89
+ 2026-07-18:15:06:54 INFO [evaluator:314] niah_multikey_1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
90
+ 2026-07-18:15:06:54 INFO [evaluator:314] niah_single_2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
91
+ 2026-07-18:15:06:54 INFO [api.task:311] Building contexts for ruler_qa_squad on rank 0...
92
 
 
93
  0%| | 0/20 [00:00<?, ?it/s]
94
+
95
  0%| | 0/20 [00:00<?, ?it/s]
96
+ 2026-07-18:15:06:54 INFO [api.task:311] Building contexts for ruler_cwe on rank 0...
97
 
 
98
  0%| | 0/20 [00:00<?, ?it/s]
99
+
100
  0%| | 0/20 [00:00<?, ?it/s]
101
+ 2026-07-18:15:06:54 INFO [api.task:311] Building contexts for niah_multivalue on rank 0...
102
 
 
103
  0%| | 0/20 [00:00<?, ?it/s]
104
+
105
  0%| | 0/20 [00:00<?, ?it/s]
106
+ 2026-07-18:15:06:54 INFO [api.task:311] Building contexts for niah_multikey_3 on rank 0...
107
 
 
108
  0%| | 0/20 [00:00<?, ?it/s]
109
+
110
  0%| | 0/20 [00:00<?, ?it/s]
111
+ 2026-07-18:15:06:54 INFO [api.task:311] Building contexts for niah_multikey_1 on rank 0...
112
 
 
113
  0%| | 0/20 [00:00<?, ?it/s]
114
+
115
  0%| | 0/20 [00:00<?, ?it/s]
116
+ 2026-07-18:15:06:54 INFO [api.task:311] Building contexts for niah_single_2 on rank 0...
117
 
 
118
  0%| | 0/20 [00:00<?, ?it/s]
119
+
120
  0%| | 0/20 [00:00<?, ?it/s]
121
+ 2026-07-18:15:06:54 INFO [evaluator:584] Running generate_until requests
122
 
123
 
124
+
125
+
126
+
127
+
128
+
129
+
130
+
131
+
132
+
133
+
134
+
135
+
136
+
137
+
138
+
139
+
140
+
141
+
142
+
143
+
144
+
145
+
146
+
147
+
148
+
149
+
150
+
151
+
152
+
153
+
154
+
155
+
156
+
157
+
158
+
159
+
160
+
161
+
162
+
163
+
164
+
165
+
166
+
167
+
168
+
169
+
170
+
171
+
172
+
173
+
174
+
175
+
176
+
177
+
178
+
179
+
180
+
181
+
182
+
183
+
184
+
185
+
186
+
187
+
188
+
189
+
190
+
191
+
192
+
193
+
194
+
195
+
196
+
197
+
198
+
199
+
200
+
201
+
202
+
203
+
204
+
205
+
206
+
207
+
208
+
209
+
210
+
211
+
212
+
213
+
214
+
215
+
216
+
217
+
218
+
219
+
220
+
221
+
222
+
223
+
224
+
225
+
226
+
227
+
228
+
229
+
230
+
231
+
232
+
233
+
234
+
235
+
236
+
237
+
238
+
239
+
240
+
241
+
242
+
243
+
244
+ 2026-07-18:15:17:52 INFO [loggers.evaluation_tracker:247] Saving results aggregated
245
+ 2026-07-18:15:17:52 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/*.jsonl
246
  hf ({'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 20.0, num_fewshot: None, batch_size: 1
247
  | Tasks |Version|Filter|n-shot|Metric| | Value | |Stderr|
248
  |---------------|------:|------|-----:|-----:|---|------:|---|------|
outputs/eval/logs/vanilla/babilong.log CHANGED
@@ -1,449 +1,449 @@
1
- 2026-07-18:11:46:01 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
- 2026-07-18:11:46:04 INFO [_cli.run:376] Selected Tasks: ['babilong_longctx']
3
- 2026-07-18:11:46:06 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
- 2026-07-18:11:46:06 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
- 2026-07-18:11:46:09 INFO [models.huggingface:161] Using device 'cuda:0'
6
- 2026-07-18:11:46:09 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
7
- 2026-07-18:11:46:11 WARNING [api.task:856] babilong_qa5: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
8
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
9
 
10
-
11
 
12
-
13
 
14
-
15
 
16
-
17
 
18
-
19
 
20
-
21
 
22
-
23
 
24
-
25
 
26
-
27
 
28
-
29
 
30
-
31
- 2026-07-18:11:46:14 WARNING [api.task:856] babilong_qa4: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
32
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
33
- 2026-07-18:11:46:14 WARNING [api.task:856] babilong_qa3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
34
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
35
- 2026-07-18:11:46:15 WARNING [api.task:856] babilong_qa2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
36
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
37
- 2026-07-18:11:46:15 WARNING [api.task:856] babilong_qa1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
38
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
39
- 2026-07-18:11:46:15 INFO [tasks:700] Selected tasks:
40
- 2026-07-18:11:46:15 INFO [tasks:703] Group: babilong_longctx
41
- 2026-07-18:11:46:15 INFO [tasks:726] ConfigurableGroup(group=babilong_longctx,group_alias=None): {'babilong_qa1': ConfigurableTask(task_name=babilong_qa1,output_type=generate_until,num_fewshot=2,num_samples=1000), 'babilong_qa2': ConfigurableTask(task_name=babilong_qa2,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa3': ConfigurableTask(task_name=babilong_qa3,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa4': ConfigurableTask(task_name=babilong_qa4,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa5': ConfigurableTask(task_name=babilong_qa5,output_type=generate_until,num_fewshot=2,num_samples=999)}
42
- 2026-07-18:11:46:15 INFO [evaluator:314] babilong_qa1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
43
- 2026-07-18:11:46:15 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa1 from 2 to 2
44
- 2026-07-18:11:46:15 INFO [evaluator:314] babilong_qa2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
45
- 2026-07-18:11:46:15 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa2 from 2 to 2
46
- 2026-07-18:11:46:15 INFO [evaluator:314] babilong_qa3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
47
- 2026-07-18:11:46:15 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa3 from 2 to 2
48
- 2026-07-18:11:46:15 INFO [evaluator:314] babilong_qa4: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
49
- 2026-07-18:11:46:15 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa4 from 2 to 2
50
- 2026-07-18:11:46:15 INFO [evaluator:314] babilong_qa5: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
51
- 2026-07-18:11:46:15 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa5 from 2 to 2
52
- 2026-07-18:11:46:15 INFO [api.task:311] Building contexts for babilong_qa1 on rank 0...
53
-
54
-
55
  0%| | 0/50 [00:00<?, ?it/s]
56
- 2026-07-18:11:46:15 INFO [api.task:311] Building contexts for babilong_qa2 on rank 0...
57
-
58
-
59
  0%| | 0/50 [00:00<?, ?it/s]
60
- 2026-07-18:11:46:15 INFO [api.task:311] Building contexts for babilong_qa3 on rank 0...
61
-
62
-
63
  0%| | 0/50 [00:00<?, ?it/s]
64
- 2026-07-18:11:46:15 INFO [api.task:311] Building contexts for babilong_qa4 on rank 0...
65
-
66
-
67
  0%| | 0/50 [00:00<?, ?it/s]
68
- 2026-07-18:11:46:15 INFO [api.task:311] Building contexts for babilong_qa5 on rank 0...
69
-
70
-
71
  0%| | 0/50 [00:00<?, ?it/s]
72
- 2026-07-18:11:46:16 INFO [evaluator:584] Running generate_until requests
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
73
  0%| | 0/50 [00:00<?, ?it/s]
 
 
 
74
  0%| | 0/50 [00:00<?, ?it/s]
 
 
 
75
  0%| | 0/50 [00:00<?, ?it/s]
 
 
 
76
  0%| | 0/50 [00:00<?, ?it/s]
 
 
 
77
  0%| | 0/50 [00:00<?, ?it/s]
 
78
 
79
 
80
 
81
-
82
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
83
 
84
-
85
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
86
 
87
-
88
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
89
 
90
-
91
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
92
 
93
-
94
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
95
 
96
-
97
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
98
 
99
-
100
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
101
 
102
-
103
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
104
 
105
-
106
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
107
 
108
-
109
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
110
 
111
-
112
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
113
 
114
-
115
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
116
 
117
-
118
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
119
 
120
-
121
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
122
 
123
-
124
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
125
 
126
-
127
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
128
 
129
-
130
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
131
 
132
-
133
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
134
 
135
-
136
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
137
 
138
-
139
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
140
 
141
-
142
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
143
 
144
-
145
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
146
 
147
-
148
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
149
 
150
-
151
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
152
 
153
-
154
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
155
 
156
-
157
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
158
 
159
-
160
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
161
 
162
-
163
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
164
 
165
-
166
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
167
 
168
-
169
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
170
 
171
-
172
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
173
 
174
-
175
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
176
 
177
-
178
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
179
 
180
-
181
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
182
 
183
-
184
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
185
 
186
-
187
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
188
 
189
-
190
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
191
 
192
-
193
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
194
 
195
-
196
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
197
 
198
-
199
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
200
 
201
-
202
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
203
 
204
-
205
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
206
 
207
-
208
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
209
 
210
-
211
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
212
 
213
-
214
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
215
 
216
-
217
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
218
 
219
-
220
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
221
 
222
-
223
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
224
 
225
-
226
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
227
 
228
-
229
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
230
 
231
-
232
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
233
 
234
-
235
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
236
 
237
-
238
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
239
 
240
-
241
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
242
 
243
-
244
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
245
 
246
-
247
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
248
 
249
-
250
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
251
 
252
-
253
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
254
 
255
-
256
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
257
 
258
-
259
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
260
 
261
-
262
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
263
 
264
-
265
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
266
 
267
-
268
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
269
 
270
-
271
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
272
 
273
-
274
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
275
 
276
-
277
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
278
 
279
-
280
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
281
 
282
-
283
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
284
 
285
-
286
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
287
 
288
-
289
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
290
 
291
-
292
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
293
 
294
-
295
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
296
 
297
-
298
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
299
 
300
-
301
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
302
 
303
-
304
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
305
 
306
-
307
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
308
 
309
-
310
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
311
 
312
-
313
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
314
 
315
-
316
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
317
 
318
-
319
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
320
 
321
-
322
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
323
 
324
-
325
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
326
 
327
-
328
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
329
 
330
-
331
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
332
 
333
-
334
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
335
 
336
-
337
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
338
 
339
-
340
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
341
 
342
-
343
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
344
 
345
-
346
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
347
 
348
-
349
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
350
 
351
-
352
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
353
 
354
-
355
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
356
 
357
-
358
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
359
 
360
-
361
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
362
 
363
-
364
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
365
 
366
-
367
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
368
 
369
-
370
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
371
 
372
-
373
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
374
 
375
-
376
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
377
 
378
-
379
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
380
 
381
-
382
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
383
 
384
-
385
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
386
 
387
-
388
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
389
 
390
-
391
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
392
 
393
-
394
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
395
 
396
-
397
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
398
 
399
-
400
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
401
 
402
-
403
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
404
 
405
-
406
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
407
 
408
-
409
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
410
 
411
-
412
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
413
 
414
-
415
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
416
 
417
-
418
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
419
 
420
-
421
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
422
 
423
-
424
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
425
 
426
-
427
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
428
 
429
-
430
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
431
 
432
-
433
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
434
 
435
-
436
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
437
 
438
-
439
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
440
 
441
-
442
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
443
 
444
-
445
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
446
 
447
-
448
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
449
 
450
-
451
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
452
 
453
-
454
-
455
- 2026-07-18:11:49:27 INFO [loggers.evaluation_tracker:247] Saving results aggregated
456
- 2026-07-18:11:49:27 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/vanilla/babilong8k_qa1_qa5_n50/__workspace/*.jsonl
457
  hf ({'pretrained': '/workspace', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 50.0, num_fewshot: 2, batch_size: 1
458
  | Tasks |Version|Filter|n-shot|Metric| |Value| |Stderr|
459
  |----------------|------:|------|-----:|------|---|----:|---|-----:|
 
1
+ 2026-07-18:14:38:06 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
+ 2026-07-18:14:38:10 INFO [_cli.run:376] Selected Tasks: ['babilong_longctx']
3
+ 2026-07-18:14:38:11 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
+ 2026-07-18:14:38:11 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
+ 2026-07-18:14:38:15 INFO [models.huggingface:161] Using device 'cuda:0'
6
+ 2026-07-18:14:38:15 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
7
+ 2026-07-18:14:38:17 WARNING [api.task:856] babilong_qa5: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
8
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
9
 
10
+
11
 
12
+
13
 
14
+
15
 
16
+
17
 
18
+
19
 
20
+
21
 
22
+
23
 
24
+
25
 
26
+
27
 
28
+
29
 
30
+
31
+ 2026-07-18:14:38:20 WARNING [api.task:856] babilong_qa4: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
32
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
33
+ 2026-07-18:14:38:20 WARNING [api.task:856] babilong_qa3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
34
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
35
+ 2026-07-18:14:38:20 WARNING [api.task:856] babilong_qa2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
36
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
37
+ 2026-07-18:14:38:20 WARNING [api.task:856] babilong_qa1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
38
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
39
  0%| | 0/50 [00:00<?, ?it/s]
 
 
 
40
  0%| | 0/50 [00:00<?, ?it/s]
 
 
 
41
  0%| | 0/50 [00:00<?, ?it/s]
 
 
 
42
  0%| | 0/50 [00:00<?, ?it/s]
 
 
 
43
  0%| | 0/50 [00:00<?, ?it/s]
44
+ 2026-07-18:14:38:20 INFO [tasks:700] Selected tasks:
45
+ 2026-07-18:14:38:20 INFO [tasks:703] Group: babilong_longctx
46
+ 2026-07-18:14:38:20 INFO [tasks:726] ConfigurableGroup(group=babilong_longctx,group_alias=None): {'babilong_qa1': ConfigurableTask(task_name=babilong_qa1,output_type=generate_until,num_fewshot=2,num_samples=1000), 'babilong_qa2': ConfigurableTask(task_name=babilong_qa2,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa3': ConfigurableTask(task_name=babilong_qa3,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa4': ConfigurableTask(task_name=babilong_qa4,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa5': ConfigurableTask(task_name=babilong_qa5,output_type=generate_until,num_fewshot=2,num_samples=999)}
47
+ 2026-07-18:14:38:20 INFO [evaluator:314] babilong_qa1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
48
+ 2026-07-18:14:38:20 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa1 from 2 to 2
49
+ 2026-07-18:14:38:20 INFO [evaluator:314] babilong_qa2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
50
+ 2026-07-18:14:38:20 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa2 from 2 to 2
51
+ 2026-07-18:14:38:20 INFO [evaluator:314] babilong_qa3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
52
+ 2026-07-18:14:38:20 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa3 from 2 to 2
53
+ 2026-07-18:14:38:20 INFO [evaluator:314] babilong_qa4: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
54
+ 2026-07-18:14:38:20 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa4 from 2 to 2
55
+ 2026-07-18:14:38:20 INFO [evaluator:314] babilong_qa5: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
56
+ 2026-07-18:14:38:20 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa5 from 2 to 2
57
+ 2026-07-18:14:38:20 INFO [api.task:311] Building contexts for babilong_qa1 on rank 0...
58
+
59
+
60
  0%| | 0/50 [00:00<?, ?it/s]
61
+ 2026-07-18:14:38:20 INFO [api.task:311] Building contexts for babilong_qa2 on rank 0...
62
+
63
+
64
  0%| | 0/50 [00:00<?, ?it/s]
65
+ 2026-07-18:14:38:21 INFO [api.task:311] Building contexts for babilong_qa3 on rank 0...
66
+
67
+
68
  0%| | 0/50 [00:00<?, ?it/s]
69
+ 2026-07-18:14:38:21 INFO [api.task:311] Building contexts for babilong_qa4 on rank 0...
70
+
71
+
72
  0%| | 0/50 [00:00<?, ?it/s]
73
+ 2026-07-18:14:38:21 INFO [api.task:311] Building contexts for babilong_qa5 on rank 0...
74
+
75
+
76
  0%| | 0/50 [00:00<?, ?it/s]
77
+ 2026-07-18:14:38:21 INFO [evaluator:584] Running generate_until requests
78
 
79
 
80
 
81
+
82
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
83
 
84
+
85
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
86
 
87
+
88
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
89
 
90
+
91
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
92
 
93
+
94
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
95
 
96
+
97
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
98
 
99
+
100
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
101
 
102
+
103
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
104
 
105
+
106
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
107
 
108
+
109
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
110
 
111
+
112
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
113
 
114
+
115
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
116
 
117
+
118
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
119
 
120
+
121
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
122
 
123
+
124
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
125
 
126
+
127
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
128
 
129
+
130
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
131
 
132
+
133
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
134
 
135
+
136
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
137
 
138
+
139
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
140
 
141
+
142
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
143
 
144
+
145
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
146
 
147
+
148
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
149
 
150
+
151
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
152
 
153
+
154
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
155
 
156
+
157
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
158
 
159
+
160
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
161
 
162
+
163
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
164
 
165
+
166
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
167
 
168
+
169
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
170
 
171
+
172
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
173
 
174
+
175
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
176
 
177
+
178
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
179
 
180
+
181
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
182
 
183
+
184
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
185
 
186
+
187
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
188
 
189
+
190
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
191
 
192
+
193
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
194
 
195
+
196
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
197
 
198
+
199
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
200
 
201
+
202
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
203
 
204
+
205
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
206
 
207
+
208
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
209
 
210
+
211
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
212
 
213
+
214
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
215
 
216
+
217
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
218
 
219
+
220
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
221
 
222
+
223
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
224
 
225
+
226
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
227
 
228
+
229
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
230
 
231
+
232
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
233
 
234
+
235
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
236
 
237
+
238
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
239
 
240
+
241
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
242
 
243
+
244
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
245
 
246
+
247
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
248
 
249
+
250
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
251
 
252
+
253
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
254
 
255
+
256
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
257
 
258
+
259
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
260
 
261
+
262
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
263
 
264
+
265
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
266
 
267
+
268
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
269
 
270
+
271
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
272
 
273
+
274
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
275
 
276
+
277
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
278
 
279
+
280
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
281
 
282
+
283
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
284
 
285
+
286
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
287
 
288
+
289
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
290
 
291
+
292
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
293
 
294
+
295
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
296
 
297
+
298
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
299
 
300
+
301
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
302
 
303
+
304
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
305
 
306
+
307
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
308
 
309
+
310
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
311
 
312
+
313
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
314
 
315
+
316
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
317
 
318
+
319
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
320
 
321
+
322
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
323
 
324
+
325
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
326
 
327
+
328
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
329
 
330
+
331
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
332
 
333
+
334
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
335
 
336
+
337
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
338
 
339
+
340
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
341
 
342
+
343
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
344
 
345
+
346
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
347
 
348
+
349
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
350
 
351
+
352
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
353
 
354
+
355
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
356
 
357
+
358
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
359
 
360
+
361
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
362
 
363
+
364
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
365
 
366
+
367
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
368
 
369
+
370
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
371
 
372
+
373
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
374
 
375
+
376
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
377
 
378
+
379
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
380
 
381
+
382
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
383
 
384
+
385
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
386
 
387
+
388
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
389
 
390
+
391
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
392
 
393
+
394
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
395
 
396
+
397
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
398
 
399
+
400
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
401
 
402
+
403
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
404
 
405
+
406
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
407
 
408
+
409
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
410
 
411
+
412
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
413
 
414
+
415
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
416
 
417
+
418
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
419
 
420
+
421
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
422
 
423
+
424
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
425
 
426
+
427
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
428
 
429
+
430
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
431
 
432
+
433
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
434
 
435
+
436
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
437
 
438
+
439
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
440
 
441
+
442
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
443
 
444
+
445
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
446
 
447
+
448
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
449
 
450
+
451
  The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
452
 
453
+
454
+
455
+ 2026-07-18:14:41:29 INFO [loggers.evaluation_tracker:247] Saving results aggregated
456
+ 2026-07-18:14:41:29 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/vanilla/babilong8k_qa1_qa5_n50/__workspace/*.jsonl
457
  hf ({'pretrained': '/workspace', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 50.0, num_fewshot: 2, batch_size: 1
458
  | Tasks |Version|Filter|n-shot|Metric| |Value| |Stderr|
459
  |----------------|------:|------|-----:|------|---|----:|---|-----:|
outputs/eval/logs/vanilla/ruler_a.log CHANGED
@@ -1,69 +1,72 @@
1
- 2026-07-18:11:25:33 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
- 2026-07-18:11:25:37 INFO [_cli.run:376] Selected Tasks: ['niah_single_1', 'niah_single_3', 'niah_multikey_2', 'niah_multiquery', 'ruler_vt', 'ruler_fwe', 'ruler_qa_hotpot']
3
- 2026-07-18:11:25:39 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
- 2026-07-18:11:25:39 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
- 2026-07-18:11:25:42 INFO [models.huggingface:161] Using device 'cuda:0'
6
- 2026-07-18:11:25:42 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
7
- 2026-07-18:11:25:52 WARNING [api.task:856] niah_single_1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
8
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
9
- 2026-07-18:11:25:52 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace for synthetic tasks.
10
 
11
 
12
-
13
-
14
-
15
-
16
-
17
-
18
-
19
- 2026-07-18:11:26:00 WARNING [api.task:856] niah_single_3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
20
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
21
 
22
-
23
 
24
-
25
 
26
 
27
-
28
-
29
-
30
-
31
-
32
-
33
-
34
-
35
- 2026-07-18:11:26:10 WARNING [api.task:856] niah_multikey_2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
36
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
37
 
38
 
39
-
40
-
41
-
42
-
43
-
44
-
45
-
46
- 2026-07-18:11:26:17 WARNING [api.task:856] niah_multiquery: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
47
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
48
 
49
 
50
-
51
-
52
-
53
-
54
-
55
-
56
-
57
-
58
- 2026-07-18:11:26:26 WARNING [api.task:856] ruler_vt: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
59
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
60
- 2026-07-18:11:26:26 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace for synthetic tasks.
61
  Max length 500 | Current length 307 | Noises: 5
62
  Max length 500 | Current length 410 | Noises: 10
63
  Max length 500 | Current length 531 | Noises: 15
64
  Num noises: 10
65
 
66
-
67
  0%| | 0/1 [00:00<?, ?it/s]
 
68
  0%| | 0/1 [00:00<?, ?it/s]
69
  Max length 8192 | Current length 791 | Noises: 10
70
  Max length 8192 | Current length 1025 | Noises: 20
71
  Max length 8192 | Current length 1273 | Noises: 30
@@ -99,372 +102,376 @@ Max length 8192 | Current length 8230 | Noises: 320
99
  Num noises: 310
100
 
101
 
102
  0%| | 0/500 [00:00<?, ?it/s]
103
-
104
  13%|█▎ | 67/500 [00:01<00:06, 66.71it/s]
105
-
106
  27%|██▋ | 134/500 [00:02<00:05, 66.50it/s]
107
-
108
  41%|████ | 203/500 [00:03<00:04, 67.30it/s]
109
-
110
  54%|█████▍ | 271/500 [00:04<00:03, 67.23it/s]
111
-
112
  68%|██████▊ | 339/500 [00:05<00:02, 67.33it/s]
113
-
114
  81%|████████▏ | 407/500 [00:06<00:01, 64.51it/s]
115
-
116
  94%|█████████▍| 472/500 [00:07<00:00, 62.86it/s]
117
- 2026-07-18:11:26:34 WARNING [api.task:856] ruler_fwe: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
118
  12%|█▏ | 59/500 [00:01<00:07, 58.56it/s]
 
119
  24%|██▎ | 118/500 [00:02<00:06, 57.94it/s]
 
120
  35%|███▌ | 176/500 [00:03<00:05, 57.91it/s]
 
121
  47%|████▋ | 234/500 [00:04<00:04, 56.14it/s]
 
122
  58%|█████▊ | 291/500 [00:05<00:03, 55.52it/s]
 
123
  69%|██████▉ | 347/500 [00:06<00:02, 54.63it/s]
 
124
  80%|████████ | 402/500 [00:07<00:01, 54.48it/s]
 
125
  91%|█████████▏| 457/500 [00:08<00:00, 54.48it/s]
 
126
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
127
 
128
 
129
-
130
-
131
-
132
-
133
-
134
-
135
-
136
-
137
-
138
-
139
-
140
-
141
- 2026-07-18:11:26:47 WARNING [api.task:856] ruler_qa_hotpot: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
142
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
143
 
144
 
145
-
146
-
147
-
148
-
149
-
150
-
151
-
152
-
153
-
154
-
155
-
156
-
157
-
158
-
159
-
160
-
161
-
162
-
163
- 2026-07-18:11:27:09 INFO [tasks:700] Selected tasks:
164
- 2026-07-18:11:27:09 INFO [tasks:691] Task: ruler_qa_hotpot (ruler/qa_hotpot.yaml)
165
- 2026-07-18:11:27:09 INFO [tasks:691] Task: ruler_fwe (ruler/fwe.yaml)
166
- 2026-07-18:11:27:09 INFO [tasks:691] Task: ruler_vt (ruler/vt.yaml)
167
- 2026-07-18:11:27:09 INFO [tasks:691] Task: niah_multiquery (ruler/niah_multiquery.yaml)
168
- 2026-07-18:11:27:09 INFO [tasks:691] Task: niah_multikey_2 (ruler/niah_multikey_2.yaml)
169
- 2026-07-18:11:27:09 INFO [tasks:691] Task: niah_single_3 (ruler/niah_single_3.yaml)
170
- 2026-07-18:11:27:09 INFO [tasks:691] Task: niah_single_1 (ruler/niah_single_1.yaml)
171
- 2026-07-18:11:27:09 INFO [evaluator:314] ruler_qa_hotpot: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 32, 'until': []}
172
- 2026-07-18:11:27:09 INFO [evaluator:314] ruler_fwe: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 50, 'until': []}
173
- 2026-07-18:11:27:09 INFO [evaluator:314] ruler_vt: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 30, 'until': []}
174
- 2026-07-18:11:27:09 INFO [evaluator:314] niah_multiquery: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
175
- 2026-07-18:11:27:09 INFO [evaluator:314] niah_multikey_2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
176
- 2026-07-18:11:27:09 INFO [evaluator:314] niah_single_3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
177
- 2026-07-18:11:27:09 INFO [evaluator:314] niah_single_1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
178
- 2026-07-18:11:27:09 INFO [api.task:311] Building contexts for ruler_qa_hotpot on rank 0...
179
-
180
-
181
  0%| | 0/20 [00:00<?, ?it/s]
182
- 2026-07-18:11:27:09 INFO [api.task:311] Building contexts for ruler_fwe on rank 0...
183
-
184
-
185
  0%| | 0/20 [00:00<?, ?it/s]
186
- 2026-07-18:11:27:09 INFO [api.task:311] Building contexts for ruler_vt on rank 0...
187
-
188
-
189
  0%| | 0/20 [00:00<?, ?it/s]
190
- 2026-07-18:11:27:09 INFO [api.task:311] Building contexts for niah_multiquery on rank 0...
191
-
192
-
193
  0%| | 0/20 [00:00<?, ?it/s]
194
- 2026-07-18:11:27:10 INFO [api.task:311] Building contexts for niah_multikey_2 on rank 0...
195
-
196
-
197
  0%| | 0/20 [00:00<?, ?it/s]
198
- 2026-07-18:11:27:10 INFO [api.task:311] Building contexts for niah_single_3 on rank 0...
199
-
200
-
201
  0%| | 0/20 [00:00<?, ?it/s]
202
- 2026-07-18:11:27:10 INFO [api.task:311] Building contexts for niah_single_1 on rank 0...
203
-
204
-
205
  0%| | 0/20 [00:00<?, ?it/s]
206
- 2026-07-18:11:27:10 INFO [evaluator:584] Running generate_until requests
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
207
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
208
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
209
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
210
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
211
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
212
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
213
  0%| | 0/20 [00:00<?, ?it/s]
 
214
 
215
 
216
 
217
-
218
 
219
-
220
 
221
-
222
 
223
-
224
 
225
-
226
 
227
-
228
 
229
-
230
 
231
-
232
 
233
-
234
 
235
-
236
 
237
-
238
 
239
-
240
 
241
-
242
 
243
-
244
 
245
-
246
 
247
-
248
 
249
-
250
 
251
-
252
 
253
-
254
 
255
-
256
 
257
-
258
 
259
-
260
 
261
-
262
 
263
-
264
 
265
-
266
 
267
 
268
 
269
-
270
 
271
-
272
 
273
-
274
 
275
-
276
 
277
-
278
 
279
-
280
 
281
-
282
 
283
-
284
 
285
-
286
 
287
-
288
 
289
-
290
 
291
-
292
 
293
-
294
 
295
-
296
 
297
-
298
 
299
-
300
 
301
-
302
 
303
-
304
 
305
-
306
 
307
-
308
 
309
-
310
 
311
-
312
 
313
-
314
 
315
-
316
 
317
-
318
 
319
-
320
 
321
-
322
 
323
-
324
 
325
-
326
 
327
-
328
 
329
-
330
 
331
 
332
 
333
-
334
 
335
 
336
 
337
-
338
 
339
-
340
 
341
-
342
 
343
 
344
 
345
-
346
 
347
-
348
 
349
-
350
 
351
-
352
 
353
-
354
 
355
-
356
 
357
-
358
 
359
-
360
 
361
-
362
 
363
-
364
 
365
-
366
 
367
-
368
 
369
-
370
 
371
-
372
 
373
-
374
 
375
-
376
 
377
-
378
 
379
-
380
 
381
-
382
 
383
-
384
 
385
-
386
 
387
-
388
 
389
-
390
 
391
-
392
 
393
-
394
 
395
-
396
 
397
-
398
 
399
-
400
 
401
-
402
 
403
-
404
 
405
-
406
 
407
-
408
 
409
-
410
 
411
-
412
 
413
-
414
 
415
-
416
 
417
-
418
 
419
-
420
 
421
-
422
 
423
-
424
 
425
-
426
 
427
-
428
 
429
-
430
 
431
-
432
 
433
-
434
 
435
-
436
 
437
-
438
 
439
-
440
 
441
-
442
 
443
-
444
 
445
-
446
 
447
-
448
 
449
-
450
 
451
-
452
 
453
-
454
 
455
-
456
 
457
-
458
 
459
-
460
 
461
-
462
 
463
-
464
 
465
-
466
 
467
-
468
 
469
-
470
 
471
-
472
 
473
-
474
 
475
-
476
 
477
-
478
 
479
-
480
 
481
-
482
 
483
-
484
 
485
-
486
 
487
-
488
 
489
-
490
 
491
-
492
 
493
-
494
 
495
-
496
- 2026-07-18:11:34:54 INFO [loggers.evaluation_tracker:247] Saving results aggregated
497
- 2026-07-18:11:34:54 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/vanilla/ruler8k_splits/a/lm_eval/__workspace/*.jsonl
498
  hf ({'pretrained': '/workspace', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 20.0, num_fewshot: None, batch_size: 1
499
  | Tasks |Version|Filter|n-shot|Metric| | Value | |Stderr|
500
  |---------------|------:|------|-----:|-----:|---|------:|---|------|
 
1
+ 2026-07-18:14:17:11 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
+ 2026-07-18:14:17:15 INFO [_cli.run:376] Selected Tasks: ['niah_single_1', 'niah_single_3', 'niah_multikey_2', 'niah_multiquery', 'ruler_vt', 'ruler_fwe', 'ruler_qa_hotpot']
3
+ 2026-07-18:14:17:16 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
+ 2026-07-18:14:17:16 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
+ 2026-07-18:14:17:20 INFO [models.huggingface:161] Using device 'cuda:0'
6
+ 2026-07-18:14:17:20 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
7
+ 2026-07-18:14:17:29 WARNING [api.task:856] niah_single_1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
8
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
9
+ 2026-07-18:14:17:29 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace for synthetic tasks.
10
 
11
 
12
+
13
+
14
+
15
+
16
+
17
+
18
+
19
+
20
+ 2026-07-18:14:17:39 WARNING [api.task:856] niah_single_3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
21
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
22
 
23
+
24
 
25
+
26
 
27
 
28
+
29
+
30
+
31
+
32
+
33
+
34
+
35
+
36
+
37
+ 2026-07-18:14:17:50 WARNING [api.task:856] niah_multikey_2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
38
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
39
 
40
 
41
+
42
+
43
+
44
+
45
+
46
+
47
+
48
+ 2026-07-18:14:17:58 WARNING [api.task:856] niah_multiquery: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
49
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
50
 
51
 
52
+
53
+
54
+
55
+
56
+
57
+
58
+
59
+
60
+
61
+ 2026-07-18:14:18:08 WARNING [api.task:856] ruler_vt: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
62
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
63
+ 2026-07-18:14:18:08 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace for synthetic tasks.
64
  Max length 500 | Current length 307 | Noises: 5
65
  Max length 500 | Current length 410 | Noises: 10
66
  Max length 500 | Current length 531 | Noises: 15
67
  Num noises: 10
68
 
 
69
  0%| | 0/1 [00:00<?, ?it/s]
70
+
71
  0%| | 0/1 [00:00<?, ?it/s]
72
  Max length 8192 | Current length 791 | Noises: 10
73
  Max length 8192 | Current length 1025 | Noises: 20
74
  Max length 8192 | Current length 1273 | Noises: 30
 
102
  Num noises: 310
103
 
104
 
105
  0%| | 0/500 [00:00<?, ?it/s]
 
106
  13%|█▎ | 67/500 [00:01<00:06, 66.71it/s]
 
107
  27%|██▋ | 134/500 [00:02<00:05, 66.50it/s]
 
108
  41%|████ | 203/500 [00:03<00:04, 67.30it/s]
 
109
  54%|█████▍ | 271/500 [00:04<00:03, 67.23it/s]
 
110
  68%|██████▊ | 339/500 [00:05<00:02, 67.33it/s]
 
111
  81%|████████▏ | 407/500 [00:06<00:01, 64.51it/s]
 
112
  94%|█████████▍| 472/500 [00:07<00:00, 62.86it/s]
113
+
114
  12%|█▏ | 59/500 [00:01<00:07, 58.56it/s]
115
+
116
  24%|██▎ | 118/500 [00:02<00:06, 57.94it/s]
117
+
118
  35%|███▌ | 176/500 [00:03<00:05, 57.91it/s]
119
+
120
  47%|████▋ | 234/500 [00:04<00:04, 56.14it/s]
121
+
122
  58%|█████▊ | 291/500 [00:05<00:03, 55.52it/s]
123
+
124
  69%|██████▉ | 347/500 [00:06<00:02, 54.63it/s]
125
+
126
  80%|████████ | 402/500 [00:07<00:01, 54.48it/s]
127
+
128
  91%|█████████▏| 457/500 [00:08<00:00, 54.48it/s]
129
+ 2026-07-18:14:18:18 WARNING [api.task:856] ruler_fwe: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
130
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
131
 
132
 
133
+
134
+
135
+
136
+
137
+
138
+
139
+
140
+
141
+
142
+
143
+
144
+
145
+
146
+ 2026-07-18:14:18:32 WARNING [api.task:856] ruler_qa_hotpot: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
147
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
148
 
149
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
150
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
151
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
152
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
153
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
154
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
155
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
156
  0%| | 0/20 [00:00<?, ?it/s]
157
+
158
+
159
+
160
+
161
+
162
+
163
+
164
+
165
+
166
+
167
+
168
+
169
+
170
+
171
+
172
+
173
+
174
+
175
+
176
+
177
+ 2026-07-18:14:18:55 INFO [tasks:700] Selected tasks:
178
+ 2026-07-18:14:18:55 INFO [tasks:691] Task: ruler_qa_hotpot (ruler/qa_hotpot.yaml)
179
+ 2026-07-18:14:18:55 INFO [tasks:691] Task: ruler_fwe (ruler/fwe.yaml)
180
+ 2026-07-18:14:18:55 INFO [tasks:691] Task: ruler_vt (ruler/vt.yaml)
181
+ 2026-07-18:14:18:55 INFO [tasks:691] Task: niah_multiquery (ruler/niah_multiquery.yaml)
182
+ 2026-07-18:14:18:55 INFO [tasks:691] Task: niah_multikey_2 (ruler/niah_multikey_2.yaml)
183
+ 2026-07-18:14:18:55 INFO [tasks:691] Task: niah_single_3 (ruler/niah_single_3.yaml)
184
+ 2026-07-18:14:18:55 INFO [tasks:691] Task: niah_single_1 (ruler/niah_single_1.yaml)
185
+ 2026-07-18:14:18:55 INFO [evaluator:314] ruler_qa_hotpot: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 32, 'until': []}
186
+ 2026-07-18:14:18:55 INFO [evaluator:314] ruler_fwe: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 50, 'until': []}
187
+ 2026-07-18:14:18:55 INFO [evaluator:314] ruler_vt: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 30, 'until': []}
188
+ 2026-07-18:14:18:55 INFO [evaluator:314] niah_multiquery: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
189
+ 2026-07-18:14:18:55 INFO [evaluator:314] niah_multikey_2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
190
+ 2026-07-18:14:18:55 INFO [evaluator:314] niah_single_3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
191
+ 2026-07-18:14:18:55 INFO [evaluator:314] niah_single_1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
192
+ 2026-07-18:14:18:55 INFO [api.task:311] Building contexts for ruler_qa_hotpot on rank 0...
193
+
194
+
195
  0%| | 0/20 [00:00<?, ?it/s]
196
+ 2026-07-18:14:18:55 INFO [api.task:311] Building contexts for ruler_fwe on rank 0...
197
+
198
+
199
  0%| | 0/20 [00:00<?, ?it/s]
200
+ 2026-07-18:14:18:56 INFO [api.task:311] Building contexts for ruler_vt on rank 0...
201
+
202
+
203
  0%| | 0/20 [00:00<?, ?it/s]
204
+ 2026-07-18:14:18:56 INFO [api.task:311] Building contexts for niah_multiquery on rank 0...
205
+
206
+
207
  0%| | 0/20 [00:00<?, ?it/s]
208
+ 2026-07-18:14:18:56 INFO [api.task:311] Building contexts for niah_multikey_2 on rank 0...
209
+
210
+
211
  0%| | 0/20 [00:00<?, ?it/s]
212
+ 2026-07-18:14:18:56 INFO [api.task:311] Building contexts for niah_single_3 on rank 0...
213
+
214
+
215
  0%| | 0/20 [00:00<?, ?it/s]
216
+ 2026-07-18:14:18:56 INFO [api.task:311] Building contexts for niah_single_1 on rank 0...
217
+
218
+
219
  0%| | 0/20 [00:00<?, ?it/s]
220
+ 2026-07-18:14:18:56 INFO [evaluator:584] Running generate_until requests
221
 
222
 
223
 
224
+
225
 
226
+
227
 
228
+
229
 
230
+
231
 
232
+
233
 
234
+
235
 
236
+
237
 
238
+
239
 
240
+
241
 
242
+
243
 
244
+
245
 
246
+
247
 
248
+
249
 
250
+
251
 
252
+
253
 
254
+
255
 
256
+
257
 
258
+
259
 
260
+
261
 
262
+
263
 
264
+
265
 
266
+
267
 
268
+
269
 
270
+
271
 
272
+
273
 
274
 
275
 
276
+
277
 
278
+
279
 
280
+
281
 
282
+
283
 
284
+
285
 
286
+
287
 
288
+
289
 
290
+
291
 
292
+
293
 
294
+
295
 
296
+
297
 
298
+
299
 
300
+
301
 
302
+
303
 
304
+
305
 
306
+
307
 
308
+
309
 
310
+
311
 
312
+
313
 
314
+
315
 
316
+
317
 
318
+
319
 
320
+
321
 
322
+
323
 
324
+
325
 
326
+
327
 
328
+
329
 
330
+
331
 
332
+
333
 
334
+
335
 
336
+
337
 
338
 
339
 
340
+
341
 
342
 
343
 
344
+
345
 
346
+
347
 
348
+
349
 
350
 
351
 
352
+
353
 
354
+
355
 
356
+
357
 
358
+
359
 
360
+
361
 
362
+
363
 
364
+
365
 
366
+
367
 
368
+
369
 
370
+
371
 
372
+
373
 
374
+
375
 
376
+
377
 
378
+
379
 
380
+
381
 
382
+
383
 
384
+
385
 
386
+
387
 
388
+
389
 
390
+
391
 
392
+
393
 
394
+
395
 
396
+
397
 
398
+
399
 
400
+
401
 
402
+
403
 
404
+
405
 
406
+
407
 
408
+
409
 
410
+
411
 
412
+
413
 
414
+
415
 
416
+
417
 
418
+
419
 
420
+
421
 
422
+
423
 
424
+
425
 
426
+
427
 
428
+
429
 
430
+
431
 
432
+
433
 
434
+
435
 
436
+
437
 
438
+
439
 
440
+
441
 
442
+
443
 
444
+
445
 
446
+
447
 
448
+
449
 
450
+
451
 
452
+
453
 
454
+
455
 
456
+
457
 
458
+
459
 
460
+
461
 
462
+
463
 
464
+
465
 
466
+
467
 
468
+
469
 
470
+
471
 
472
+
473
 
474
+
475
 
476
+
477
 
478
+
479
 
480
+
481
 
482
+
483
 
484
+
485
 
486
+
487
 
488
+
489
 
490
+
491
 
492
+
493
 
494
+
495
 
496
+
497
 
498
+
499
 
500
+
501
 
502
+
503
+ 2026-07-18:14:27:50 INFO [loggers.evaluation_tracker:247] Saving results aggregated
504
+ 2026-07-18:14:27:50 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/vanilla/ruler8k_splits/a/lm_eval/__workspace/*.jsonl
505
  hf ({'pretrained': '/workspace', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 20.0, num_fewshot: None, batch_size: 1
506
  | Tasks |Version|Filter|n-shot|Metric| | Value | |Stderr|
507
  |---------------|------:|------|-----:|-----:|---|------:|---|------|
outputs/eval/logs/vanilla/ruler_b.log CHANGED
@@ -1,353 +1,356 @@
1
- 2026-07-18:11:34:57 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
- 2026-07-18:11:35:00 INFO [_cli.run:376] Selected Tasks: ['niah_single_2', 'niah_multikey_1', 'niah_multikey_3', 'niah_multivalue', 'ruler_cwe', 'ruler_qa_squad']
3
- 2026-07-18:11:35:02 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
- 2026-07-18:11:35:02 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
- 2026-07-18:11:35:05 INFO [models.huggingface:161] Using device 'cuda:0'
6
- 2026-07-18:11:35:05 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
7
- 2026-07-18:11:35:14 WARNING [api.task:856] niah_single_2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
8
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
9
- 2026-07-18:11:35:15 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace for synthetic tasks.
10
 
11
 
12
-
13
-
14
-
15
-
16
-
17
-
18
-
19
-
20
- 2026-07-18:11:35:25 WARNING [api.task:856] niah_multikey_1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
21
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
22
 
23
 
24
-
25
-
26
-
27
-
28
-
29
-
30
-
31
-
32
- 2026-07-18:11:35:33 WARNING [api.task:856] niah_multikey_3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
33
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
34
 
35
 
36
-
37
-
38
-
39
-
40
-
41
- 2026-07-18:11:35:39 WARNING [api.task:856] niah_multivalue: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
42
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
43
 
44
 
45
-
46
-
47
-
48
-
49
-
50
-
51
-
52
-
53
- 2026-07-18:11:35:49 WARNING [api.task:856] ruler_cwe: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
54
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
55
- 2026-07-18:11:35:49 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace for synthetic tasks.
56
 
57
 
58
-
59
-
60
-
61
-
62
-
63
- 2026-07-18:11:35:55 WARNING [api.task:856] ruler_qa_squad: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
64
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
65
 
66
 
67
-
68
-
69
-
70
-
71
-
72
-
73
-
74
-
75
- 2026-07-18:11:36:05 INFO [tasks:700] Selected tasks:
76
- 2026-07-18:11:36:05 INFO [tasks:691] Task: ruler_qa_squad (ruler/qa_squad.yaml)
77
- 2026-07-18:11:36:05 INFO [tasks:691] Task: ruler_cwe (ruler/cwe.yaml)
78
- 2026-07-18:11:36:05 INFO [tasks:691] Task: niah_multivalue (ruler/niah_multivalue.yaml)
79
- 2026-07-18:11:36:05 INFO [tasks:691] Task: niah_multikey_3 (ruler/niah_multikey_3.yaml)
80
- 2026-07-18:11:36:05 INFO [tasks:691] Task: niah_multikey_1 (ruler/niah_multikey_1.yaml)
81
- 2026-07-18:11:36:05 INFO [tasks:691] Task: niah_single_2 (ruler/niah_single_2.yaml)
82
- 2026-07-18:11:36:05 INFO [evaluator:314] ruler_qa_squad: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 32, 'until': []}
83
- 2026-07-18:11:36:05 INFO [evaluator:314] ruler_cwe: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 120, 'until': []}
84
- 2026-07-18:11:36:05 INFO [evaluator:314] niah_multivalue: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
85
- 2026-07-18:11:36:05 INFO [evaluator:314] niah_multikey_3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
86
- 2026-07-18:11:36:05 INFO [evaluator:314] niah_multikey_1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
87
- 2026-07-18:11:36:05 INFO [evaluator:314] niah_single_2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
88
- 2026-07-18:11:36:05 INFO [api.task:311] Building contexts for ruler_qa_squad on rank 0...
89
-
90
-
91
  0%| | 0/20 [00:00<?, ?it/s]
92
- 2026-07-18:11:36:05 INFO [api.task:311] Building contexts for ruler_cwe on rank 0...
93
-
94
-
95
  0%| | 0/20 [00:00<?, ?it/s]
96
- 2026-07-18:11:36:05 INFO [api.task:311] Building contexts for niah_multivalue on rank 0...
97
-
98
-
99
  0%| | 0/20 [00:00<?, ?it/s]
100
- 2026-07-18:11:36:05 INFO [api.task:311] Building contexts for niah_multikey_3 on rank 0...
101
-
102
-
103
  0%| | 0/20 [00:00<?, ?it/s]
104
- 2026-07-18:11:36:05 INFO [api.task:311] Building contexts for niah_multikey_1 on rank 0...
105
-
106
-
107
  0%| | 0/20 [00:00<?, ?it/s]
108
- 2026-07-18:11:36:06 INFO [api.task:311] Building contexts for niah_single_2 on rank 0...
109
-
110
-
111
  0%| | 0/20 [00:00<?, ?it/s]
112
- 2026-07-18:11:36:06 INFO [evaluator:584] Running generate_until requests
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
113
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
114
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
115
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
116
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
117
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
118
  0%| | 0/20 [00:00<?, ?it/s]
 
119
 
120
 
121
 
122
-
123
 
124
-
125
 
126
-
127
 
128
-
129
 
130
-
131
 
132
-
133
 
134
-
135
 
136
-
137
 
138
-
139
 
140
-
141
 
142
-
143
 
144
-
145
 
146
-
147
 
148
-
149
 
150
-
151
 
152
-
153
 
154
-
155
 
156
-
157
 
158
-
159
 
160
-
161
 
162
-
163
 
164
-
165
 
166
-
167
 
168
-
169
 
170
-
171
 
172
-
173
 
174
-
175
 
176
-
177
 
178
-
179
 
180
-
181
 
182
-
183
 
184
-
185
 
186
-
187
 
188
-
189
 
190
-
191
 
192
-
193
 
194
-
195
 
196
-
197
 
198
-
199
 
200
-
201
 
202
-
203
 
204
-
205
 
206
-
207
 
208
-
209
 
210
-
211
 
212
-
213
 
214
-
215
 
216
-
217
 
218
-
219
 
220
-
221
 
222
-
223
 
224
-
225
 
226
-
227
 
228
-
229
 
230
-
231
 
232
-
233
 
234
-
235
 
236
-
237
 
238
-
239
 
240
-
241
 
242
-
243
 
244
-
245
 
246
-
247
 
248
-
249
 
250
-
251
 
252
-
253
 
254
-
255
 
256
-
257
 
258
-
259
 
260
-
261
 
262
-
263
 
264
-
265
 
266
-
267
 
268
-
269
 
270
-
271
 
272
-
273
 
274
-
275
 
276
-
277
 
278
-
279
 
280
-
281
 
282
-
283
 
284
-
285
 
286
-
287
 
288
-
289
 
290
-
291
 
292
-
293
 
294
-
295
 
296
-
297
 
298
-
299
 
300
-
301
 
302
-
303
 
304
-
305
 
306
-
307
 
308
-
309
 
310
-
311
 
312
-
313
 
314
-
315
 
316
-
317
 
318
-
319
 
320
-
321
 
322
-
323
 
324
-
325
 
326
-
327
 
328
-
329
 
330
-
331
 
332
-
333
 
334
-
335
 
336
-
337
 
338
-
339
 
340
-
341
 
342
-
343
 
344
-
345
 
346
-
347
 
348
-
349
 
350
-
351
 
352
-
353
 
354
-
355
 
356
-
357
 
358
-
359
 
360
-
361
- 2026-07-18:11:45:58 INFO [loggers.evaluation_tracker:247] Saving results aggregated
362
- 2026-07-18:11:45:58 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/vanilla/ruler8k_splits/b/lm_eval/__workspace/*.jsonl
363
  hf ({'pretrained': '/workspace', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 20.0, num_fewshot: None, batch_size: 1
364
  | Tasks |Version|Filter|n-shot|Metric| | Value | |Stderr|
365
  |---------------|------:|------|-----:|-----:|---|------:|---|------|
 
1
+ 2026-07-18:14:27:53 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
+ 2026-07-18:14:27:57 INFO [_cli.run:376] Selected Tasks: ['niah_single_2', 'niah_multikey_1', 'niah_multikey_3', 'niah_multivalue', 'ruler_cwe', 'ruler_qa_squad']
3
+ 2026-07-18:14:27:58 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
+ 2026-07-18:14:27:58 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
+ 2026-07-18:14:28:01 INFO [models.huggingface:161] Using device 'cuda:0'
6
+ 2026-07-18:14:28:02 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
7
+ 2026-07-18:14:28:11 WARNING [api.task:856] niah_single_2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
8
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
9
+ 2026-07-18:14:28:11 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace for synthetic tasks.
10
 
11
 
12
+
13
+
14
+
15
+
16
+
17
+
18
+
19
+
20
+ 2026-07-18:14:28:21 WARNING [api.task:856] niah_multikey_1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
21
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
22
 
23
 
24
+
25
+
26
+
27
+
28
+
29
+
30
+
31
+
32
+ 2026-07-18:14:28:30 WARNING [api.task:856] niah_multikey_3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
33
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
34
 
35
 
36
+
37
+
38
+
39
+
40
+
41
+ 2026-07-18:14:28:36 WARNING [api.task:856] niah_multivalue: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
42
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
43
 
44
 
45
+
46
+
47
+
48
+
49
+
50
+
51
+
52
+
53
+ 2026-07-18:14:28:46 WARNING [api.task:856] ruler_cwe: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
54
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
55
+ 2026-07-18:14:28:46 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace for synthetic tasks.
56
 
57
 
58
+
59
+
60
+
61
+
62
+
63
+
64
+ 2026-07-18:14:28:53 WARNING [api.task:856] ruler_qa_squad: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
65
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
66
 
67
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
68
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
69
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
70
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
71
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
72
  0%| | 0/20 [00:00<?, ?it/s]
 
 
 
73
  0%| | 0/20 [00:00<?, ?it/s]
74
+
75
+
76
+
77
+
78
+
79
+
80
+
81
+
82
+
83
+
84
+ 2026-07-18:14:29:05 INFO [tasks:700] Selected tasks:
85
+ 2026-07-18:14:29:05 INFO [tasks:691] Task: ruler_qa_squad (ruler/qa_squad.yaml)
86
+ 2026-07-18:14:29:05 INFO [tasks:691] Task: ruler_cwe (ruler/cwe.yaml)
87
+ 2026-07-18:14:29:05 INFO [tasks:691] Task: niah_multivalue (ruler/niah_multivalue.yaml)
88
+ 2026-07-18:14:29:05 INFO [tasks:691] Task: niah_multikey_3 (ruler/niah_multikey_3.yaml)
89
+ 2026-07-18:14:29:05 INFO [tasks:691] Task: niah_multikey_1 (ruler/niah_multikey_1.yaml)
90
+ 2026-07-18:14:29:05 INFO [tasks:691] Task: niah_single_2 (ruler/niah_single_2.yaml)
91
+ 2026-07-18:14:29:05 INFO [evaluator:314] ruler_qa_squad: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 32, 'until': []}
92
+ 2026-07-18:14:29:05 INFO [evaluator:314] ruler_cwe: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 120, 'until': []}
93
+ 2026-07-18:14:29:05 INFO [evaluator:314] niah_multivalue: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
94
+ 2026-07-18:14:29:05 INFO [evaluator:314] niah_multikey_3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
95
+ 2026-07-18:14:29:05 INFO [evaluator:314] niah_multikey_1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
96
+ 2026-07-18:14:29:05 INFO [evaluator:314] niah_single_2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
97
+ 2026-07-18:14:29:05 INFO [api.task:311] Building contexts for ruler_qa_squad on rank 0...
98
+
99
+
100
  0%| | 0/20 [00:00<?, ?it/s]
101
+ 2026-07-18:14:29:05 INFO [api.task:311] Building contexts for ruler_cwe on rank 0...
102
+
103
+
104
  0%| | 0/20 [00:00<?, ?it/s]
105
+ 2026-07-18:14:29:05 INFO [api.task:311] Building contexts for niah_multivalue on rank 0...
106
+
107
+
108
  0%| | 0/20 [00:00<?, ?it/s]
109
+ 2026-07-18:14:29:05 INFO [api.task:311] Building contexts for niah_multikey_3 on rank 0...
110
+
111
+
112
  0%| | 0/20 [00:00<?, ?it/s]
113
+ 2026-07-18:14:29:05 INFO [api.task:311] Building contexts for niah_multikey_1 on rank 0...
114
+
115
+
116
  0%| | 0/20 [00:00<?, ?it/s]
117
+ 2026-07-18:14:29:05 INFO [api.task:311] Building contexts for niah_single_2 on rank 0...
118
+
119
+
120
  0%| | 0/20 [00:00<?, ?it/s]
121
+ 2026-07-18:14:29:05 INFO [evaluator:584] Running generate_until requests
122
 
123
 
124
 
125
+
126
 
127
+
128
 
129
+
130
 
131
+
132
 
133
+
134
 
135
+
136
 
137
+
138
 
139
+
140
 
141
+
142
 
143
+
144
 
145
+
146
 
147
+
148
 
149
+
150
 
151
+
152
 
153
+
154
 
155
+
156
 
157
+
158
 
159
+
160
 
161
+
162
 
163
+
164
 
165
+
166
 
167
+
168
 
169
+
170
 
171
+
172
 
173
+
174
 
175
+
176
 
177
+
178
 
179
+
180
 
181
+
182
 
183
+
184
 
185
+
186
 
187
+
188
 
189
+
190
 
191
+
192
 
193
+
194
 
195
+
196
 
197
+
198
 
199
+
200
 
201
+
202
 
203
+
204
 
205
+
206
 
207
+
208
 
209
+
210
 
211
+
212
 
213
+
214
 
215
+
216
 
217
+
218
 
219
+
220
 
221
+
222
 
223
+
224
 
225
+
226
 
227
+
228
 
229
+
230
 
231
+
232
 
233
+
234
 
235
+
236
 
237
+
238
 
239
+
240
 
241
+
242
 
243
+
244
 
245
+
246
 
247
+
248
 
249
+
250
 
251
+
252
 
253
+
254
 
255
+
256
 
257
+
258
 
259
+
260
 
261
+
262
 
263
+
264
 
265
+
266
 
267
+
268
 
269
+
270
 
271
+
272
 
273
+
274
 
275
+
276
 
277
+
278
 
279
+
280
 
281
+
282
 
283
+
284
 
285
+
286
 
287
+
288
 
289
+
290
 
291
+
292
 
293
+
294
 
295
+
296
 
297
+
298
 
299
+
300
 
301
+
302
 
303
+
304
 
305
+
306
 
307
+
308
 
309
+
310
 
311
+
312
 
313
+
314
 
315
+
316
 
317
+
318
 
319
+
320
 
321
+
322
 
323
+
324
 
325
+
326
 
327
+
328
 
329
+
330
 
331
+
332
 
333
+
334
 
335
+
336
 
337
+
338
 
339
+
340
 
341
+
342
 
343
+
344
 
345
+
346
 
347
+
348
 
349
+
350
 
351
+
352
 
353
+
354
 
355
+
356
 
357
+
358
 
359
+
360
 
361
+
362
 
363
+
364
+ 2026-07-18:14:38:04 INFO [loggers.evaluation_tracker:247] Saving results aggregated
365
+ 2026-07-18:14:38:04 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/vanilla/ruler8k_splits/b/lm_eval/__workspace/*.jsonl
366
  hf ({'pretrained': '/workspace', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 20.0, num_fewshot: None, batch_size: 1
367
  | Tasks |Version|Filter|n-shot|Metric| | Value | |Stderr|
368
  |---------------|------:|------|-----:|-----:|---|------:|---|------|
outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/results_2026-07-18T15-23-38.141056.json ADDED
@@ -0,0 +1,505 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "results": {
3
+ "babilong_longctx": {
4
+ "acc,none": 0.588,
5
+ "acc_stderr,none": 0.029807545955988512,
6
+ "alias": "babilong_longctx"
7
+ },
8
+ "babilong_qa1": {
9
+ "alias": " - babilong_qa1",
10
+ "acc,none": 0.64,
11
+ "acc_stderr,none": 0.06857142857142856
12
+ },
13
+ "babilong_qa2": {
14
+ "alias": " - babilong_qa2",
15
+ "acc,none": 0.48,
16
+ "acc_stderr,none": 0.07137140569598172
17
+ },
18
+ "babilong_qa3": {
19
+ "alias": " - babilong_qa3",
20
+ "acc,none": 0.34,
21
+ "acc_stderr,none": 0.06767268161329718
22
+ },
23
+ "babilong_qa4": {
24
+ "alias": " - babilong_qa4",
25
+ "acc,none": 0.72,
26
+ "acc_stderr,none": 0.06414269805898185
27
+ },
28
+ "babilong_qa5": {
29
+ "alias": " - babilong_qa5",
30
+ "acc,none": 0.76,
31
+ "acc_stderr,none": 0.06101187572589322
32
+ }
33
+ },
34
+ "groups": {
35
+ "babilong_longctx": {
36
+ "acc,none": 0.588,
37
+ "acc_stderr,none": 0.029807545955988512,
38
+ "alias": "babilong_longctx"
39
+ }
40
+ },
41
+ "group_subtasks": {
42
+ "babilong_longctx": [
43
+ "babilong_qa1",
44
+ "babilong_qa2",
45
+ "babilong_qa3",
46
+ "babilong_qa4",
47
+ "babilong_qa5"
48
+ ]
49
+ },
50
+ "configs": {
51
+ "babilong_qa1": {
52
+ "task": "babilong_qa1",
53
+ "custom_dataset": "def load_dataset(**kwargs):\n config_name = kwargs.get(\"max_seq_lengths\", \"0k\")\n\n # Get specific qa split\n qa_split = kwargs.get(\"qa_split\")\n\n eval_logger.info(\n f\"Loading babilong dataset: max_seq_lengths={config_name}, split={qa_split}\"\n )\n dataset = datasets.load_dataset(\n \"RMT-team/babilong-1k-samples\", name=config_name, split=qa_split\n )\n return {qa_split: dataset}\n",
54
+ "dataset_path": "RMT-team/babilong-1k-samples",
55
+ "dataset_kwargs": {
56
+ "qa_split": "qa1"
57
+ },
58
+ "test_split": "qa1",
59
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
60
+ "doc_to_target": "{{target}}",
61
+ "unsafe_code": false,
62
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n pred = postprocess_pred(results)\n target = doc.get(\"target\", \"\").strip()\n\n # String match\n score = 1.0 if target.lower() in pred[0].lower() else 0.0\n\n return {\"acc\": score}\n",
63
+ "description": "I will give you context with the facts about positions of different persons hidden in some random text and a question. You need to answer the question based only on the information from the facts. If a person was in different locations, use the latest location to answer the question.\nAlways return your answer in the following format:\nThe most recent location of 'person' is 'location'. Do not write anything else after that.\n\n",
64
+ "target_delimiter": " ",
65
+ "fewshot_delimiter": "\n\n",
66
+ "fewshot_config": {
67
+ "sampler": "first_n",
68
+ "split": null,
69
+ "process_docs": null,
70
+ "fewshot_indices": null,
71
+ "samples": [
72
+ {
73
+ "input": "Charlie went to the hallway. Judith come back to the kitchen. Charlie travelled to balcony.",
74
+ "question": "Where is Charlie?",
75
+ "target": "The most recent location of Charlie is balcony."
76
+ },
77
+ {
78
+ "input": "Alan moved to the garage. Charlie went to the beach. Alan went to the shop. Rouse travelled to balcony.",
79
+ "question": "Where is Alan?",
80
+ "target": "The most recent location of Alan is shop."
81
+ }
82
+ ],
83
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
84
+ "doc_to_choice": null,
85
+ "doc_to_target": "{{target}}",
86
+ "gen_prefix": null,
87
+ "fewshot_delimiter": "\n\n",
88
+ "target_delimiter": " "
89
+ },
90
+ "num_fewshot": 2,
91
+ "metric_list": [
92
+ {
93
+ "metric": "acc",
94
+ "aggregation": "mean",
95
+ "higher_is_better": true
96
+ }
97
+ ],
98
+ "output_type": "generate_until",
99
+ "generation_kwargs": {
100
+ "do_sample": false,
101
+ "temperature": 0.0,
102
+ "max_gen_toks": 16,
103
+ "until": []
104
+ },
105
+ "repeats": 1,
106
+ "should_decontaminate": false,
107
+ "metadata": {
108
+ "version": 0.0,
109
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
110
+ "trust_remote_code": true,
111
+ "dtype": "bfloat16",
112
+ "max_length": 16384,
113
+ "attn_implementation": "sdpa",
114
+ "max_seq_lengths": "8k"
115
+ }
116
+ },
117
+ "babilong_qa2": {
118
+ "task": "babilong_qa2",
119
+ "custom_dataset": "def load_dataset(**kwargs):\n config_name = kwargs.get(\"max_seq_lengths\", \"0k\")\n\n # Get specific qa split\n qa_split = kwargs.get(\"qa_split\")\n\n eval_logger.info(\n f\"Loading babilong dataset: max_seq_lengths={config_name}, split={qa_split}\"\n )\n dataset = datasets.load_dataset(\n \"RMT-team/babilong-1k-samples\", name=config_name, split=qa_split\n )\n return {qa_split: dataset}\n",
120
+ "dataset_path": "RMT-team/babilong-1k-samples",
121
+ "dataset_kwargs": {
122
+ "qa_split": "qa2"
123
+ },
124
+ "test_split": "qa2",
125
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
126
+ "doc_to_target": "{{target}}",
127
+ "unsafe_code": false,
128
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n pred = postprocess_pred(results)\n target = doc.get(\"target\", \"\").strip()\n\n # String match\n score = 1.0 if target.lower() in pred[0].lower() else 0.0\n\n return {\"acc\": score}\n",
129
+ "description": "I will give you context with the facts about locations and actions of different persons hidden in some random text and a question. You need to answer the question based only on the information from the facts. If a person got an item in the first location and travelled to the second location the item is also in the second location. If a person dropped an item in the first location and moved to the second location the item remains in the first location.\nAlways return your answer in the following format:\nThe 'item' is in 'location'. Do not write anything else after that.\n\n",
130
+ "target_delimiter": " ",
131
+ "fewshot_delimiter": "\n\n",
132
+ "fewshot_config": {
133
+ "sampler": "first_n",
134
+ "split": null,
135
+ "process_docs": null,
136
+ "fewshot_indices": null,
137
+ "samples": [
138
+ {
139
+ "input": "Charlie went to the kitchen. Charlie got a bottle. Charlie moved to the balcony.",
140
+ "question": "Where is the bottle?",
141
+ "target": "The bottle is in the balcony."
142
+ },
143
+ {
144
+ "input": "Alan moved to the garage. Alan got a screw driver. Alan moved to the kitchen.",
145
+ "question": "Where is the screw driver?",
146
+ "target": "The screw driver is in the kitchen."
147
+ }
148
+ ],
149
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
150
+ "doc_to_choice": null,
151
+ "doc_to_target": "{{target}}",
152
+ "gen_prefix": null,
153
+ "fewshot_delimiter": "\n\n",
154
+ "target_delimiter": " "
155
+ },
156
+ "num_fewshot": 2,
157
+ "metric_list": [
158
+ {
159
+ "metric": "acc",
160
+ "aggregation": "mean",
161
+ "higher_is_better": true
162
+ }
163
+ ],
164
+ "output_type": "generate_until",
165
+ "generation_kwargs": {
166
+ "do_sample": false,
167
+ "temperature": 0.0,
168
+ "max_gen_toks": 16,
169
+ "until": []
170
+ },
171
+ "repeats": 1,
172
+ "should_decontaminate": false,
173
+ "metadata": {
174
+ "version": 0.0,
175
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
176
+ "trust_remote_code": true,
177
+ "dtype": "bfloat16",
178
+ "max_length": 16384,
179
+ "attn_implementation": "sdpa",
180
+ "max_seq_lengths": "8k"
181
+ }
182
+ },
183
+ "babilong_qa3": {
184
+ "task": "babilong_qa3",
185
+ "custom_dataset": "def load_dataset(**kwargs):\n config_name = kwargs.get(\"max_seq_lengths\", \"0k\")\n\n # Get specific qa split\n qa_split = kwargs.get(\"qa_split\")\n\n eval_logger.info(\n f\"Loading babilong dataset: max_seq_lengths={config_name}, split={qa_split}\"\n )\n dataset = datasets.load_dataset(\n \"RMT-team/babilong-1k-samples\", name=config_name, split=qa_split\n )\n return {qa_split: dataset}\n",
186
+ "dataset_path": "RMT-team/babilong-1k-samples",
187
+ "dataset_kwargs": {
188
+ "qa_split": "qa3"
189
+ },
190
+ "test_split": "qa3",
191
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
192
+ "doc_to_target": "{{target}}",
193
+ "unsafe_code": false,
194
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n pred = postprocess_pred(results)\n target = doc.get(\"target\", \"\").strip()\n\n # String match\n score = 1.0 if target.lower() in pred[0].lower() else 0.0\n\n return {\"acc\": score}\n",
195
+ "description": "I give you context with the facts about locations and actions of different persons hidden in some random text and a question. You need to answer the question based only on the information from the facts. If a person got an item in the first location and travelled to the second location the item is also in the second location. If a person dropped an item in the first location and moved to the second location the item remains in the first location.\nAlways return your answer in the following format:\nBefore the $location_1$ the $item$ was in the $location_2$. Do not write anything else after that.\n\n",
196
+ "target_delimiter": " ",
197
+ "fewshot_delimiter": "\n\n",
198
+ "fewshot_config": {
199
+ "sampler": "first_n",
200
+ "split": null,
201
+ "process_docs": null,
202
+ "fewshot_indices": null,
203
+ "samples": [
204
+ {
205
+ "input": "John journeyed to the bedroom. Mary grabbed the apple. Mary went back to the bathroom. Daniel journeyed to the bedroom. Daniel moved to the garden. Mary travelled to the kitchen.",
206
+ "question": "Where was the apple before the kitchen?",
207
+ "target": "Before the kitchen the apple was in the bathroom."
208
+ },
209
+ {
210
+ "input": "John went back to the bedroom. John went back to the garden. John went back to the kitchen. Sandra took the football. Sandra travelled to the garden. Sandra journeyed to the bedroom.",
211
+ "question": "Where was the football before the bedroom?",
212
+ "target": "Before the bedroom the football was in the garden."
213
+ }
214
+ ],
215
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
216
+ "doc_to_choice": null,
217
+ "doc_to_target": "{{target}}",
218
+ "gen_prefix": null,
219
+ "fewshot_delimiter": "\n\n",
220
+ "target_delimiter": " "
221
+ },
222
+ "num_fewshot": 2,
223
+ "metric_list": [
224
+ {
225
+ "metric": "acc",
226
+ "aggregation": "mean",
227
+ "higher_is_better": true
228
+ }
229
+ ],
230
+ "output_type": "generate_until",
231
+ "generation_kwargs": {
232
+ "do_sample": false,
233
+ "temperature": 0.0,
234
+ "max_gen_toks": 16,
235
+ "until": []
236
+ },
237
+ "repeats": 1,
238
+ "should_decontaminate": false,
239
+ "metadata": {
240
+ "version": 0.0,
241
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
242
+ "trust_remote_code": true,
243
+ "dtype": "bfloat16",
244
+ "max_length": 16384,
245
+ "attn_implementation": "sdpa",
246
+ "max_seq_lengths": "8k"
247
+ }
248
+ },
249
+ "babilong_qa4": {
250
+ "task": "babilong_qa4",
251
+ "custom_dataset": "def load_dataset(**kwargs):\n config_name = kwargs.get(\"max_seq_lengths\", \"0k\")\n\n # Get specific qa split\n qa_split = kwargs.get(\"qa_split\")\n\n eval_logger.info(\n f\"Loading babilong dataset: max_seq_lengths={config_name}, split={qa_split}\"\n )\n dataset = datasets.load_dataset(\n \"RMT-team/babilong-1k-samples\", name=config_name, split=qa_split\n )\n return {qa_split: dataset}\n",
252
+ "dataset_path": "RMT-team/babilong-1k-samples",
253
+ "dataset_kwargs": {
254
+ "qa_split": "qa4"
255
+ },
256
+ "test_split": "qa4",
257
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
258
+ "doc_to_target": "{{target}}",
259
+ "unsafe_code": false,
260
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n pred = postprocess_pred(results)\n target = doc.get(\"target\", \"\").strip()\n\n # String match\n score = 1.0 if target.lower() in pred[0].lower() else 0.0\n\n return {\"acc\": score}\n",
261
+ "description": "I will give you context with the facts about different people, their location and actions, hidden in some random text and a question. You need to answer the question based only on the information from the facts.\nYour answer should contain only one word - location. Do not write anything else after that.\n\n",
262
+ "target_delimiter": " ",
263
+ "fewshot_delimiter": "\n\n",
264
+ "fewshot_config": {
265
+ "sampler": "first_n",
266
+ "split": null,
267
+ "process_docs": null,
268
+ "fewshot_indices": null,
269
+ "samples": [
270
+ {
271
+ "input": "The hallway is south of the kitchen. The bedroom is north of the kitchen.",
272
+ "question": "What is the kitchen south of?",
273
+ "target": "bedroom"
274
+ },
275
+ {
276
+ "input": "The garden is west of the bedroom. The bedroom is west of the kitchen.",
277
+ "question": "What is west of the bedroom?",
278
+ "target": "garden"
279
+ }
280
+ ],
281
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
282
+ "doc_to_choice": null,
283
+ "doc_to_target": "{{target}}",
284
+ "gen_prefix": null,
285
+ "fewshot_delimiter": "\n\n",
286
+ "target_delimiter": " "
287
+ },
288
+ "num_fewshot": 2,
289
+ "metric_list": [
290
+ {
291
+ "metric": "acc",
292
+ "aggregation": "mean",
293
+ "higher_is_better": true
294
+ }
295
+ ],
296
+ "output_type": "generate_until",
297
+ "generation_kwargs": {
298
+ "do_sample": false,
299
+ "temperature": 0.0,
300
+ "max_gen_toks": 16,
301
+ "until": []
302
+ },
303
+ "repeats": 1,
304
+ "should_decontaminate": false,
305
+ "metadata": {
306
+ "version": 0.0,
307
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
308
+ "trust_remote_code": true,
309
+ "dtype": "bfloat16",
310
+ "max_length": 16384,
311
+ "attn_implementation": "sdpa",
312
+ "max_seq_lengths": "8k"
313
+ }
314
+ },
315
+ "babilong_qa5": {
316
+ "task": "babilong_qa5",
317
+ "custom_dataset": "def load_dataset(**kwargs):\n config_name = kwargs.get(\"max_seq_lengths\", \"0k\")\n\n # Get specific qa split\n qa_split = kwargs.get(\"qa_split\")\n\n eval_logger.info(\n f\"Loading babilong dataset: max_seq_lengths={config_name}, split={qa_split}\"\n )\n dataset = datasets.load_dataset(\n \"RMT-team/babilong-1k-samples\", name=config_name, split=qa_split\n )\n return {qa_split: dataset}\n",
318
+ "dataset_path": "RMT-team/babilong-1k-samples",
319
+ "dataset_kwargs": {
320
+ "qa_split": "qa5"
321
+ },
322
+ "test_split": "qa5",
323
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
324
+ "doc_to_target": "{{target}}",
325
+ "unsafe_code": false,
326
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n pred = postprocess_pred(results)\n target = doc.get(\"target\", \"\").strip()\n\n # String match\n score = 1.0 if target.lower() in pred[0].lower() else 0.0\n\n return {\"acc\": score}\n",
327
+ "description": "I will give you context with the facts about locations and their relations hidden in some random text and a question. You need to answer the question based only on the information from the facts.\nYour answer should contain only one word. Do not write anything else after that. Do not explain your answer.\n\n",
328
+ "target_delimiter": " ",
329
+ "fewshot_delimiter": "\n\n",
330
+ "fewshot_config": {
331
+ "sampler": "first_n",
332
+ "split": null,
333
+ "process_docs": null,
334
+ "fewshot_indices": null,
335
+ "samples": [
336
+ {
337
+ "input": "Mary picked up the apple there. Mary gave the apple to Fred. Mary moved to the bedroom. Bill took the milk there.",
338
+ "question": "Who did Mary give the apple to?",
339
+ "target": "Fred"
340
+ },
341
+ {
342
+ "input": "Jeff took the football there. Jeff passed the football to Fred. Jeff got the milk there. Bill travelled to the bedroom.",
343
+ "question": "Who gave the football?",
344
+ "target": "Jeff"
345
+ },
346
+ {
347
+ "input": "Fred picked up the apple there. Fred handed the apple to Bill. Bill journeyed to the bedroom. Jeff went back to the garden.",
348
+ "question": "What did Fred give to Bill?",
349
+ "target": "apple"
350
+ }
351
+ ],
352
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
353
+ "doc_to_choice": null,
354
+ "doc_to_target": "{{target}}",
355
+ "gen_prefix": null,
356
+ "fewshot_delimiter": "\n\n",
357
+ "target_delimiter": " "
358
+ },
359
+ "num_fewshot": 2,
360
+ "metric_list": [
361
+ {
362
+ "metric": "acc",
363
+ "aggregation": "mean",
364
+ "higher_is_better": true
365
+ }
366
+ ],
367
+ "output_type": "generate_until",
368
+ "generation_kwargs": {
369
+ "do_sample": false,
370
+ "temperature": 0.0,
371
+ "max_gen_toks": 16,
372
+ "until": []
373
+ },
374
+ "repeats": 1,
375
+ "should_decontaminate": false,
376
+ "metadata": {
377
+ "version": 0.0,
378
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
379
+ "trust_remote_code": true,
380
+ "dtype": "bfloat16",
381
+ "max_length": 16384,
382
+ "attn_implementation": "sdpa",
383
+ "max_seq_lengths": "8k"
384
+ }
385
+ }
386
+ },
387
+ "versions": {
388
+ "babilong_longctx": 0.0,
389
+ "babilong_qa1": 0.0,
390
+ "babilong_qa2": 0.0,
391
+ "babilong_qa3": 0.0,
392
+ "babilong_qa4": 0.0,
393
+ "babilong_qa5": 0.0
394
+ },
395
+ "n-shot": {
396
+ "babilong_qa1": 2,
397
+ "babilong_qa2": 2,
398
+ "babilong_qa3": 2,
399
+ "babilong_qa4": 2,
400
+ "babilong_qa5": 2
401
+ },
402
+ "higher_is_better": {
403
+ "babilong_longctx": {
404
+ "acc": true
405
+ },
406
+ "babilong_qa1": {
407
+ "acc": true
408
+ },
409
+ "babilong_qa2": {
410
+ "acc": true
411
+ },
412
+ "babilong_qa3": {
413
+ "acc": true
414
+ },
415
+ "babilong_qa4": {
416
+ "acc": true
417
+ },
418
+ "babilong_qa5": {
419
+ "acc": true
420
+ }
421
+ },
422
+ "n-samples": {
423
+ "babilong_qa1": {
424
+ "original": 1000,
425
+ "effective": 50
426
+ },
427
+ "babilong_qa2": {
428
+ "original": 999,
429
+ "effective": 50
430
+ },
431
+ "babilong_qa3": {
432
+ "original": 999,
433
+ "effective": 50
434
+ },
435
+ "babilong_qa4": {
436
+ "original": 999,
437
+ "effective": 50
438
+ },
439
+ "babilong_qa5": {
440
+ "original": 999,
441
+ "effective": 50
442
+ }
443
+ },
444
+ "config": {
445
+ "model": "hf",
446
+ "model_args": {
447
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
448
+ "trust_remote_code": true,
449
+ "dtype": "bfloat16",
450
+ "max_length": 16384,
451
+ "attn_implementation": "sdpa"
452
+ },
453
+ "model_num_parameters": 2031969308,
454
+ "model_dtype": "torch.bfloat16",
455
+ "model_revision": "main",
456
+ "model_sha": "",
457
+ "batch_size": "1",
458
+ "batch_sizes": [],
459
+ "device": "cuda:0",
460
+ "use_cache": null,
461
+ "limit": 50.0,
462
+ "bootstrap_iters": 100000,
463
+ "gen_kwargs": {},
464
+ "random_seed": 0,
465
+ "numpy_seed": 1234,
466
+ "torch_seed": 1234,
467
+ "fewshot_seed": 1234
468
+ },
469
+ "git_hash": null,
470
+ "date": 1784387878.9986007,
471
+ "pretty_env_info": "PyTorch version: 2.9.1+cu128\nIs debug build: False\nCUDA used to build PyTorch: 12.8\nROCM used to build PyTorch: N/A\n\nOS: Ubuntu 22.04.5 LTS (x86_64)\nGCC version: Could not collect\nClang version: Could not collect\nCMake version: version 4.1.2\nLibc version: glibc-2.35\n\nPython version: 3.11.14 | packaged by conda-forge | (main, Oct 22 2025, 22:46:25) [GCC 14.3.0] (64-bit runtime)\nPython platform: Linux-6.12.90-120.164.amzn2023.x86_64-x86_64-with-glibc2.35\nIs CUDA available: True\nCUDA runtime version: Could not collect\nCUDA_MODULE_LOADING set to: \nGPU models and configuration: GPU 0: NVIDIA A100-SXM4-80GB\nNvidia driver version: 580.159.03\ncuDNN version: Could not collect\nIs XPU available: False\nHIP runtime version: N/A\nMIOpen runtime version: N/A\nIs XNNPACK available: True\n\nCPU:\nArchitecture: x86_64\nCPU op-mode(s): 32-bit, 64-bit\nAddress sizes: 46 bits physical, 48 bits virtual\nByte Order: Little Endian\nCPU(s): 96\nOn-line CPU(s) list: 0-95\nVendor ID: GenuineIntel\nModel name: Intel(R) Xeon(R) Platinum 8275CL CPU @ 3.00GHz\nCPU family: 6\nModel: 85\nThread(s) per core: 2\nCore(s) per socket: 24\nSocket(s): 2\nStepping: 7\nBogoMIPS: 5999.99\nFlags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon rep_good nopl xtopology nonstop_tsc cpuid aperfmperf tsc_known_freq pni pclmulqdq monitor ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand hypervisor lahf_lm abm 3dnowprefetch pti fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid mpx avx512f avx512dq rdseed adx smap clflushopt clwb avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves ida arat pku ospke\nHypervisor vendor: KVM\nVirtualization type: full\nL1d cache: 1.5 MiB (48 instances)\nL1i cache: 1.5 MiB (48 instances)\nL2 cache: 48 MiB (48 instances)\nL3 cache: 71.5 MiB (2 instances)\nNUMA node(s): 2\nNUMA node0 CPU(s): 0-23,48-71\nNUMA node1 CPU(s): 24-47,72-95\nVulnerability Gather data sampling: Unknown: Dependent on hypervisor status\nVulnerability Indirect target selection: Mitigation; Aligned branch/return thunks\nVulnerability Itlb multihit: KVM: Mitigation: VMX unsupported\nVulnerability L1tf: Mitigation; PTE Inversion\nVulnerability Mds: Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown\nVulnerability Meltdown: Mitigation; PTI\nVulnerability Mmio stale data: Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown\nVulnerability Reg file data sampling: Not affected\nVulnerability Retbleed: Vulnerable\nVulnerability Spec rstack overflow: Not affected\nVulnerability Spec store bypass: Vulnerable\nVulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization\nVulnerability Spectre v2: Mitigation; Retpolines; STIBP disabled; RSB filling; PBRSB-eIBRS Not affected; BHI Retpoline\nVulnerability Srbds: Not affected\nVulnerability Tsa: Not affected\nVulnerability Tsx async abort: Not affected\nVulnerability Vmscape: Not affected\n\nVersions of relevant libraries:\n[pip3] numpy==2.3.4\n[pip3] nvidia-cublas-cu12==12.8.4.1\n[pip3] nvidia-cuda-cupti-cu12==12.8.90\n[pip3] nvidia-cuda-nvrtc-cu12==12.8.93\n[pip3] nvidia-cuda-runtime-cu12==12.8.90\n[pip3] nvidia-cudnn-cu12==9.10.2.21\n[pip3] nvidia-cufft-cu12==11.3.3.83\n[pip3] nvidia-curand-cu12==10.3.9.90\n[pip3] nvidia-cusolver-cu12==11.7.3.90\n[pip3] nvidia-cusparse-cu12==12.5.8.93\n[pip3] nvidia-cusparselt-cu12==0.7.1\n[pip3] nvidia-nccl-cu12==2.27.5\n[pip3] nvidia-nvjitlink-cu12==12.8.93\n[pip3] nvidia-nvtx-cu12==12.8.90\n[pip3] optree==0.17.0\n[pip3] torch==2.9.1+cu128\n[pip3] torchaudio==2.9.1+cu128\n[pip3] torchelastic==0.2.2\n[pip3] torchvision==0.24.1+cu128\n[pip3] triton==3.5.1\n[conda] numpy 2.3.4 py311h2e04523_0 conda-forge\n[conda] nvidia-cublas-cu12 12.8.4.1 pypi_0 pypi\n[conda] nvidia-cuda-cupti-cu12 12.8.90 pypi_0 pypi\n[conda] nvidia-cuda-nvrtc-cu12 12.8.93 pypi_0 pypi\n[conda] nvidia-cuda-runtime-cu12 12.8.90 pypi_0 pypi\n[conda] nvidia-cudnn-cu12 9.10.2.21 pypi_0 pypi\n[conda] nvidia-cufft-cu12 11.3.3.83 pypi_0 pypi\n[conda] nvidia-curand-cu12 10.3.9.90 pypi_0 pypi\n[conda] nvidia-cusolver-cu12 11.7.3.90 pypi_0 pypi\n[conda] nvidia-cusparse-cu12 12.5.8.93 pypi_0 pypi\n[conda] nvidia-cusparselt-cu12 0.7.1 pypi_0 pypi\n[conda] nvidia-nccl-cu12 2.27.5 pypi_0 pypi\n[conda] nvidia-nvjitlink-cu12 12.8.93 pypi_0 pypi\n[conda] nvidia-nvtx-cu12 12.8.90 pypi_0 pypi\n[conda] optree 0.17.0 pypi_0 pypi\n[conda] torch 2.9.1+cu128 pypi_0 pypi\n[conda] torchaudio 2.9.1+cu128 pypi_0 pypi\n[conda] torchelastic 0.2.2 pypi_0 pypi\n[conda] torchvision 0.24.1+cu128 pypi_0 pypi\n[conda] triton 3.5.1 pypi_0 pypi",
472
+ "transformers_version": "4.54.0",
473
+ "lm_eval_version": "0.4.11",
474
+ "upper_git_hash": null,
475
+ "tokenizer_pad_token": [
476
+ "<|endoftext|>",
477
+ "151643"
478
+ ],
479
+ "tokenizer_eos_token": [
480
+ "<|im_end|>",
481
+ "151645"
482
+ ],
483
+ "tokenizer_bos_token": [
484
+ null,
485
+ "None"
486
+ ],
487
+ "eot_token_id": 151645,
488
+ "max_length": 16384,
489
+ "task_hashes": {
490
+ "babilong_qa1": "bbdcd109bea1bb245e118e814bd0563de6384bc875a21df5218ab4ed75b27380",
491
+ "babilong_qa2": "fa4f8d12bdd481e02a1051414304811f648caf7bce8f89dcf03ecbf5b86f1e92",
492
+ "babilong_qa3": "737b64396c8b2e5f96da9796bcd5e448f08e5a13e32ae7b610d31a60269b6e48",
493
+ "babilong_qa4": "298a25727c69c9babddcd78f438cae1aa21a388e14e7db90b2206eff753c90c0",
494
+ "babilong_qa5": "e60afcdcdd1c4b3e24b002ffbf8f87085f5ec4a02fcd72318f042dcdafe230f4"
495
+ },
496
+ "model_source": "hf",
497
+ "model_name": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
498
+ "model_name_sanitized": "__workspace__outputs__l2a_style__stage2__checkpoint-25",
499
+ "system_instruction": null,
500
+ "system_instruction_sha": null,
501
+ "fewshot_as_multiturn": null,
502
+ "chat_template": null,
503
+ "chat_template_sha": null,
504
+ "total_evaluation_time_seconds": "342.35258994298056"
505
+ }
outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa1_2026-07-18T15-23-38.141056.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4bb665d4b030c9de83733515a244cabc2345f83990ccc55d24110b3ec5bfc739
3
+ size 3164084
outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa2_2026-07-18T15-23-38.141056.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a41a461533929734a423ed2f311990f3c1d38a85ada1f9730acc45b1bb70e1be
3
+ size 3147502
outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa3_2026-07-18T15-23-38.141056.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:71f73bb1d538565321b650de5af906d40c1206e407a27c69cb2554101954b1c2
3
+ size 3126982
outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa4_2026-07-18T15-23-38.141056.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6d3469275d1131254499c2c7ffd046c1c3f6014ba45f3580f3fb4149d00c2429
3
+ size 3137614
outputs/eval/token_t045/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa5_2026-07-18T15-23-38.141056.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5d6d02c0fb088e4bf897d06df24d9d6969adddcd4873a16dc18a405d2d8e5a38
3
+ size 3135534
outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/results_2026-07-18T15-05-41.414854.json ADDED
@@ -0,0 +1,821 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "results": {
3
+ "niah_multikey_2": {
4
+ "alias": "niah_multikey_2",
5
+ "4096,none": -1,
6
+ "4096_stderr,none": "N/A",
7
+ "8192,none": 1.0,
8
+ "8192_stderr,none": "N/A"
9
+ },
10
+ "niah_multiquery": {
11
+ "alias": "niah_multiquery",
12
+ "4096,none": -1,
13
+ "4096_stderr,none": "N/A",
14
+ "8192,none": 0.9875,
15
+ "8192_stderr,none": "N/A"
16
+ },
17
+ "niah_single_1": {
18
+ "alias": "niah_single_1",
19
+ "4096,none": -1,
20
+ "4096_stderr,none": "N/A",
21
+ "8192,none": 1.0,
22
+ "8192_stderr,none": "N/A"
23
+ },
24
+ "niah_single_3": {
25
+ "alias": "niah_single_3",
26
+ "4096,none": -1,
27
+ "4096_stderr,none": "N/A",
28
+ "8192,none": 0.95,
29
+ "8192_stderr,none": "N/A"
30
+ },
31
+ "ruler_fwe": {
32
+ "alias": "ruler_fwe",
33
+ "4096,none": -1,
34
+ "4096_stderr,none": "N/A",
35
+ "8192,none": 0.7666666666666666,
36
+ "8192_stderr,none": "N/A"
37
+ },
38
+ "ruler_qa_hotpot": {
39
+ "alias": "ruler_qa_hotpot",
40
+ "4096,none": -1,
41
+ "4096_stderr,none": "N/A",
42
+ "8192,none": 0.4,
43
+ "8192_stderr,none": "N/A"
44
+ },
45
+ "ruler_vt": {
46
+ "alias": "ruler_vt",
47
+ "4096,none": -1,
48
+ "4096_stderr,none": "N/A",
49
+ "8192,none": 0.8700000000000003,
50
+ "8192_stderr,none": "N/A"
51
+ }
52
+ },
53
+ "group_subtasks": {
54
+ "niah_single_1": [],
55
+ "niah_single_3": [],
56
+ "niah_multikey_2": [],
57
+ "niah_multiquery": [],
58
+ "ruler_vt": [],
59
+ "ruler_fwe": [],
60
+ "ruler_qa_hotpot": []
61
+ },
62
+ "configs": {
63
+ "niah_multikey_2": {
64
+ "task": "niah_multikey_2",
65
+ "tag": [
66
+ "longcxt"
67
+ ],
68
+ "custom_dataset": "def niah_multikey_2(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"needle\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"needle\",\n type_needle_k=\"words\",\n type_needle_v=\"numbers\",\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
69
+ "dataset_path": "",
70
+ "dataset_name": "",
71
+ "test_split": "test",
72
+ "doc_to_text": "{{input}}",
73
+ "doc_to_target": "{{outputs}}",
74
+ "unsafe_code": false,
75
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
76
+ "description": "",
77
+ "target_delimiter": " ",
78
+ "fewshot_delimiter": "\n\n",
79
+ "fewshot_config": {
80
+ "sampler": "default",
81
+ "split": null,
82
+ "process_docs": null,
83
+ "fewshot_indices": null,
84
+ "samples": null,
85
+ "doc_to_text": "{{input}}",
86
+ "doc_to_choice": null,
87
+ "doc_to_target": "{{outputs}}",
88
+ "gen_prefix": "{{gen_prefix}}",
89
+ "fewshot_delimiter": "\n\n",
90
+ "target_delimiter": " "
91
+ },
92
+ "num_fewshot": 0,
93
+ "metric_list": [
94
+ {
95
+ "metric": "4096",
96
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
97
+ "higher_is_better": true
98
+ },
99
+ {
100
+ "metric": "8192",
101
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
102
+ "higher_is_better": true
103
+ },
104
+ {
105
+ "metric": "16384",
106
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
107
+ "higher_is_better": true
108
+ },
109
+ {
110
+ "metric": "32768",
111
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
112
+ "higher_is_better": true
113
+ },
114
+ {
115
+ "metric": "65536",
116
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
117
+ "higher_is_better": true
118
+ },
119
+ {
120
+ "metric": "131072",
121
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
122
+ "higher_is_better": true
123
+ }
124
+ ],
125
+ "output_type": "generate_until",
126
+ "generation_kwargs": {
127
+ "do_sample": false,
128
+ "temperature": 0.0,
129
+ "max_gen_toks": 128,
130
+ "until": []
131
+ },
132
+ "repeats": 1,
133
+ "should_decontaminate": false,
134
+ "gen_prefix": "{{gen_prefix}}",
135
+ "metadata": {
136
+ "version": 1.0,
137
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
138
+ "trust_remote_code": true,
139
+ "dtype": "bfloat16",
140
+ "max_length": 16384,
141
+ "attn_implementation": "sdpa",
142
+ "max_seq_lengths": [
143
+ 8192
144
+ ]
145
+ }
146
+ },
147
+ "niah_multiquery": {
148
+ "task": "niah_multiquery",
149
+ "tag": [
150
+ "longcxt"
151
+ ],
152
+ "custom_dataset": "def niah_multiquery(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"essay\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"essay\",\n type_needle_k=\"words\",\n type_needle_v=\"numbers\",\n num_needle_q=4,\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
153
+ "dataset_path": "",
154
+ "dataset_name": "",
155
+ "test_split": "test",
156
+ "doc_to_text": "{{input}}",
157
+ "doc_to_target": "{{outputs}}",
158
+ "unsafe_code": false,
159
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
160
+ "description": "",
161
+ "target_delimiter": " ",
162
+ "fewshot_delimiter": "\n\n",
163
+ "fewshot_config": {
164
+ "sampler": "default",
165
+ "split": null,
166
+ "process_docs": null,
167
+ "fewshot_indices": null,
168
+ "samples": null,
169
+ "doc_to_text": "{{input}}",
170
+ "doc_to_choice": null,
171
+ "doc_to_target": "{{outputs}}",
172
+ "gen_prefix": "{{gen_prefix}}",
173
+ "fewshot_delimiter": "\n\n",
174
+ "target_delimiter": " "
175
+ },
176
+ "num_fewshot": 0,
177
+ "metric_list": [
178
+ {
179
+ "metric": "4096",
180
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
181
+ "higher_is_better": true
182
+ },
183
+ {
184
+ "metric": "8192",
185
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
186
+ "higher_is_better": true
187
+ },
188
+ {
189
+ "metric": "16384",
190
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
191
+ "higher_is_better": true
192
+ },
193
+ {
194
+ "metric": "32768",
195
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
196
+ "higher_is_better": true
197
+ },
198
+ {
199
+ "metric": "65536",
200
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
201
+ "higher_is_better": true
202
+ },
203
+ {
204
+ "metric": "131072",
205
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
206
+ "higher_is_better": true
207
+ }
208
+ ],
209
+ "output_type": "generate_until",
210
+ "generation_kwargs": {
211
+ "do_sample": false,
212
+ "temperature": 0.0,
213
+ "max_gen_toks": 128,
214
+ "until": []
215
+ },
216
+ "repeats": 1,
217
+ "should_decontaminate": false,
218
+ "gen_prefix": "{{gen_prefix}}",
219
+ "metadata": {
220
+ "version": 1.0,
221
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
222
+ "trust_remote_code": true,
223
+ "dtype": "bfloat16",
224
+ "max_length": 16384,
225
+ "attn_implementation": "sdpa",
226
+ "max_seq_lengths": [
227
+ 8192
228
+ ]
229
+ }
230
+ },
231
+ "niah_single_1": {
232
+ "task": "niah_single_1",
233
+ "tag": [
234
+ "longcxt"
235
+ ],
236
+ "custom_dataset": "def niah_single_1(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"repeat\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"repeat\",\n type_needle_k=\"words\",\n type_needle_v=\"numbers\",\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
237
+ "dataset_path": "",
238
+ "dataset_name": "",
239
+ "test_split": "test",
240
+ "doc_to_text": "{{input}}",
241
+ "doc_to_target": "{{outputs}}",
242
+ "unsafe_code": false,
243
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
244
+ "description": "",
245
+ "target_delimiter": " ",
246
+ "fewshot_delimiter": "\n\n",
247
+ "fewshot_config": {
248
+ "sampler": "default",
249
+ "split": null,
250
+ "process_docs": null,
251
+ "fewshot_indices": null,
252
+ "samples": null,
253
+ "doc_to_text": "{{input}}",
254
+ "doc_to_choice": null,
255
+ "doc_to_target": "{{outputs}}",
256
+ "gen_prefix": "{{gen_prefix}}",
257
+ "fewshot_delimiter": "\n\n",
258
+ "target_delimiter": " "
259
+ },
260
+ "num_fewshot": 0,
261
+ "metric_list": [
262
+ {
263
+ "metric": "4096",
264
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
265
+ "higher_is_better": true
266
+ },
267
+ {
268
+ "metric": "8192",
269
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
270
+ "higher_is_better": true
271
+ },
272
+ {
273
+ "metric": "16384",
274
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
275
+ "higher_is_better": true
276
+ },
277
+ {
278
+ "metric": "32768",
279
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
280
+ "higher_is_better": true
281
+ },
282
+ {
283
+ "metric": "65536",
284
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
285
+ "higher_is_better": true
286
+ },
287
+ {
288
+ "metric": "131072",
289
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
290
+ "higher_is_better": true
291
+ }
292
+ ],
293
+ "output_type": "generate_until",
294
+ "generation_kwargs": {
295
+ "do_sample": false,
296
+ "temperature": 0.0,
297
+ "max_gen_toks": 128,
298
+ "until": []
299
+ },
300
+ "repeats": 1,
301
+ "should_decontaminate": false,
302
+ "gen_prefix": "{{gen_prefix}}",
303
+ "metadata": {
304
+ "version": 1.0,
305
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
306
+ "trust_remote_code": true,
307
+ "dtype": "bfloat16",
308
+ "max_length": 16384,
309
+ "attn_implementation": "sdpa",
310
+ "max_seq_lengths": [
311
+ 8192
312
+ ]
313
+ }
314
+ },
315
+ "niah_single_3": {
316
+ "task": "niah_single_3",
317
+ "tag": [
318
+ "longcxt"
319
+ ],
320
+ "custom_dataset": "def niah_single_3(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"essay\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"essay\",\n type_needle_k=\"words\",\n type_needle_v=\"uuids\",\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
321
+ "dataset_path": "",
322
+ "dataset_name": "",
323
+ "test_split": "test",
324
+ "doc_to_text": "{{input}}",
325
+ "doc_to_target": "{{outputs}}",
326
+ "unsafe_code": false,
327
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
328
+ "description": "",
329
+ "target_delimiter": " ",
330
+ "fewshot_delimiter": "\n\n",
331
+ "fewshot_config": {
332
+ "sampler": "default",
333
+ "split": null,
334
+ "process_docs": null,
335
+ "fewshot_indices": null,
336
+ "samples": null,
337
+ "doc_to_text": "{{input}}",
338
+ "doc_to_choice": null,
339
+ "doc_to_target": "{{outputs}}",
340
+ "gen_prefix": "{{gen_prefix}}",
341
+ "fewshot_delimiter": "\n\n",
342
+ "target_delimiter": " "
343
+ },
344
+ "num_fewshot": 0,
345
+ "metric_list": [
346
+ {
347
+ "metric": "4096",
348
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
349
+ "higher_is_better": true
350
+ },
351
+ {
352
+ "metric": "8192",
353
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
354
+ "higher_is_better": true
355
+ },
356
+ {
357
+ "metric": "16384",
358
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
359
+ "higher_is_better": true
360
+ },
361
+ {
362
+ "metric": "32768",
363
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
364
+ "higher_is_better": true
365
+ },
366
+ {
367
+ "metric": "65536",
368
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
369
+ "higher_is_better": true
370
+ },
371
+ {
372
+ "metric": "131072",
373
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
374
+ "higher_is_better": true
375
+ }
376
+ ],
377
+ "output_type": "generate_until",
378
+ "generation_kwargs": {
379
+ "do_sample": false,
380
+ "temperature": 0.0,
381
+ "max_gen_toks": 128,
382
+ "until": []
383
+ },
384
+ "repeats": 1,
385
+ "should_decontaminate": false,
386
+ "gen_prefix": "{{gen_prefix}}",
387
+ "metadata": {
388
+ "version": 1.0,
389
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
390
+ "trust_remote_code": true,
391
+ "dtype": "bfloat16",
392
+ "max_length": 16384,
393
+ "attn_implementation": "sdpa",
394
+ "max_seq_lengths": [
395
+ 8192
396
+ ]
397
+ }
398
+ },
399
+ "ruler_fwe": {
400
+ "task": "ruler_fwe",
401
+ "tag": [
402
+ "longcxt"
403
+ ],
404
+ "custom_dataset": "def fwe_download(**kwargs):\n pretrained = kwargs.get(\"tokenizer\", kwargs.get(\"pretrained\", {}))\n df = (\n get_dataset(pretrained, max_seq_length=seq)\n for seq in kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n )\n\n return {\n \"test\": datasets.Dataset.from_list(\n list(itertools.chain.from_iterable(df)), split=datasets.Split.TEST\n )\n }\n",
405
+ "dataset_path": "",
406
+ "dataset_name": "",
407
+ "test_split": "test",
408
+ "doc_to_text": "{{input}}",
409
+ "doc_to_target": "{{outputs}}",
410
+ "unsafe_code": false,
411
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
412
+ "description": "",
413
+ "target_delimiter": " ",
414
+ "fewshot_delimiter": "\n\n",
415
+ "fewshot_config": {
416
+ "sampler": "default",
417
+ "split": null,
418
+ "process_docs": null,
419
+ "fewshot_indices": null,
420
+ "samples": null,
421
+ "doc_to_text": "{{input}}",
422
+ "doc_to_choice": null,
423
+ "doc_to_target": "{{outputs}}",
424
+ "gen_prefix": "{{gen_prefix}}",
425
+ "fewshot_delimiter": "\n\n",
426
+ "target_delimiter": " "
427
+ },
428
+ "num_fewshot": 0,
429
+ "metric_list": [
430
+ {
431
+ "metric": "4096",
432
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
433
+ "higher_is_better": true
434
+ },
435
+ {
436
+ "metric": "8192",
437
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
438
+ "higher_is_better": true
439
+ },
440
+ {
441
+ "metric": "16384",
442
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
443
+ "higher_is_better": true
444
+ },
445
+ {
446
+ "metric": "32768",
447
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
448
+ "higher_is_better": true
449
+ },
450
+ {
451
+ "metric": "65536",
452
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
453
+ "higher_is_better": true
454
+ },
455
+ {
456
+ "metric": "131072",
457
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
458
+ "higher_is_better": true
459
+ }
460
+ ],
461
+ "output_type": "generate_until",
462
+ "generation_kwargs": {
463
+ "do_sample": false,
464
+ "temperature": 0.0,
465
+ "max_gen_toks": 50,
466
+ "until": []
467
+ },
468
+ "repeats": 1,
469
+ "should_decontaminate": false,
470
+ "gen_prefix": "{{gen_prefix}}",
471
+ "metadata": {
472
+ "version": 1.0,
473
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
474
+ "trust_remote_code": true,
475
+ "dtype": "bfloat16",
476
+ "max_length": 16384,
477
+ "attn_implementation": "sdpa",
478
+ "max_seq_lengths": [
479
+ 8192
480
+ ]
481
+ }
482
+ },
483
+ "ruler_qa_hotpot": {
484
+ "task": "ruler_qa_hotpot",
485
+ "tag": [
486
+ "longcxt"
487
+ ],
488
+ "custom_dataset": "def get_hotpotqa(**kwargs):\n return get_qa_dataset(\"hotpotqa\", **kwargs)\n",
489
+ "dataset_path": "",
490
+ "dataset_name": "",
491
+ "test_split": "test",
492
+ "doc_to_text": "{{input}}",
493
+ "doc_to_target": "{{outputs}}",
494
+ "unsafe_code": false,
495
+ "process_results": "def process_results_part(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_part(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
496
+ "description": "",
497
+ "target_delimiter": " ",
498
+ "fewshot_delimiter": "\n\n",
499
+ "fewshot_config": {
500
+ "sampler": "default",
501
+ "split": null,
502
+ "process_docs": null,
503
+ "fewshot_indices": null,
504
+ "samples": null,
505
+ "doc_to_text": "{{input}}",
506
+ "doc_to_choice": null,
507
+ "doc_to_target": "{{outputs}}",
508
+ "gen_prefix": "{{gen_prefix}}",
509
+ "fewshot_delimiter": "\n\n",
510
+ "target_delimiter": " "
511
+ },
512
+ "num_fewshot": 0,
513
+ "metric_list": [
514
+ {
515
+ "metric": "4096",
516
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
517
+ "higher_is_better": true
518
+ },
519
+ {
520
+ "metric": "8192",
521
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
522
+ "higher_is_better": true
523
+ },
524
+ {
525
+ "metric": "16384",
526
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
527
+ "higher_is_better": true
528
+ },
529
+ {
530
+ "metric": "32768",
531
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
532
+ "higher_is_better": true
533
+ },
534
+ {
535
+ "metric": "65536",
536
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
537
+ "higher_is_better": true
538
+ },
539
+ {
540
+ "metric": "131072",
541
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
542
+ "higher_is_better": true
543
+ }
544
+ ],
545
+ "output_type": "generate_until",
546
+ "generation_kwargs": {
547
+ "do_sample": false,
548
+ "temperature": 0.0,
549
+ "max_gen_toks": 32,
550
+ "until": []
551
+ },
552
+ "repeats": 1,
553
+ "should_decontaminate": false,
554
+ "gen_prefix": "{{gen_prefix}}",
555
+ "metadata": {
556
+ "version": 1.0,
557
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
558
+ "trust_remote_code": true,
559
+ "dtype": "bfloat16",
560
+ "max_length": 16384,
561
+ "attn_implementation": "sdpa",
562
+ "max_seq_lengths": [
563
+ 8192
564
+ ]
565
+ }
566
+ },
567
+ "ruler_vt": {
568
+ "task": "ruler_vt",
569
+ "tag": [
570
+ "longcxt"
571
+ ],
572
+ "custom_dataset": "def get_vt_dataset(**kwargs) -> dict[str, datasets.Dataset]:\n pretrained = kwargs.get(\"tokenizer\", kwargs.get(\"pretrained\", \"\"))\n df = (\n get_dataset(tokenizer=get_tokenizer(pretrained), seq=seq)\n for seq in kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n )\n\n return {\n \"test\": datasets.Dataset.from_list(\n list(itertools.chain.from_iterable(df)), split=datasets.Split.TEST\n )\n }\n",
573
+ "dataset_path": "",
574
+ "dataset_name": "",
575
+ "test_split": "test",
576
+ "doc_to_text": "{{input}}",
577
+ "doc_to_target": "{{outputs}}",
578
+ "unsafe_code": false,
579
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
580
+ "description": "",
581
+ "target_delimiter": " ",
582
+ "fewshot_delimiter": "\n\n",
583
+ "fewshot_config": {
584
+ "sampler": "default",
585
+ "split": null,
586
+ "process_docs": null,
587
+ "fewshot_indices": null,
588
+ "samples": null,
589
+ "doc_to_text": "{{input}}",
590
+ "doc_to_choice": null,
591
+ "doc_to_target": "{{outputs}}",
592
+ "gen_prefix": "{{gen_prefix}}",
593
+ "fewshot_delimiter": "\n\n",
594
+ "target_delimiter": " "
595
+ },
596
+ "num_fewshot": 0,
597
+ "metric_list": [
598
+ {
599
+ "metric": "4096",
600
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
601
+ "higher_is_better": true
602
+ },
603
+ {
604
+ "metric": "8192",
605
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
606
+ "higher_is_better": true
607
+ },
608
+ {
609
+ "metric": "16384",
610
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
611
+ "higher_is_better": true
612
+ },
613
+ {
614
+ "metric": "32768",
615
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
616
+ "higher_is_better": true
617
+ },
618
+ {
619
+ "metric": "65536",
620
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
621
+ "higher_is_better": true
622
+ },
623
+ {
624
+ "metric": "131072",
625
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
626
+ "higher_is_better": true
627
+ }
628
+ ],
629
+ "output_type": "generate_until",
630
+ "generation_kwargs": {
631
+ "do_sample": false,
632
+ "temperature": 0.0,
633
+ "max_gen_toks": 30,
634
+ "until": []
635
+ },
636
+ "repeats": 1,
637
+ "should_decontaminate": false,
638
+ "gen_prefix": "{{gen_prefix}}",
639
+ "metadata": {
640
+ "version": 1.0,
641
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
642
+ "trust_remote_code": true,
643
+ "dtype": "bfloat16",
644
+ "max_length": 16384,
645
+ "attn_implementation": "sdpa",
646
+ "max_seq_lengths": [
647
+ 8192
648
+ ]
649
+ }
650
+ }
651
+ },
652
+ "versions": {
653
+ "niah_multikey_2": 1.0,
654
+ "niah_multiquery": 1.0,
655
+ "niah_single_1": 1.0,
656
+ "niah_single_3": 1.0,
657
+ "ruler_fwe": 1.0,
658
+ "ruler_qa_hotpot": 1.0,
659
+ "ruler_vt": 1.0
660
+ },
661
+ "n-shot": {
662
+ "niah_multikey_2": 0,
663
+ "niah_multiquery": 0,
664
+ "niah_single_1": 0,
665
+ "niah_single_3": 0,
666
+ "ruler_fwe": 0,
667
+ "ruler_qa_hotpot": 0,
668
+ "ruler_vt": 0
669
+ },
670
+ "higher_is_better": {
671
+ "niah_multikey_2": {
672
+ "4096": true,
673
+ "8192": true,
674
+ "16384": true,
675
+ "32768": true,
676
+ "65536": true,
677
+ "131072": true
678
+ },
679
+ "niah_multiquery": {
680
+ "4096": true,
681
+ "8192": true,
682
+ "16384": true,
683
+ "32768": true,
684
+ "65536": true,
685
+ "131072": true
686
+ },
687
+ "niah_single_1": {
688
+ "4096": true,
689
+ "8192": true,
690
+ "16384": true,
691
+ "32768": true,
692
+ "65536": true,
693
+ "131072": true
694
+ },
695
+ "niah_single_3": {
696
+ "4096": true,
697
+ "8192": true,
698
+ "16384": true,
699
+ "32768": true,
700
+ "65536": true,
701
+ "131072": true
702
+ },
703
+ "ruler_fwe": {
704
+ "4096": true,
705
+ "8192": true,
706
+ "16384": true,
707
+ "32768": true,
708
+ "65536": true,
709
+ "131072": true
710
+ },
711
+ "ruler_qa_hotpot": {
712
+ "4096": true,
713
+ "8192": true,
714
+ "16384": true,
715
+ "32768": true,
716
+ "65536": true,
717
+ "131072": true
718
+ },
719
+ "ruler_vt": {
720
+ "4096": true,
721
+ "8192": true,
722
+ "16384": true,
723
+ "32768": true,
724
+ "65536": true,
725
+ "131072": true
726
+ }
727
+ },
728
+ "n-samples": {
729
+ "ruler_qa_hotpot": {
730
+ "original": 500,
731
+ "effective": 20
732
+ },
733
+ "ruler_fwe": {
734
+ "original": 500,
735
+ "effective": 20
736
+ },
737
+ "ruler_vt": {
738
+ "original": 500,
739
+ "effective": 20
740
+ },
741
+ "niah_multiquery": {
742
+ "original": 500,
743
+ "effective": 20
744
+ },
745
+ "niah_multikey_2": {
746
+ "original": 500,
747
+ "effective": 20
748
+ },
749
+ "niah_single_3": {
750
+ "original": 500,
751
+ "effective": 20
752
+ },
753
+ "niah_single_1": {
754
+ "original": 500,
755
+ "effective": 20
756
+ }
757
+ },
758
+ "config": {
759
+ "model": "hf",
760
+ "model_args": {
761
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
762
+ "trust_remote_code": true,
763
+ "dtype": "bfloat16",
764
+ "max_length": 16384,
765
+ "attn_implementation": "sdpa"
766
+ },
767
+ "model_num_parameters": 2031969308,
768
+ "model_dtype": "torch.bfloat16",
769
+ "model_revision": "main",
770
+ "model_sha": "",
771
+ "batch_size": "1",
772
+ "batch_sizes": [],
773
+ "device": "cuda:0",
774
+ "use_cache": null,
775
+ "limit": 20.0,
776
+ "bootstrap_iters": 100000,
777
+ "gen_kwargs": {},
778
+ "random_seed": 0,
779
+ "numpy_seed": 1234,
780
+ "torch_seed": 1234,
781
+ "fewshot_seed": 1234
782
+ },
783
+ "git_hash": null,
784
+ "date": 1784386405.0936582,
785
+ "pretty_env_info": "PyTorch version: 2.9.1+cu128\nIs debug build: False\nCUDA used to build PyTorch: 12.8\nROCM used to build PyTorch: N/A\n\nOS: Ubuntu 22.04.5 LTS (x86_64)\nGCC version: Could not collect\nClang version: Could not collect\nCMake version: version 4.1.2\nLibc version: glibc-2.35\n\nPython version: 3.11.14 | packaged by conda-forge | (main, Oct 22 2025, 22:46:25) [GCC 14.3.0] (64-bit runtime)\nPython platform: Linux-6.12.90-120.164.amzn2023.x86_64-x86_64-with-glibc2.35\nIs CUDA available: True\nCUDA runtime version: Could not collect\nCUDA_MODULE_LOADING set to: \nGPU models and configuration: GPU 0: NVIDIA A100-SXM4-80GB\nNvidia driver version: 580.159.03\ncuDNN version: Could not collect\nIs XPU available: False\nHIP runtime version: N/A\nMIOpen runtime version: N/A\nIs XNNPACK available: True\n\nCPU:\nArchitecture: x86_64\nCPU op-mode(s): 32-bit, 64-bit\nAddress sizes: 46 bits physical, 48 bits virtual\nByte Order: Little Endian\nCPU(s): 96\nOn-line CPU(s) list: 0-95\nVendor ID: GenuineIntel\nModel name: Intel(R) Xeon(R) Platinum 8275CL CPU @ 3.00GHz\nCPU family: 6\nModel: 85\nThread(s) per core: 2\nCore(s) per socket: 24\nSocket(s): 2\nStepping: 7\nBogoMIPS: 5999.99\nFlags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon rep_good nopl xtopology nonstop_tsc cpuid aperfmperf tsc_known_freq pni pclmulqdq monitor ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand hypervisor lahf_lm abm 3dnowprefetch pti fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid mpx avx512f avx512dq rdseed adx smap clflushopt clwb avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves ida arat pku ospke\nHypervisor vendor: KVM\nVirtualization type: full\nL1d cache: 1.5 MiB (48 instances)\nL1i cache: 1.5 MiB (48 instances)\nL2 cache: 48 MiB (48 instances)\nL3 cache: 71.5 MiB (2 instances)\nNUMA node(s): 2\nNUMA node0 CPU(s): 0-23,48-71\nNUMA node1 CPU(s): 24-47,72-95\nVulnerability Gather data sampling: Unknown: Dependent on hypervisor status\nVulnerability Indirect target selection: Mitigation; Aligned branch/return thunks\nVulnerability Itlb multihit: KVM: Mitigation: VMX unsupported\nVulnerability L1tf: Mitigation; PTE Inversion\nVulnerability Mds: Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown\nVulnerability Meltdown: Mitigation; PTI\nVulnerability Mmio stale data: Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown\nVulnerability Reg file data sampling: Not affected\nVulnerability Retbleed: Vulnerable\nVulnerability Spec rstack overflow: Not affected\nVulnerability Spec store bypass: Vulnerable\nVulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization\nVulnerability Spectre v2: Mitigation; Retpolines; STIBP disabled; RSB filling; PBRSB-eIBRS Not affected; BHI Retpoline\nVulnerability Srbds: Not affected\nVulnerability Tsa: Not affected\nVulnerability Tsx async abort: Not affected\nVulnerability Vmscape: Not affected\n\nVersions of relevant libraries:\n[pip3] numpy==2.3.4\n[pip3] nvidia-cublas-cu12==12.8.4.1\n[pip3] nvidia-cuda-cupti-cu12==12.8.90\n[pip3] nvidia-cuda-nvrtc-cu12==12.8.93\n[pip3] nvidia-cuda-runtime-cu12==12.8.90\n[pip3] nvidia-cudnn-cu12==9.10.2.21\n[pip3] nvidia-cufft-cu12==11.3.3.83\n[pip3] nvidia-curand-cu12==10.3.9.90\n[pip3] nvidia-cusolver-cu12==11.7.3.90\n[pip3] nvidia-cusparse-cu12==12.5.8.93\n[pip3] nvidia-cusparselt-cu12==0.7.1\n[pip3] nvidia-nccl-cu12==2.27.5\n[pip3] nvidia-nvjitlink-cu12==12.8.93\n[pip3] nvidia-nvtx-cu12==12.8.90\n[pip3] optree==0.17.0\n[pip3] torch==2.9.1+cu128\n[pip3] torchaudio==2.9.1+cu128\n[pip3] torchelastic==0.2.2\n[pip3] torchvision==0.24.1+cu128\n[pip3] triton==3.5.1\n[conda] numpy 2.3.4 py311h2e04523_0 conda-forge\n[conda] nvidia-cublas-cu12 12.8.4.1 pypi_0 pypi\n[conda] nvidia-cuda-cupti-cu12 12.8.90 pypi_0 pypi\n[conda] nvidia-cuda-nvrtc-cu12 12.8.93 pypi_0 pypi\n[conda] nvidia-cuda-runtime-cu12 12.8.90 pypi_0 pypi\n[conda] nvidia-cudnn-cu12 9.10.2.21 pypi_0 pypi\n[conda] nvidia-cufft-cu12 11.3.3.83 pypi_0 pypi\n[conda] nvidia-curand-cu12 10.3.9.90 pypi_0 pypi\n[conda] nvidia-cusolver-cu12 11.7.3.90 pypi_0 pypi\n[conda] nvidia-cusparse-cu12 12.5.8.93 pypi_0 pypi\n[conda] nvidia-cusparselt-cu12 0.7.1 pypi_0 pypi\n[conda] nvidia-nccl-cu12 2.27.5 pypi_0 pypi\n[conda] nvidia-nvjitlink-cu12 12.8.93 pypi_0 pypi\n[conda] nvidia-nvtx-cu12 12.8.90 pypi_0 pypi\n[conda] optree 0.17.0 pypi_0 pypi\n[conda] torch 2.9.1+cu128 pypi_0 pypi\n[conda] torchaudio 2.9.1+cu128 pypi_0 pypi\n[conda] torchelastic 0.2.2 pypi_0 pypi\n[conda] torchvision 0.24.1+cu128 pypi_0 pypi\n[conda] triton 3.5.1 pypi_0 pypi",
786
+ "transformers_version": "4.54.0",
787
+ "lm_eval_version": "0.4.11",
788
+ "upper_git_hash": null,
789
+ "tokenizer_pad_token": [
790
+ "<|endoftext|>",
791
+ "151643"
792
+ ],
793
+ "tokenizer_eos_token": [
794
+ "<|im_end|>",
795
+ "151645"
796
+ ],
797
+ "tokenizer_bos_token": [
798
+ null,
799
+ "None"
800
+ ],
801
+ "eot_token_id": 151645,
802
+ "max_length": 16384,
803
+ "task_hashes": {
804
+ "ruler_qa_hotpot": "fd48cb605efc8c5adac7e3771533db4d4c354b71bf91ce00841c814604ba6404",
805
+ "ruler_fwe": "c4633bcd0acf3ba4b96f905b5ba269ba94cf45a4c8ca0c13aae1468227cc3a10",
806
+ "ruler_vt": "af011da859c8b86f63f57f6256cf59077d3d3d8536084ab67ccf39e18e9a615a",
807
+ "niah_multiquery": "f9a36d2854dfd2bb52ffc7a6bf3f53490fb5303ca9da0046256bc28b55c2adb2",
808
+ "niah_multikey_2": "86e49b8f875843454b2923143b9124ba82e8a2e23b145f792726cd82f4e49c24",
809
+ "niah_single_3": "905e40146a287e5f0c8e2a81f64f082108c45a23862949e6a99bfcbbf561cad2",
810
+ "niah_single_1": "426e8138e992480aa416d01e6b9cef5865a68c00f532400e448a0c20ee026cff"
811
+ },
812
+ "model_source": "hf",
813
+ "model_name": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
814
+ "model_name_sanitized": "__workspace__outputs__l2a_style__stage2__checkpoint-25",
815
+ "system_instruction": null,
816
+ "system_instruction_sha": null,
817
+ "fewshot_as_multiturn": null,
818
+ "chat_template": null,
819
+ "chat_template_sha": null,
820
+ "total_evaluation_time_seconds": "739.5297531939868"
821
+ }
outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multikey_2_2026-07-18T15-05-41.414854.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ec7ad717b95e4d5fa923a2627ae2eaf57f1f9db47d459d64eb1ce5f9d6647f8c
3
+ size 987526
outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multiquery_2026-07-18T15-05-41.414854.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:dac9a6008ea1a5889eaeb1f76cacfe65c4c4c5a7abd6f5608b3dadddaa221648
3
+ size 1440697
outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_single_1_2026-07-18T15-05-41.414854.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b01ced4cc6722de3a956473b8ae0600525ca6e5485940b5a5e2267633cc3bd5d
3
+ size 1231588
outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_single_3_2026-07-18T15-05-41.414854.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f20ebd346ce5932c6d4cbcdf5f0330574b9b13b9b52951d58b08d2b974886e0c
3
+ size 1435128
outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_fwe_2026-07-18T15-05-41.414854.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:84bda7768fbe2ab66d514e01fa24e4dadff944d9c721e69274765df0bcdd8864
3
+ size 827295
outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_qa_hotpot_2026-07-18T15-05-41.414854.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2541d5179a144eda9e0c39f423f5ec895fad28e7188eeadae9317754b88954f1
3
+ size 1172252
outputs/eval/token_t045/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_vt_2026-07-18T15-05-41.414854.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b7c6124651beb57d18609eccd0f1db78d94c9b96851d2e0759422a5477af8a06
3
+ size 1228994
outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/results_2026-07-18T15-17-52.616068.json ADDED
@@ -0,0 +1,714 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "results": {
3
+ "niah_multikey_1": {
4
+ "alias": "niah_multikey_1",
5
+ "4096,none": -1,
6
+ "4096_stderr,none": "N/A",
7
+ "8192,none": 1.0,
8
+ "8192_stderr,none": "N/A"
9
+ },
10
+ "niah_multikey_3": {
11
+ "alias": "niah_multikey_3",
12
+ "4096,none": -1,
13
+ "4096_stderr,none": "N/A",
14
+ "8192,none": 0.75,
15
+ "8192_stderr,none": "N/A"
16
+ },
17
+ "niah_multivalue": {
18
+ "alias": "niah_multivalue",
19
+ "4096,none": -1,
20
+ "4096_stderr,none": "N/A",
21
+ "8192,none": 1.0,
22
+ "8192_stderr,none": "N/A"
23
+ },
24
+ "niah_single_2": {
25
+ "alias": "niah_single_2",
26
+ "4096,none": -1,
27
+ "4096_stderr,none": "N/A",
28
+ "8192,none": 1.0,
29
+ "8192_stderr,none": "N/A"
30
+ },
31
+ "ruler_cwe": {
32
+ "alias": "ruler_cwe",
33
+ "4096,none": -1,
34
+ "4096_stderr,none": "N/A",
35
+ "8192,none": 0.3949999999999999,
36
+ "8192_stderr,none": "N/A"
37
+ },
38
+ "ruler_qa_squad": {
39
+ "alias": "ruler_qa_squad",
40
+ "4096,none": -1,
41
+ "4096_stderr,none": "N/A",
42
+ "8192,none": 0.4208333333333334,
43
+ "8192_stderr,none": "N/A"
44
+ }
45
+ },
46
+ "group_subtasks": {
47
+ "niah_single_2": [],
48
+ "niah_multikey_1": [],
49
+ "niah_multikey_3": [],
50
+ "niah_multivalue": [],
51
+ "ruler_cwe": [],
52
+ "ruler_qa_squad": []
53
+ },
54
+ "configs": {
55
+ "niah_multikey_1": {
56
+ "task": "niah_multikey_1",
57
+ "tag": [
58
+ "longcxt"
59
+ ],
60
+ "custom_dataset": "def niah_multikey_1(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"essay\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"essay\",\n type_needle_k=\"words\",\n type_needle_v=\"numbers\",\n num_needle_k=4,\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
61
+ "dataset_path": "",
62
+ "dataset_name": "",
63
+ "test_split": "test",
64
+ "doc_to_text": "{{input}}",
65
+ "doc_to_target": "{{outputs}}",
66
+ "unsafe_code": false,
67
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
68
+ "description": "",
69
+ "target_delimiter": " ",
70
+ "fewshot_delimiter": "\n\n",
71
+ "fewshot_config": {
72
+ "sampler": "default",
73
+ "split": null,
74
+ "process_docs": null,
75
+ "fewshot_indices": null,
76
+ "samples": null,
77
+ "doc_to_text": "{{input}}",
78
+ "doc_to_choice": null,
79
+ "doc_to_target": "{{outputs}}",
80
+ "gen_prefix": "{{gen_prefix}}",
81
+ "fewshot_delimiter": "\n\n",
82
+ "target_delimiter": " "
83
+ },
84
+ "num_fewshot": 0,
85
+ "metric_list": [
86
+ {
87
+ "metric": "4096",
88
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
89
+ "higher_is_better": true
90
+ },
91
+ {
92
+ "metric": "8192",
93
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
94
+ "higher_is_better": true
95
+ },
96
+ {
97
+ "metric": "16384",
98
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
99
+ "higher_is_better": true
100
+ },
101
+ {
102
+ "metric": "32768",
103
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
104
+ "higher_is_better": true
105
+ },
106
+ {
107
+ "metric": "65536",
108
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
109
+ "higher_is_better": true
110
+ },
111
+ {
112
+ "metric": "131072",
113
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
114
+ "higher_is_better": true
115
+ }
116
+ ],
117
+ "output_type": "generate_until",
118
+ "generation_kwargs": {
119
+ "do_sample": false,
120
+ "temperature": 0.0,
121
+ "max_gen_toks": 128,
122
+ "until": []
123
+ },
124
+ "repeats": 1,
125
+ "should_decontaminate": false,
126
+ "gen_prefix": "{{gen_prefix}}",
127
+ "metadata": {
128
+ "version": 1.0,
129
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
130
+ "trust_remote_code": true,
131
+ "dtype": "bfloat16",
132
+ "max_length": 16384,
133
+ "attn_implementation": "sdpa",
134
+ "max_seq_lengths": [
135
+ 8192
136
+ ]
137
+ }
138
+ },
139
+ "niah_multikey_3": {
140
+ "task": "niah_multikey_3",
141
+ "tag": [
142
+ "longcxt"
143
+ ],
144
+ "custom_dataset": "def niah_multikey_3(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"needle\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"needle\",\n type_needle_k=\"uuids\",\n type_needle_v=\"uuids\",\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
145
+ "dataset_path": "",
146
+ "dataset_name": "",
147
+ "test_split": "test",
148
+ "doc_to_text": "{{input}}",
149
+ "doc_to_target": "{{outputs}}",
150
+ "unsafe_code": false,
151
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
152
+ "description": "",
153
+ "target_delimiter": " ",
154
+ "fewshot_delimiter": "\n\n",
155
+ "fewshot_config": {
156
+ "sampler": "default",
157
+ "split": null,
158
+ "process_docs": null,
159
+ "fewshot_indices": null,
160
+ "samples": null,
161
+ "doc_to_text": "{{input}}",
162
+ "doc_to_choice": null,
163
+ "doc_to_target": "{{outputs}}",
164
+ "gen_prefix": "{{gen_prefix}}",
165
+ "fewshot_delimiter": "\n\n",
166
+ "target_delimiter": " "
167
+ },
168
+ "num_fewshot": 0,
169
+ "metric_list": [
170
+ {
171
+ "metric": "4096",
172
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
173
+ "higher_is_better": true
174
+ },
175
+ {
176
+ "metric": "8192",
177
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
178
+ "higher_is_better": true
179
+ },
180
+ {
181
+ "metric": "16384",
182
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
183
+ "higher_is_better": true
184
+ },
185
+ {
186
+ "metric": "32768",
187
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
188
+ "higher_is_better": true
189
+ },
190
+ {
191
+ "metric": "65536",
192
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
193
+ "higher_is_better": true
194
+ },
195
+ {
196
+ "metric": "131072",
197
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
198
+ "higher_is_better": true
199
+ }
200
+ ],
201
+ "output_type": "generate_until",
202
+ "generation_kwargs": {
203
+ "do_sample": false,
204
+ "temperature": 0.0,
205
+ "max_gen_toks": 128,
206
+ "until": []
207
+ },
208
+ "repeats": 1,
209
+ "should_decontaminate": false,
210
+ "gen_prefix": "{{gen_prefix}}",
211
+ "metadata": {
212
+ "version": 1.0,
213
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
214
+ "trust_remote_code": true,
215
+ "dtype": "bfloat16",
216
+ "max_length": 16384,
217
+ "attn_implementation": "sdpa",
218
+ "max_seq_lengths": [
219
+ 8192
220
+ ]
221
+ }
222
+ },
223
+ "niah_multivalue": {
224
+ "task": "niah_multivalue",
225
+ "tag": [
226
+ "longcxt"
227
+ ],
228
+ "custom_dataset": "def niah_multivalue(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"essay\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"essay\",\n type_needle_k=\"words\",\n type_needle_v=\"numbers\",\n num_needle_v=4,\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
229
+ "dataset_path": "",
230
+ "dataset_name": "",
231
+ "test_split": "test",
232
+ "doc_to_text": "{{input}}",
233
+ "doc_to_target": "{{outputs}}",
234
+ "unsafe_code": false,
235
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
236
+ "description": "",
237
+ "target_delimiter": " ",
238
+ "fewshot_delimiter": "\n\n",
239
+ "fewshot_config": {
240
+ "sampler": "default",
241
+ "split": null,
242
+ "process_docs": null,
243
+ "fewshot_indices": null,
244
+ "samples": null,
245
+ "doc_to_text": "{{input}}",
246
+ "doc_to_choice": null,
247
+ "doc_to_target": "{{outputs}}",
248
+ "gen_prefix": "{{gen_prefix}}",
249
+ "fewshot_delimiter": "\n\n",
250
+ "target_delimiter": " "
251
+ },
252
+ "num_fewshot": 0,
253
+ "metric_list": [
254
+ {
255
+ "metric": "4096",
256
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
257
+ "higher_is_better": true
258
+ },
259
+ {
260
+ "metric": "8192",
261
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
262
+ "higher_is_better": true
263
+ },
264
+ {
265
+ "metric": "16384",
266
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
267
+ "higher_is_better": true
268
+ },
269
+ {
270
+ "metric": "32768",
271
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
272
+ "higher_is_better": true
273
+ },
274
+ {
275
+ "metric": "65536",
276
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
277
+ "higher_is_better": true
278
+ },
279
+ {
280
+ "metric": "131072",
281
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
282
+ "higher_is_better": true
283
+ }
284
+ ],
285
+ "output_type": "generate_until",
286
+ "generation_kwargs": {
287
+ "do_sample": false,
288
+ "temperature": 0.0,
289
+ "max_gen_toks": 128,
290
+ "until": []
291
+ },
292
+ "repeats": 1,
293
+ "should_decontaminate": false,
294
+ "gen_prefix": "{{gen_prefix}}",
295
+ "metadata": {
296
+ "version": 1.0,
297
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
298
+ "trust_remote_code": true,
299
+ "dtype": "bfloat16",
300
+ "max_length": 16384,
301
+ "attn_implementation": "sdpa",
302
+ "max_seq_lengths": [
303
+ 8192
304
+ ]
305
+ }
306
+ },
307
+ "niah_single_2": {
308
+ "task": "niah_single_2",
309
+ "tag": [
310
+ "longcxt"
311
+ ],
312
+ "custom_dataset": "def niah_single_2(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"essay\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"essay\",\n type_needle_k=\"words\",\n type_needle_v=\"numbers\",\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
313
+ "dataset_path": "",
314
+ "dataset_name": "",
315
+ "test_split": "test",
316
+ "doc_to_text": "{{input}}",
317
+ "doc_to_target": "{{outputs}}",
318
+ "unsafe_code": false,
319
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
320
+ "description": "",
321
+ "target_delimiter": " ",
322
+ "fewshot_delimiter": "\n\n",
323
+ "fewshot_config": {
324
+ "sampler": "default",
325
+ "split": null,
326
+ "process_docs": null,
327
+ "fewshot_indices": null,
328
+ "samples": null,
329
+ "doc_to_text": "{{input}}",
330
+ "doc_to_choice": null,
331
+ "doc_to_target": "{{outputs}}",
332
+ "gen_prefix": "{{gen_prefix}}",
333
+ "fewshot_delimiter": "\n\n",
334
+ "target_delimiter": " "
335
+ },
336
+ "num_fewshot": 0,
337
+ "metric_list": [
338
+ {
339
+ "metric": "4096",
340
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
341
+ "higher_is_better": true
342
+ },
343
+ {
344
+ "metric": "8192",
345
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
346
+ "higher_is_better": true
347
+ },
348
+ {
349
+ "metric": "16384",
350
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
351
+ "higher_is_better": true
352
+ },
353
+ {
354
+ "metric": "32768",
355
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
356
+ "higher_is_better": true
357
+ },
358
+ {
359
+ "metric": "65536",
360
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
361
+ "higher_is_better": true
362
+ },
363
+ {
364
+ "metric": "131072",
365
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
366
+ "higher_is_better": true
367
+ }
368
+ ],
369
+ "output_type": "generate_until",
370
+ "generation_kwargs": {
371
+ "do_sample": false,
372
+ "temperature": 0.0,
373
+ "max_gen_toks": 128,
374
+ "until": []
375
+ },
376
+ "repeats": 1,
377
+ "should_decontaminate": false,
378
+ "gen_prefix": "{{gen_prefix}}",
379
+ "metadata": {
380
+ "version": 1.0,
381
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
382
+ "trust_remote_code": true,
383
+ "dtype": "bfloat16",
384
+ "max_length": 16384,
385
+ "attn_implementation": "sdpa",
386
+ "max_seq_lengths": [
387
+ 8192
388
+ ]
389
+ }
390
+ },
391
+ "ruler_cwe": {
392
+ "task": "ruler_cwe",
393
+ "tag": [
394
+ "longcxt"
395
+ ],
396
+ "custom_dataset": "def get_cw_dataset(**kwargs):\n pretrained = kwargs.get(\"tokenizer\", kwargs.get(\"pretrained\", {}))\n df = (\n get_dataset(pretrained, seq=seq)\n for seq in kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n )\n\n return {\n \"test\": datasets.Dataset.from_list(\n list(itertools.chain.from_iterable(df)), split=datasets.Split.TEST\n )\n }\n",
397
+ "dataset_path": "",
398
+ "dataset_name": "",
399
+ "test_split": "test",
400
+ "doc_to_text": "{{input}}",
401
+ "doc_to_target": "{{outputs}}",
402
+ "unsafe_code": false,
403
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
404
+ "description": "",
405
+ "target_delimiter": "\n\n",
406
+ "fewshot_delimiter": "\n\n",
407
+ "fewshot_config": {
408
+ "sampler": "default",
409
+ "split": null,
410
+ "process_docs": null,
411
+ "fewshot_indices": null,
412
+ "samples": null,
413
+ "doc_to_text": "{{input}}",
414
+ "doc_to_choice": null,
415
+ "doc_to_target": "{{outputs}}",
416
+ "gen_prefix": "{{gen_prefix}}",
417
+ "fewshot_delimiter": "\n\n",
418
+ "target_delimiter": "\n\n"
419
+ },
420
+ "num_fewshot": 0,
421
+ "metric_list": [
422
+ {
423
+ "metric": "4096",
424
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
425
+ "higher_is_better": true
426
+ },
427
+ {
428
+ "metric": "8192",
429
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
430
+ "higher_is_better": true
431
+ },
432
+ {
433
+ "metric": "16384",
434
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
435
+ "higher_is_better": true
436
+ },
437
+ {
438
+ "metric": "32768",
439
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
440
+ "higher_is_better": true
441
+ },
442
+ {
443
+ "metric": "65536",
444
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
445
+ "higher_is_better": true
446
+ },
447
+ {
448
+ "metric": "131072",
449
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
450
+ "higher_is_better": true
451
+ }
452
+ ],
453
+ "output_type": "generate_until",
454
+ "generation_kwargs": {
455
+ "do_sample": false,
456
+ "temperature": 0.0,
457
+ "max_gen_toks": 120,
458
+ "until": []
459
+ },
460
+ "repeats": 1,
461
+ "should_decontaminate": false,
462
+ "gen_prefix": "{{gen_prefix}}",
463
+ "metadata": {
464
+ "version": 1.0,
465
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
466
+ "trust_remote_code": true,
467
+ "dtype": "bfloat16",
468
+ "max_length": 16384,
469
+ "attn_implementation": "sdpa",
470
+ "max_seq_lengths": [
471
+ 8192
472
+ ]
473
+ }
474
+ },
475
+ "ruler_qa_squad": {
476
+ "task": "ruler_qa_squad",
477
+ "tag": [
478
+ "longcxt"
479
+ ],
480
+ "custom_dataset": "def get_squad(**kwargs):\n return get_qa_dataset(\"squad\", **kwargs)\n",
481
+ "dataset_path": "",
482
+ "dataset_name": "",
483
+ "test_split": "test",
484
+ "doc_to_text": "{{input}}",
485
+ "doc_to_target": "{{outputs}}",
486
+ "unsafe_code": false,
487
+ "process_results": "def process_results_part(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_part(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
488
+ "description": "",
489
+ "target_delimiter": " ",
490
+ "fewshot_delimiter": "\n\n",
491
+ "fewshot_config": {
492
+ "sampler": "default",
493
+ "split": null,
494
+ "process_docs": null,
495
+ "fewshot_indices": null,
496
+ "samples": null,
497
+ "doc_to_text": "{{input}}",
498
+ "doc_to_choice": null,
499
+ "doc_to_target": "{{outputs}}",
500
+ "gen_prefix": "{{gen_prefix}}",
501
+ "fewshot_delimiter": "\n\n",
502
+ "target_delimiter": " "
503
+ },
504
+ "num_fewshot": 0,
505
+ "metric_list": [
506
+ {
507
+ "metric": "4096",
508
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
509
+ "higher_is_better": true
510
+ },
511
+ {
512
+ "metric": "8192",
513
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
514
+ "higher_is_better": true
515
+ },
516
+ {
517
+ "metric": "16384",
518
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
519
+ "higher_is_better": true
520
+ },
521
+ {
522
+ "metric": "32768",
523
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
524
+ "higher_is_better": true
525
+ },
526
+ {
527
+ "metric": "65536",
528
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
529
+ "higher_is_better": true
530
+ },
531
+ {
532
+ "metric": "131072",
533
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
534
+ "higher_is_better": true
535
+ }
536
+ ],
537
+ "output_type": "generate_until",
538
+ "generation_kwargs": {
539
+ "do_sample": false,
540
+ "temperature": 0.0,
541
+ "max_gen_toks": 32,
542
+ "until": []
543
+ },
544
+ "repeats": 1,
545
+ "should_decontaminate": false,
546
+ "gen_prefix": "{{gen_prefix}}",
547
+ "metadata": {
548
+ "version": 1.0,
549
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
550
+ "trust_remote_code": true,
551
+ "dtype": "bfloat16",
552
+ "max_length": 16384,
553
+ "attn_implementation": "sdpa",
554
+ "max_seq_lengths": [
555
+ 8192
556
+ ]
557
+ }
558
+ }
559
+ },
560
+ "versions": {
561
+ "niah_multikey_1": 1.0,
562
+ "niah_multikey_3": 1.0,
563
+ "niah_multivalue": 1.0,
564
+ "niah_single_2": 1.0,
565
+ "ruler_cwe": 1.0,
566
+ "ruler_qa_squad": 1.0
567
+ },
568
+ "n-shot": {
569
+ "niah_multikey_1": 0,
570
+ "niah_multikey_3": 0,
571
+ "niah_multivalue": 0,
572
+ "niah_single_2": 0,
573
+ "ruler_cwe": 0,
574
+ "ruler_qa_squad": 0
575
+ },
576
+ "higher_is_better": {
577
+ "niah_multikey_1": {
578
+ "4096": true,
579
+ "8192": true,
580
+ "16384": true,
581
+ "32768": true,
582
+ "65536": true,
583
+ "131072": true
584
+ },
585
+ "niah_multikey_3": {
586
+ "4096": true,
587
+ "8192": true,
588
+ "16384": true,
589
+ "32768": true,
590
+ "65536": true,
591
+ "131072": true
592
+ },
593
+ "niah_multivalue": {
594
+ "4096": true,
595
+ "8192": true,
596
+ "16384": true,
597
+ "32768": true,
598
+ "65536": true,
599
+ "131072": true
600
+ },
601
+ "niah_single_2": {
602
+ "4096": true,
603
+ "8192": true,
604
+ "16384": true,
605
+ "32768": true,
606
+ "65536": true,
607
+ "131072": true
608
+ },
609
+ "ruler_cwe": {
610
+ "4096": true,
611
+ "8192": true,
612
+ "16384": true,
613
+ "32768": true,
614
+ "65536": true,
615
+ "131072": true
616
+ },
617
+ "ruler_qa_squad": {
618
+ "4096": true,
619
+ "8192": true,
620
+ "16384": true,
621
+ "32768": true,
622
+ "65536": true,
623
+ "131072": true
624
+ }
625
+ },
626
+ "n-samples": {
627
+ "ruler_qa_squad": {
628
+ "original": 500,
629
+ "effective": 20
630
+ },
631
+ "ruler_cwe": {
632
+ "original": 500,
633
+ "effective": 20
634
+ },
635
+ "niah_multivalue": {
636
+ "original": 500,
637
+ "effective": 20
638
+ },
639
+ "niah_multikey_3": {
640
+ "original": 500,
641
+ "effective": 20
642
+ },
643
+ "niah_multikey_1": {
644
+ "original": 500,
645
+ "effective": 20
646
+ },
647
+ "niah_single_2": {
648
+ "original": 500,
649
+ "effective": 20
650
+ }
651
+ },
652
+ "config": {
653
+ "model": "hf",
654
+ "model_args": {
655
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
656
+ "trust_remote_code": true,
657
+ "dtype": "bfloat16",
658
+ "max_length": 16384,
659
+ "attn_implementation": "sdpa"
660
+ },
661
+ "model_num_parameters": 2031969308,
662
+ "model_dtype": "torch.bfloat16",
663
+ "model_revision": "main",
664
+ "model_sha": "",
665
+ "batch_size": "1",
666
+ "batch_sizes": [],
667
+ "device": "cuda:0",
668
+ "use_cache": null,
669
+ "limit": 20.0,
670
+ "bootstrap_iters": 100000,
671
+ "gen_kwargs": {},
672
+ "random_seed": 0,
673
+ "numpy_seed": 1234,
674
+ "torch_seed": 1234,
675
+ "fewshot_seed": 1234
676
+ },
677
+ "git_hash": null,
678
+ "date": 1784387147.8366756,
679
+ "pretty_env_info": "PyTorch version: 2.9.1+cu128\nIs debug build: False\nCUDA used to build PyTorch: 12.8\nROCM used to build PyTorch: N/A\n\nOS: Ubuntu 22.04.5 LTS (x86_64)\nGCC version: Could not collect\nClang version: Could not collect\nCMake version: version 4.1.2\nLibc version: glibc-2.35\n\nPython version: 3.11.14 | packaged by conda-forge | (main, Oct 22 2025, 22:46:25) [GCC 14.3.0] (64-bit runtime)\nPython platform: Linux-6.12.90-120.164.amzn2023.x86_64-x86_64-with-glibc2.35\nIs CUDA available: True\nCUDA runtime version: Could not collect\nCUDA_MODULE_LOADING set to: \nGPU models and configuration: GPU 0: NVIDIA A100-SXM4-80GB\nNvidia driver version: 580.159.03\ncuDNN version: Could not collect\nIs XPU available: False\nHIP runtime version: N/A\nMIOpen runtime version: N/A\nIs XNNPACK available: True\n\nCPU:\nArchitecture: x86_64\nCPU op-mode(s): 32-bit, 64-bit\nAddress sizes: 46 bits physical, 48 bits virtual\nByte Order: Little Endian\nCPU(s): 96\nOn-line CPU(s) list: 0-95\nVendor ID: GenuineIntel\nModel name: Intel(R) Xeon(R) Platinum 8275CL CPU @ 3.00GHz\nCPU family: 6\nModel: 85\nThread(s) per core: 2\nCore(s) per socket: 24\nSocket(s): 2\nStepping: 7\nBogoMIPS: 5999.99\nFlags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon rep_good nopl xtopology nonstop_tsc cpuid aperfmperf tsc_known_freq pni pclmulqdq monitor ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand hypervisor lahf_lm abm 3dnowprefetch pti fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid mpx avx512f avx512dq rdseed adx smap clflushopt clwb avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves ida arat pku ospke\nHypervisor vendor: KVM\nVirtualization type: full\nL1d cache: 1.5 MiB (48 instances)\nL1i cache: 1.5 MiB (48 instances)\nL2 cache: 48 MiB (48 instances)\nL3 cache: 71.5 MiB (2 instances)\nNUMA node(s): 2\nNUMA node0 CPU(s): 0-23,48-71\nNUMA node1 CPU(s): 24-47,72-95\nVulnerability Gather data sampling: Unknown: Dependent on hypervisor status\nVulnerability Indirect target selection: Mitigation; Aligned branch/return thunks\nVulnerability Itlb multihit: KVM: Mitigation: VMX unsupported\nVulnerability L1tf: Mitigation; PTE Inversion\nVulnerability Mds: Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown\nVulnerability Meltdown: Mitigation; PTI\nVulnerability Mmio stale data: Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown\nVulnerability Reg file data sampling: Not affected\nVulnerability Retbleed: Vulnerable\nVulnerability Spec rstack overflow: Not affected\nVulnerability Spec store bypass: Vulnerable\nVulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization\nVulnerability Spectre v2: Mitigation; Retpolines; STIBP disabled; RSB filling; PBRSB-eIBRS Not affected; BHI Retpoline\nVulnerability Srbds: Not affected\nVulnerability Tsa: Not affected\nVulnerability Tsx async abort: Not affected\nVulnerability Vmscape: Not affected\n\nVersions of relevant libraries:\n[pip3] numpy==2.3.4\n[pip3] nvidia-cublas-cu12==12.8.4.1\n[pip3] nvidia-cuda-cupti-cu12==12.8.90\n[pip3] nvidia-cuda-nvrtc-cu12==12.8.93\n[pip3] nvidia-cuda-runtime-cu12==12.8.90\n[pip3] nvidia-cudnn-cu12==9.10.2.21\n[pip3] nvidia-cufft-cu12==11.3.3.83\n[pip3] nvidia-curand-cu12==10.3.9.90\n[pip3] nvidia-cusolver-cu12==11.7.3.90\n[pip3] nvidia-cusparse-cu12==12.5.8.93\n[pip3] nvidia-cusparselt-cu12==0.7.1\n[pip3] nvidia-nccl-cu12==2.27.5\n[pip3] nvidia-nvjitlink-cu12==12.8.93\n[pip3] nvidia-nvtx-cu12==12.8.90\n[pip3] optree==0.17.0\n[pip3] torch==2.9.1+cu128\n[pip3] torchaudio==2.9.1+cu128\n[pip3] torchelastic==0.2.2\n[pip3] torchvision==0.24.1+cu128\n[pip3] triton==3.5.1\n[conda] numpy 2.3.4 py311h2e04523_0 conda-forge\n[conda] nvidia-cublas-cu12 12.8.4.1 pypi_0 pypi\n[conda] nvidia-cuda-cupti-cu12 12.8.90 pypi_0 pypi\n[conda] nvidia-cuda-nvrtc-cu12 12.8.93 pypi_0 pypi\n[conda] nvidia-cuda-runtime-cu12 12.8.90 pypi_0 pypi\n[conda] nvidia-cudnn-cu12 9.10.2.21 pypi_0 pypi\n[conda] nvidia-cufft-cu12 11.3.3.83 pypi_0 pypi\n[conda] nvidia-curand-cu12 10.3.9.90 pypi_0 pypi\n[conda] nvidia-cusolver-cu12 11.7.3.90 pypi_0 pypi\n[conda] nvidia-cusparse-cu12 12.5.8.93 pypi_0 pypi\n[conda] nvidia-cusparselt-cu12 0.7.1 pypi_0 pypi\n[conda] nvidia-nccl-cu12 2.27.5 pypi_0 pypi\n[conda] nvidia-nvjitlink-cu12 12.8.93 pypi_0 pypi\n[conda] nvidia-nvtx-cu12 12.8.90 pypi_0 pypi\n[conda] optree 0.17.0 pypi_0 pypi\n[conda] torch 2.9.1+cu128 pypi_0 pypi\n[conda] torchaudio 2.9.1+cu128 pypi_0 pypi\n[conda] torchelastic 0.2.2 pypi_0 pypi\n[conda] torchvision 0.24.1+cu128 pypi_0 pypi\n[conda] triton 3.5.1 pypi_0 pypi",
680
+ "transformers_version": "4.54.0",
681
+ "lm_eval_version": "0.4.11",
682
+ "upper_git_hash": null,
683
+ "tokenizer_pad_token": [
684
+ "<|endoftext|>",
685
+ "151643"
686
+ ],
687
+ "tokenizer_eos_token": [
688
+ "<|im_end|>",
689
+ "151645"
690
+ ],
691
+ "tokenizer_bos_token": [
692
+ null,
693
+ "None"
694
+ ],
695
+ "eot_token_id": 151645,
696
+ "max_length": 16384,
697
+ "task_hashes": {
698
+ "ruler_qa_squad": "8c658a728ccd3d67ddd3f45e8e76527cf15a1ae4119033d8246f293ecc68cf11",
699
+ "ruler_cwe": "f68206129fc7d709951d550ec9269fed05e73bea2fc86ce9f2c758974db5db71",
700
+ "niah_multivalue": "29646655c230b4f155a8de1fbc7ddf22c49e43045bd8fdab5770e7cb45026c7c",
701
+ "niah_multikey_3": "de2cec52885e38d56e0bda0e344e989909969c71ab5f8048b9368a011dd3616b",
702
+ "niah_multikey_1": "338f81e9efa98dbe6979d4f178f7f06cdbd6d367072318667a932b43c5032ab8",
703
+ "niah_single_2": "88ba10e924c2fc6c47543eb646b9e6f7487c7ca026c9391d6cd7054fde59eaad"
704
+ },
705
+ "model_source": "hf",
706
+ "model_name": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
707
+ "model_name_sanitized": "__workspace__outputs__l2a_style__stage2__checkpoint-25",
708
+ "system_instruction": null,
709
+ "system_instruction_sha": null,
710
+ "fewshot_as_multiturn": null,
711
+ "chat_template": null,
712
+ "chat_template_sha": null,
713
+ "total_evaluation_time_seconds": "728.001551622001"
714
+ }
outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multikey_1_2026-07-18T15-17-52.616068.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a33e240ce70e10445a34fea2d42efc0519a02e24a140de317904411fc5f1ef03
3
+ size 1445444
outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multikey_3_2026-07-18T15-17-52.616068.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6d2e7da44df1fc5f05d1a36a2576a831b9b7488a08a2a3749c885bbf1d34d6cd
3
+ size 502030
outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multivalue_2026-07-18T15-17-52.616068.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6bb6d9fdfc2c27b4b41ba425c1be458b8c3adcc4bf7d376447622c3772e7ffbb
3
+ size 1433534
outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_single_2_2026-07-18T15-17-52.616068.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c9f28fb7ee751ed53efff7df7ca83a1a5924c3ab6903c091ebc104c50ab765d5
3
+ size 1438222
outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_cwe_2026-07-18T15-17-52.616068.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6a5399a30e42235fe359c466f8b5a56f2ab4da5a81b78fadb4b9c8e353308fb1
3
+ size 635553
outputs/eval/token_t045/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_qa_squad_2026-07-18T15-17-52.616068.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c072e0fc458030ab99bf9a9cf9443b4393f6b410ef501be053bdc7be752c60c8
3
+ size 1387270