w-ahmad commited on
Commit
f4022be
Β·
verified Β·
1 Parent(s): e24420f

Auto upload zain 2026-08-19T17:15:02.582332 (part 3)

Browse files
zain/Activation/wandb/run-20260819_171215-2nil4rqn/files/config.yaml ADDED
@@ -0,0 +1,435 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ _name_or_path:
2
+ value: ""
3
+ _wandb:
4
+ value:
5
+ cli_version: 0.28.1
6
+ e:
7
+ rjsr5aiivqlwnjxtkxlgjkrgdfxgvs5f:
8
+ args:
9
+ - --config
10
+ - /mnt/data/zainulabideen/zain-exp/notebooks/zain/Activation/configs/baseline100L.yaml
11
+ - --variants
12
+ - mlp-linear-3L
13
+ codePath: sweep.py
14
+ codePathLocal: sweep.py
15
+ cpu_count: 112
16
+ cpu_count_logical: 224
17
+ cudaVersion: "12.4"
18
+ disk:
19
+ /:
20
+ total: "1560765693952"
21
+ used: "708477448192"
22
+ email: deepnevro@gmail.com
23
+ executable: /mnt/data/zainulabideen/zain-exp/notebooks/my_env/bin/python
24
+ git:
25
+ commit: 463c9961366755fc55f02df9a0d471b3cbcb025e
26
+ remote: https://github.com/w-ahmad1a10/Activation.git
27
+ gpu: NVIDIA H100 80GB HBM3
28
+ gpu_count: 8
29
+ gpu_nvidia:
30
+ - architecture: Hopper
31
+ cudaCores: 16896
32
+ memoryTotal: "85520809984"
33
+ name: NVIDIA H100 80GB HBM3
34
+ uuid: GPU-39c684a5-fde6-83d7-1663-0859795881ae
35
+ - architecture: Hopper
36
+ cudaCores: 16896
37
+ memoryTotal: "85520809984"
38
+ name: NVIDIA H100 80GB HBM3
39
+ uuid: GPU-68012e5a-38b6-b643-0ca6-62fb66720bf3
40
+ - architecture: Hopper
41
+ cudaCores: 16896
42
+ memoryTotal: "85520809984"
43
+ name: NVIDIA H100 80GB HBM3
44
+ uuid: GPU-132944c4-b689-2b5f-89a4-d730401677ab
45
+ - architecture: Hopper
46
+ cudaCores: 16896
47
+ memoryTotal: "85520809984"
48
+ name: NVIDIA H100 80GB HBM3
49
+ uuid: GPU-2df386cc-6d26-d0e2-7a2d-a057b0d95864
50
+ - architecture: Hopper
51
+ cudaCores: 16896
52
+ memoryTotal: "85520809984"
53
+ name: NVIDIA H100 80GB HBM3
54
+ uuid: GPU-bfa16575-1d94-1aa2-4537-2c93433f42ef
55
+ - architecture: Hopper
56
+ cudaCores: 16896
57
+ memoryTotal: "85520809984"
58
+ name: NVIDIA H100 80GB HBM3
59
+ uuid: GPU-bc6c3e3c-9b90-09ca-c034-774961847c54
60
+ - architecture: Hopper
61
+ cudaCores: 16896
62
+ memoryTotal: "85520809984"
63
+ name: NVIDIA H100 80GB HBM3
64
+ uuid: GPU-00a441e1-7c95-e7d6-4c35-43d6b291aea9
65
+ - architecture: Hopper
66
+ cudaCores: 16896
67
+ memoryTotal: "85520809984"
68
+ name: NVIDIA H100 80GB HBM3
69
+ uuid: GPU-1c4d29a2-4647-6fce-d8fc-0c5ecfbbd6ea
70
+ host: deeplens-k3s-node1
71
+ memory:
72
+ total: "2164089937920"
73
+ os: Linux-5.15.0-126-generic-x86_64-with-glibc2.35
74
+ program: /mnt/data/zainulabideen/zain-exp/notebooks/zain/Activation/sweep.py
75
+ python: CPython 3.11.15
76
+ root: /mnt/data/zainulabideen/zain-exp/notebooks/zain/Activation
77
+ startedAt: "2026-08-19T17:12:15.316941Z"
78
+ writerId: rjsr5aiivqlwnjxtkxlgjkrgdfxgvs5f
79
+ m:
80
+ - "1": train/global_step
81
+ "6":
82
+ - 3
83
+ "7": []
84
+ - "2": '*'
85
+ "5": 1
86
+ "6":
87
+ - 1
88
+ "7": []
89
+ python_version: 3.11.15
90
+ t:
91
+ "1":
92
+ - 1
93
+ - 5
94
+ - 11
95
+ - 41
96
+ - 49
97
+ - 51
98
+ - 53
99
+ - 71
100
+ "2":
101
+ - 1
102
+ - 5
103
+ - 11
104
+ - 41
105
+ - 49
106
+ - 51
107
+ - 53
108
+ - 71
109
+ "3":
110
+ - 2
111
+ - 7
112
+ - 13
113
+ - 19
114
+ - 62
115
+ - 66
116
+ "4": 3.11.15
117
+ "5": 0.28.1
118
+ "6": 5.16.0.dev0
119
+ "9":
120
+ "1": transformers_trainer
121
+ "12": 0.28.1
122
+ "13": linux-x86_64
123
+ accelerator_config:
124
+ value:
125
+ dispatch_batches: null
126
+ even_batches: true
127
+ gradient_accumulation_kwargs: null
128
+ non_blocking: false
129
+ split_batches: false
130
+ use_seedable_sampler: true
131
+ activation:
132
+ value: linear
133
+ adam_beta1:
134
+ value: 0.9
135
+ adam_beta2:
136
+ value: 0.999
137
+ adam_epsilon:
138
+ value: 1e-08
139
+ architectures:
140
+ value: null
141
+ attention_bias:
142
+ value: false
143
+ attention_dropout:
144
+ value: 0
145
+ auto_find_batch_size:
146
+ value: false
147
+ average_tokens_across_devices:
148
+ value: true
149
+ batch_eval_metrics:
150
+ value: false
151
+ bf16:
152
+ value: true
153
+ bf16_full_eval:
154
+ value: false
155
+ bos_token_id:
156
+ value: 1
157
+ chunk_size_feed_forward:
158
+ value: 0
159
+ data_seed:
160
+ value: 42
161
+ dataloader_drop_last:
162
+ value: false
163
+ dataloader_in_order:
164
+ value: true
165
+ dataloader_multiprocessing_context:
166
+ value: null
167
+ dataloader_num_workers:
168
+ value: 0
169
+ dataloader_persistent_workers:
170
+ value: false
171
+ dataloader_pin_memory:
172
+ value: true
173
+ dataloader_prefetch_factor:
174
+ value: null
175
+ ddp_backend:
176
+ value: null
177
+ ddp_broadcast_buffers:
178
+ value: null
179
+ ddp_bucket_cap_mb:
180
+ value: null
181
+ ddp_find_unused_parameters:
182
+ value: null
183
+ ddp_static_graph:
184
+ value: null
185
+ ddp_timeout:
186
+ value: 1800
187
+ debug:
188
+ value: []
189
+ deepspeed:
190
+ value: null
191
+ disable_tqdm:
192
+ value: false
193
+ do_eval:
194
+ value: true
195
+ do_predict:
196
+ value: false
197
+ do_train:
198
+ value: false
199
+ dtype:
200
+ value: null
201
+ enable_jit_checkpoint:
202
+ value: false
203
+ eos_token_id:
204
+ value: 2
205
+ eval_accumulation_steps:
206
+ value: null
207
+ eval_delay:
208
+ value: 0
209
+ eval_do_concat_batches:
210
+ value: true
211
+ eval_on_start:
212
+ value: false
213
+ eval_steps:
214
+ value: 100
215
+ eval_strategy:
216
+ value: steps
217
+ eval_use_gather_object:
218
+ value: false
219
+ fp16:
220
+ value: false
221
+ fp16_full_eval:
222
+ value: false
223
+ fsdp:
224
+ value: null
225
+ fsdp_config:
226
+ value: null
227
+ full_determinism:
228
+ value: false
229
+ gradient_accumulation_steps:
230
+ value: 1
231
+ gradient_checkpointing:
232
+ value: false
233
+ gradient_checkpointing_kwargs:
234
+ value: null
235
+ greater_is_better:
236
+ value: null
237
+ head_dim:
238
+ value: 32
239
+ hidden_act:
240
+ value: silu
241
+ hidden_size:
242
+ value: 128
243
+ hub_always_push:
244
+ value: false
245
+ hub_model_id:
246
+ value: w-ahmad/6L-mlp-linear-3L
247
+ hub_private_repo:
248
+ value: null
249
+ hub_revision:
250
+ value: null
251
+ hub_strategy:
252
+ value: every_save
253
+ hub_token:
254
+ value: <HUB_TOKEN>
255
+ id2label:
256
+ value:
257
+ "0": LABEL_0
258
+ "1": LABEL_1
259
+ ignore_data_skip:
260
+ value: false
261
+ include_for_metrics:
262
+ value: []
263
+ include_num_input_tokens_seen:
264
+ value: "no"
265
+ initializer_range:
266
+ value: 0.02
267
+ intermediate_size:
268
+ value: 256
269
+ is_encoder_decoder:
270
+ value: false
271
+ label_names:
272
+ value: null
273
+ label_smoothing_factor:
274
+ value: 0
275
+ label2id:
276
+ value:
277
+ LABEL_0: 0
278
+ LABEL_1: 1
279
+ learning_rate:
280
+ value: 0.001
281
+ length_column_name:
282
+ value: length
283
+ liger_kernel_config:
284
+ value: null
285
+ load_best_model_at_end:
286
+ value: false
287
+ local_rank:
288
+ value: -1
289
+ log_level:
290
+ value: passive
291
+ log_level_replica:
292
+ value: warning
293
+ log_on_each_node:
294
+ value: true
295
+ logging_first_step:
296
+ value: false
297
+ logging_nan_inf_filter:
298
+ value: true
299
+ logging_steps:
300
+ value: 20
301
+ logging_strategy:
302
+ value: steps
303
+ lr_scheduler_kwargs:
304
+ value: null
305
+ lr_scheduler_type:
306
+ value: constant_with_warmup
307
+ max_grad_norm:
308
+ value: 1
309
+ max_position_embeddings:
310
+ value: 512
311
+ max_steps:
312
+ value: 1000
313
+ metric_for_best_model:
314
+ value: null
315
+ mlp_bias:
316
+ value: false
317
+ mlp_type:
318
+ value: mlp
319
+ model/num_parameters:
320
+ value: 1016704
321
+ model_type:
322
+ value: tiny_llama
323
+ neftune_noise_alpha:
324
+ value: null
325
+ num_attention_heads:
326
+ value: 4
327
+ num_hidden_layers:
328
+ value: 3
329
+ num_key_value_heads:
330
+ value: 4
331
+ num_train_epochs:
332
+ value: 1
333
+ optim:
334
+ value: adamw_torch_fused
335
+ optim_args:
336
+ value: null
337
+ optim_target_modules:
338
+ value: null
339
+ output_attentions:
340
+ value: false
341
+ output_dir:
342
+ value: out/mlp-linear-3L_run
343
+ output_hidden_states:
344
+ value: false
345
+ pad_token_id:
346
+ value: 0
347
+ parallelism_config:
348
+ value: null
349
+ per_device_eval_batch_size:
350
+ value: 1500
351
+ per_device_train_batch_size:
352
+ value: 80
353
+ powlu_m:
354
+ value: 3
355
+ prediction_loss_only:
356
+ value: false
357
+ pretraining_tp:
358
+ value: 1
359
+ problem_type:
360
+ value: null
361
+ project:
362
+ value: huggingface
363
+ push_to_hub:
364
+ value: false
365
+ remove_unused_columns:
366
+ value: false
367
+ report_to:
368
+ value:
369
+ - wandb
370
+ restore_callback_states_from_checkpoint:
371
+ value: false
372
+ resume_from_checkpoint:
373
+ value: null
374
+ return_dict:
375
+ value: true
376
+ rms_norm_eps:
377
+ value: 1e-06
378
+ rope_parameters:
379
+ value:
380
+ rope_theta: 10000
381
+ rope_type: default
382
+ run_name:
383
+ value: LM-mlp-linear-3L-1.0M-20260819-171214
384
+ save_on_each_node:
385
+ value: false
386
+ save_only_model:
387
+ value: false
388
+ save_steps:
389
+ value: 100
390
+ save_strategy:
391
+ value: steps
392
+ save_total_limit:
393
+ value: null
394
+ seed:
395
+ value: 42
396
+ skip_memory_metrics:
397
+ value: true
398
+ tf32:
399
+ value: null
400
+ tie_word_embeddings:
401
+ value: true
402
+ tokenizer_name:
403
+ value: w-ahmad/tiny-stories-tokenizer
404
+ torch_compile:
405
+ value: false
406
+ torch_compile_backend:
407
+ value: null
408
+ torch_compile_mode:
409
+ value: null
410
+ torch_empty_cache_steps:
411
+ value: null
412
+ trackio_bucket_id:
413
+ value: null
414
+ trackio_space_id:
415
+ value: null
416
+ trackio_static_space_id:
417
+ value: null
418
+ train_sampling_strategy:
419
+ value: random
420
+ transformers_version:
421
+ value: 5.16.0.dev0
422
+ use_cache:
423
+ value: false
424
+ use_cpu:
425
+ value: false
426
+ use_liger_kernel:
427
+ value: false
428
+ vocab_size:
429
+ value: 4096
430
+ waleed_beta:
431
+ value: 10
432
+ warmup_steps:
433
+ value: 200
434
+ weight_decay:
435
+ value: 0
zain/Activation/wandb/run-20260819_171215-2nil4rqn/files/output.log CHANGED
@@ -88,8 +88,34 @@ Writing model shards: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1/1 [00:00<00:00, 340
88
  - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception).
89
  - If you are not the owner of the model architecture class, please contact the model code owner to update it.
90
  Writing model shards: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1/1 [00:00<00:00, 338.44it/s]
91
- 89%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 889/1000 [02:09<00:08, 13.32it/s]?, ?it/s]
92
  {'loss': '2.927', 'grad_norm': '0.7422', 'learning_rate': '0.001', 'epoch': '0.06909', 'train/total_time_seconds': '14.41', 'train/time_per_step_avg': '0.01651', 'train/epoch_time_elapsed': '124.6', 'train/estimated_remaining_minutes': '0.05273'}
93
  {'loss': '2.923', 'grad_norm': '0.8828', 'learning_rate': '0.001', 'epoch': '0.07078', 'train/total_time_seconds': '14.74', 'train/time_per_step_avg': '0.01652', 'train/epoch_time_elapsed': '126.2', 'train/estimated_remaining_minutes': '0.04681'}
94
  {'loss': '2.909', 'grad_norm': '0.8008', 'learning_rate': '0.001', 'epoch': '0.07246', 'train/total_time_seconds': '15.08', 'train/time_per_step_avg': '0.01655', 'train/epoch_time_elapsed': '127.7', 'train/estimated_remaining_minutes': '0.04091'}
95
  {'loss': '2.907', 'grad_norm': '1.133', 'learning_rate': '0.001', 'epoch': '0.07415', 'train/total_time_seconds': '15.41', 'train/time_per_step_avg': '0.01656', 'train/epoch_time_elapsed': '129.2', 'train/estimated_remaining_minutes': '0.03502'}
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
88
  - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception).
89
  - If you are not the owner of the model architecture class, please contact the model code owner to update it.
90
  Writing model shards: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1/1 [00:00<00:00, 338.44it/s]
91
+ 90%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 900/1000 [02:18<00:07, 13.35[transformers] TinyLlamaForCausalLM has generative capabilities, as `prepare_inputs_for_generation` is explicitly defined. However, it doesn't directly inherit from `GenerationMixin`. From πŸ‘‰v4.50πŸ‘ˆ onwards, `PreTrainedModel` will NOT inherit from `GenerationMixin`, and this model will lose the ability to call `generate` and other related functions.
92
  {'loss': '2.927', 'grad_norm': '0.7422', 'learning_rate': '0.001', 'epoch': '0.06909', 'train/total_time_seconds': '14.41', 'train/time_per_step_avg': '0.01651', 'train/epoch_time_elapsed': '124.6', 'train/estimated_remaining_minutes': '0.05273'}
93
  {'loss': '2.923', 'grad_norm': '0.8828', 'learning_rate': '0.001', 'epoch': '0.07078', 'train/total_time_seconds': '14.74', 'train/time_per_step_avg': '0.01652', 'train/epoch_time_elapsed': '126.2', 'train/estimated_remaining_minutes': '0.04681'}
94
  {'loss': '2.909', 'grad_norm': '0.8008', 'learning_rate': '0.001', 'epoch': '0.07246', 'train/total_time_seconds': '15.08', 'train/time_per_step_avg': '0.01655', 'train/epoch_time_elapsed': '127.7', 'train/estimated_remaining_minutes': '0.04091'}
95
  {'loss': '2.907', 'grad_norm': '1.133', 'learning_rate': '0.001', 'epoch': '0.07415', 'train/total_time_seconds': '15.41', 'train/time_per_step_avg': '0.01656', 'train/epoch_time_elapsed': '129.2', 'train/estimated_remaining_minutes': '0.03502'}
96
+ {'loss': '2.868', 'grad_norm': '0.9922', 'learning_rate': '0.001', 'epoch': '0.07583', 'train/total_time_seconds': '15.74', 'train/time_per_step_avg': '0.01656', 'train/epoch_time_elapsed': '130.7', 'train/estimated_remaining_minutes': '0.02914'}
97
+ - If you're using `trust_remote_code=True`, you can get rid of this warning by loading the model with an auto class. See https://huggingface.co/docs/transformers/en/model_doc/auto#auto-classes
98
+ {'eval_loss': '2.89', 'eval_runtime': '7.735', 'eval_samples_per_second': '1232', 'eval_steps_per_second': '0.905', 'epoch': '0.07583', 'train/total_time_seconds': '15.74', 'train/time_per_step_avg': '0.01656', 'train/epoch_time_elapsed': '138.4', 'train/estimated_remaining_minutes': '0.02914'}
99
+ - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception).
100
+ - If you are not the owner of the model architecture class, please contact the model code owner to update it.
101
+ Writing model shards: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1/1 [00:00<00:00, 326.63it/s]
102
+ 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [02:33<00:00, 11.6[transformers] TinyLlamaForCausalLM has generative capabilities, as `prepare_inputs_for_generation` is explicitly defined. However, it doesn't directly inherit from `GenerationMixin`. From πŸ‘‰v4.50πŸ‘ˆ onwards, `PreTrainedModel` will NOT inherit from `GenerationMixin`, and this model will lose the ability to call `generate` and other related functions.
103
+ {'loss': '2.875', 'grad_norm': '0.8125', 'learning_rate': '0.001', 'epoch': '0.07752', 'train/total_time_seconds': '16.07', 'train/time_per_step_avg': '0.01656', 'train/epoch_time_elapsed': '140', 'train/estimated_remaining_minutes': '0.02329'}
104
+ {'loss': '2.869', 'grad_norm': '0.8047', 'learning_rate': '0.001', 'epoch': '0.0792', 'train/total_time_seconds': '16.4', 'train/time_per_step_avg': '0.01656', 'train/epoch_time_elapsed': '141.5', 'train/estimated_remaining_minutes': '0.01745'}
105
+ {'loss': '2.866', 'grad_norm': '0.8867', 'learning_rate': '0.001', 'epoch': '0.08089', 'train/total_time_seconds': '16.73', 'train/time_per_step_avg': '0.01653', 'train/epoch_time_elapsed': '143', 'train/estimated_remaining_minutes': '0.01162'}
106
+ {'loss': '2.861', 'grad_norm': '0.7383', 'learning_rate': '0.001', 'epoch': '0.08257', 'train/total_time_seconds': '17.06', 'train/time_per_step_avg': '0.01653', 'train/epoch_time_elapsed': '144.5', 'train/estimated_remaining_minutes': '0.005803'}
107
+ {'loss': '2.839', 'grad_norm': '0.8867', 'learning_rate': '0.001', 'epoch': '0.08426', 'train/total_time_seconds': '17.39', 'train/time_per_step_avg': '0.01654', 'train/epoch_time_elapsed': '146', 'train/estimated_remaining_minutes': '0'}
108
+ - If you're using `trust_remote_code=True`, you can get rid of this warning by loading the model with an auto class. See https://huggingface.co/docs/transformers/en/model_doc/auto#auto-classes
109
+ {'eval_loss': '2.841', 'eval_runtime': '7.595', 'eval_samples_per_second': '1254', 'eval_steps_per_second': '0.922', 'epoch': '0.08426', 'train/total_time_seconds': '17.39', 'train/time_per_step_avg': '0.01654', 'train/epoch_time_elapsed': '153.6', 'train/estimated_remaining_minutes': '0'}
110
+ - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception).
111
+ - If you are not the owner of the model architecture class, please contact the model code owner to update it.
112
+ Writing model shards: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1/1 [00:00<00:00, 184.90it/s]
113
+ 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [02:33<00:00, 6.51it/s], ?it/s]
114
+ {'train_runtime': '154.6', 'train_samples_per_second': '517.5', 'train_steps_per_second': '6.468', 'train_loss': '3.728', 'epoch': '0.08426', 'train/total_time_seconds': '17.39', 'train/time_per_step_avg': '0.01654', 'train/epoch_time_elapsed': '153.7', 'train/estimated_remaining_minutes': '0'}
115
+ 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 7/7 [00:05<00:00, 1.32it/s]
116
+ [transformers] TinyLlamaForCausalLM has generative capabilities, as `prepare_inputs_for_generation` is explicitly defined. However, it doesn't directly inherit from `GenerationMixin`. From πŸ‘‰v4.50πŸ‘ˆ onwards, `PreTrainedModel` will NOT inherit from `GenerationMixin`, and this model will lose the ability to call `generate` and other related functions.
117
+ - If you're using `trust_remote_code=True`, you can get rid of this warning by loading the model with an auto class. See https://huggingface.co/docs/transformers/en/model_doc/auto#auto-classes
118
+ - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception).
119
+ - If you are not the owner of the model architecture class, please contact the model code owner to update it.
120
+ Writing model shards: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1/1 [00:00<00:00, 324.94it/s]
121
+ >>> FINISHED mlp-linear-3L successfully
zain/Activation/wandb/run-20260819_171215-2nil4rqn/files/wandb-summary.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"train_steps_per_second":6.468,"_wandb":{"runtime":161},"train/learning_rate":0.001,"train/global_step":1000,"eval/steps_per_second":0.927,"train/epoch":0.08426019548365352,"_timestamp":1.787159697355621e+09,"train_runtime":154.6026,"_runtime":161,"train/loss":2.8389766693115233,"train/grad_norm":0.88671875,"eval/loss":2.841240167617798,"eval/runtime":7.5505,"train_loss":3.7283405075073244,"eval/samples_per_second":1261.772,"_step":61,"total_flos":1.2101615616e+14,"train_samples_per_second":517.456}
zain/Activation/wandb/run-20260819_171215-2nil4rqn/logs/debug-core.log CHANGED
@@ -54,3 +54,11 @@
54
  {"time":"2026-08-19T17:12:15.319455303Z","level":"INFO","msg":"handleInformInit: received","streamId":"2nil4rqn","id":"6(@)"}
55
  {"time":"2026-08-19T17:12:15.582732337Z","level":"INFO","msg":"handleInformInit: stream started","streamId":"2nil4rqn","id":"6(@)"}
56
  {"time":"2026-08-19T17:12:21.153010638Z","level":"INFO","msg":"connection: cancelling request","id":"6(@)","requestId":"g8ym1re70rep"}
 
 
 
 
 
 
 
 
 
54
  {"time":"2026-08-19T17:12:15.319455303Z","level":"INFO","msg":"handleInformInit: received","streamId":"2nil4rqn","id":"6(@)"}
55
  {"time":"2026-08-19T17:12:15.582732337Z","level":"INFO","msg":"handleInformInit: stream started","streamId":"2nil4rqn","id":"6(@)"}
56
  {"time":"2026-08-19T17:12:21.153010638Z","level":"INFO","msg":"connection: cancelling request","id":"6(@)","requestId":"g8ym1re70rep"}
57
+ {"time":"2026-08-19T17:14:57.367689028Z","level":"INFO","msg":"connection: cancelling request","id":"6(@)","requestId":"g8ym1re70rep"}
58
+ {"time":"2026-08-19T17:14:58.279859891Z","level":"INFO","msg":"connection: cancelling request","id":"6(@)","requestId":"g8ym1re70rep"}
59
+ {"time":"2026-08-19T17:14:58.28171116Z","level":"INFO","msg":"handleInformFinish: finish message received","streamId":"2nil4rqn","id":"6(@)"}
60
+ {"time":"2026-08-19T17:14:58.282403203Z","level":"INFO","msg":"handleInformFinish: stream closed","streamId":"2nil4rqn","id":"6(@)"}
61
+ {"time":"2026-08-19T17:15:00.260323511Z","level":"INFO","msg":"connection: closing","id":"6(@)"}
62
+ {"time":"2026-08-19T17:15:00.26039463Z","level":"INFO","msg":"connection: closed successfully","id":"6(@)"}
63
+ {"time":"2026-08-19T17:15:00.260332968Z","level":"INFO","msg":"processOutgoingData: finished","id":"6(@)"}
64
+ {"time":"2026-08-19T17:15:00.260403826Z","level":"INFO","msg":"connection: ManageConnectionData: connection closed","id":"6(@)"}
zain/Activation/wandb/run-20260819_171215-2nil4rqn/logs/debug-internal.log CHANGED
@@ -23,3 +23,15 @@
23
  {"time":"2026-08-19T17:14:01.314421382Z","level":"INFO","msg":"filestream: request sent","status":"200 OK"}
24
  {"time":"2026-08-19T17:14:16.125013671Z","level":"INFO","msg":"filestream: sending request","total_files":4,"history_offset":41,"history_lines":6,"events_offset":13,"events_lines":2,"console_offset":68,"console_lines":1}
25
  {"time":"2026-08-19T17:14:16.73098493Z","level":"INFO","msg":"filestream: request sent","status":"200 OK"}
 
 
 
 
 
 
 
 
 
 
 
 
 
23
  {"time":"2026-08-19T17:14:01.314421382Z","level":"INFO","msg":"filestream: request sent","status":"200 OK"}
24
  {"time":"2026-08-19T17:14:16.125013671Z","level":"INFO","msg":"filestream: sending request","total_files":4,"history_offset":41,"history_lines":6,"events_offset":13,"events_lines":2,"console_offset":68,"console_lines":1}
25
  {"time":"2026-08-19T17:14:16.73098493Z","level":"INFO","msg":"filestream: request sent","status":"200 OK"}
26
+ {"time":"2026-08-19T17:14:31.124676234Z","level":"INFO","msg":"filestream: sending request","total_files":4,"history_offset":47,"history_lines":6,"events_offset":15,"events_lines":2,"console_offset":74,"console_lines":23}
27
+ {"time":"2026-08-19T17:14:31.25448083Z","level":"INFO","msg":"filestream: request sent","status":"200 OK"}
28
+ {"time":"2026-08-19T17:14:46.124973259Z","level":"INFO","msg":"filestream: sending request","total_files":4,"history_offset":53,"history_lines":6,"events_offset":17,"events_lines":2,"console_offset":90,"console_lines":1}
29
+ {"time":"2026-08-19T17:14:46.269794228Z","level":"INFO","msg":"filestream: request sent","status":"200 OK"}
30
+ {"time":"2026-08-19T17:14:58.136247949Z","level":"INFO","msg":"fileTransfer: Close: file transfer manager closed"}
31
+ {"time":"2026-08-19T17:14:58.136517121Z","level":"INFO","msg":"filestream: sending request","total_files":4,"history_offset":59,"history_lines":3,"events_offset":19,"events_lines":1,"console_offset":96,"console_lines":25,"uploaded_len":3,"complete":true,"exit_code":0}
32
+ {"time":"2026-08-19T17:14:58.277374581Z","level":"INFO","msg":"filestream: request sent","status":"200 OK"}
33
+ {"time":"2026-08-19T17:14:58.278584446Z","level":"INFO","msg":"handler: operation stats","stats":{}}
34
+ {"time":"2026-08-19T17:14:58.281745009Z","level":"INFO","msg":"stream: finishing up"}
35
+ {"time":"2026-08-19T17:14:58.281775882Z","level":"INFO","msg":"handler: closed"}
36
+ {"time":"2026-08-19T17:14:58.281829879Z","level":"INFO","msg":"sender: closed"}
37
+ {"time":"2026-08-19T17:14:58.281833448Z","level":"INFO","msg":"stream: all finished"}
zain/Activation/wandb/run-20260819_171215-2nil4rqn/logs/debug.log CHANGED
@@ -21,3 +21,8 @@ config: {'_wandb': {}}
21
  2026-08-19 17:12:16,120 INFO MainThread:3920475 [wandb_run.py:_config_callback():1346] config_cb None None {'transformers_version': '5.16.0.dev0', 'architectures': None, 'output_hidden_states': False, 'return_dict': True, 'dtype': None, 'chunk_size_feed_forward': 0, 'is_encoder_decoder': False, 'id2label': {0: 'LABEL_0', 1: 'LABEL_1'}, 'label2id': {'LABEL_0': 0, 'LABEL_1': 1}, 'problem_type': None, 'vocab_size': 4096, 'hidden_size': 128, 'intermediate_size': 256, 'num_hidden_layers': 3, 'num_attention_heads': 4, 'num_key_value_heads': 4, 'hidden_act': 'silu', 'max_position_embeddings': 512, 'initializer_range': 0.02, 'rms_norm_eps': 1e-06, 'use_cache': False, 'pad_token_id': 0, 'bos_token_id': 1, 'eos_token_id': 2, 'pretraining_tp': 1, 'tie_word_embeddings': True, 'rope_parameters': {'rope_theta': 10000.0, 'rope_type': 'default'}, 'attention_bias': False, 'attention_dropout': 0.0, 'mlp_bias': False, 'head_dim': 32, '_name_or_path': '', 'tokenizer_name': 'w-ahmad/tiny-stories-tokenizer', 'mlp_type': 'mlp', 'activation': 'linear', 'waleed_beta': 10.0, 'powlu_m': 3.0, 'model_type': 'tiny_llama', 'output_attentions': False, 'output_dir': 'out/mlp-linear-3L_run', 'per_device_train_batch_size': 80, 'num_train_epochs': 1, 'max_steps': 1000, 'learning_rate': 0.001, 'lr_scheduler_type': 'constant_with_warmup', 'lr_scheduler_kwargs': None, 'warmup_steps': 200, 'optim': 'adamw_torch_fused', 'optim_args': None, 'weight_decay': 0.0, 'adam_beta1': 0.9, 'adam_beta2': 0.999, 'adam_epsilon': 1e-08, 'optim_target_modules': None, 'gradient_accumulation_steps': 1, 'average_tokens_across_devices': True, 'max_grad_norm': 1.0, 'label_smoothing_factor': 0.0, 'bf16': True, 'fp16': False, 'bf16_full_eval': False, 'fp16_full_eval': False, 'tf32': None, 'gradient_checkpointing': False, 'gradient_checkpointing_kwargs': None, 'torch_compile': False, 'torch_compile_backend': None, 'torch_compile_mode': None, 'use_liger_kernel': False, 'liger_kernel_config': None, 'neftune_noise_alpha': None, 'torch_empty_cache_steps': None, 'auto_find_batch_size': False, 'logging_strategy': 'steps', 'logging_steps': 20, 'logging_first_step': False, 'log_on_each_node': True, 'logging_nan_inf_filter': True, 'include_num_input_tokens_seen': 'no', 'log_level': 'passive', 'log_level_replica': 'warning', 'disable_tqdm': False, 'report_to': ['wandb'], 'run_name': 'LM-mlp-linear-3L-1.0M-20260819-171214', 'project': 'huggingface', 'trackio_space_id': None, 'trackio_bucket_id': None, 'trackio_static_space_id': None, 'eval_strategy': 'steps', 'eval_steps': 100, 'eval_delay': 0, 'per_device_eval_batch_size': 1500, 'prediction_loss_only': False, 'eval_on_start': False, 'eval_do_concat_batches': True, 'eval_use_gather_object': False, 'eval_accumulation_steps': None, 'include_for_metrics': [], 'batch_eval_metrics': False, 'save_only_model': False, 'save_strategy': 'steps', 'save_steps': 100, 'save_on_each_node': False, 'save_total_limit': None, 'enable_jit_checkpoint': False, 'push_to_hub': False, 'hub_token': '<HUB_TOKEN>', 'hub_private_repo': None, 'hub_model_id': 'w-ahmad/6L-mlp-linear-3L', 'hub_strategy': 'every_save', 'hub_always_push': False, 'hub_revision': None, 'load_best_model_at_end': False, 'metric_for_best_model': None, 'greater_is_better': None, 'ignore_data_skip': False, 'restore_callback_states_from_checkpoint': False, 'full_determinism': False, 'seed': 42, 'data_seed': 42, 'use_cpu': False, 'accelerator_config': {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}, 'parallelism_config': None, 'dataloader_drop_last': False, 'dataloader_num_workers': 0, 'dataloader_pin_memory': True, 'dataloader_persistent_workers': False, 'dataloader_prefetch_factor': None, 'dataloader_multiprocessing_context': None, 'dataloader_in_order': True, 'remove_unused_columns': False, 'label_names': None, 'train_sampling_strategy': 'random', 'length_column_name': 'length', 'ddp_find_unused_parameters': None, 'ddp_bucket_cap_mb': None, 'ddp_broadcast_buffers': None, 'ddp_static_graph': None, 'ddp_backend': None, 'ddp_timeout': 1800, 'fsdp': None, 'fsdp_config': None, 'deepspeed': None, 'debug': [], 'skip_memory_metrics': True, 'do_train': False, 'do_eval': True, 'do_predict': False, 'resume_from_checkpoint': None, 'local_rank': -1}
22
  2026-08-19 17:12:16,122 INFO MainThread:3920475 [wandb_config.py:__setitem__():155] [no run ID] config set model/num_parameters = 1016704 - <bound method Run._config_callback of <wandb.sdk.wandb_run.Run object at 0x14f219615650>>
23
  2026-08-19 17:12:16,122 INFO MainThread:3920475 [wandb_run.py:_config_callback():1346] config_cb model/num_parameters 1016704 None
 
 
 
 
 
 
21
  2026-08-19 17:12:16,120 INFO MainThread:3920475 [wandb_run.py:_config_callback():1346] config_cb None None {'transformers_version': '5.16.0.dev0', 'architectures': None, 'output_hidden_states': False, 'return_dict': True, 'dtype': None, 'chunk_size_feed_forward': 0, 'is_encoder_decoder': False, 'id2label': {0: 'LABEL_0', 1: 'LABEL_1'}, 'label2id': {'LABEL_0': 0, 'LABEL_1': 1}, 'problem_type': None, 'vocab_size': 4096, 'hidden_size': 128, 'intermediate_size': 256, 'num_hidden_layers': 3, 'num_attention_heads': 4, 'num_key_value_heads': 4, 'hidden_act': 'silu', 'max_position_embeddings': 512, 'initializer_range': 0.02, 'rms_norm_eps': 1e-06, 'use_cache': False, 'pad_token_id': 0, 'bos_token_id': 1, 'eos_token_id': 2, 'pretraining_tp': 1, 'tie_word_embeddings': True, 'rope_parameters': {'rope_theta': 10000.0, 'rope_type': 'default'}, 'attention_bias': False, 'attention_dropout': 0.0, 'mlp_bias': False, 'head_dim': 32, '_name_or_path': '', 'tokenizer_name': 'w-ahmad/tiny-stories-tokenizer', 'mlp_type': 'mlp', 'activation': 'linear', 'waleed_beta': 10.0, 'powlu_m': 3.0, 'model_type': 'tiny_llama', 'output_attentions': False, 'output_dir': 'out/mlp-linear-3L_run', 'per_device_train_batch_size': 80, 'num_train_epochs': 1, 'max_steps': 1000, 'learning_rate': 0.001, 'lr_scheduler_type': 'constant_with_warmup', 'lr_scheduler_kwargs': None, 'warmup_steps': 200, 'optim': 'adamw_torch_fused', 'optim_args': None, 'weight_decay': 0.0, 'adam_beta1': 0.9, 'adam_beta2': 0.999, 'adam_epsilon': 1e-08, 'optim_target_modules': None, 'gradient_accumulation_steps': 1, 'average_tokens_across_devices': True, 'max_grad_norm': 1.0, 'label_smoothing_factor': 0.0, 'bf16': True, 'fp16': False, 'bf16_full_eval': False, 'fp16_full_eval': False, 'tf32': None, 'gradient_checkpointing': False, 'gradient_checkpointing_kwargs': None, 'torch_compile': False, 'torch_compile_backend': None, 'torch_compile_mode': None, 'use_liger_kernel': False, 'liger_kernel_config': None, 'neftune_noise_alpha': None, 'torch_empty_cache_steps': None, 'auto_find_batch_size': False, 'logging_strategy': 'steps', 'logging_steps': 20, 'logging_first_step': False, 'log_on_each_node': True, 'logging_nan_inf_filter': True, 'include_num_input_tokens_seen': 'no', 'log_level': 'passive', 'log_level_replica': 'warning', 'disable_tqdm': False, 'report_to': ['wandb'], 'run_name': 'LM-mlp-linear-3L-1.0M-20260819-171214', 'project': 'huggingface', 'trackio_space_id': None, 'trackio_bucket_id': None, 'trackio_static_space_id': None, 'eval_strategy': 'steps', 'eval_steps': 100, 'eval_delay': 0, 'per_device_eval_batch_size': 1500, 'prediction_loss_only': False, 'eval_on_start': False, 'eval_do_concat_batches': True, 'eval_use_gather_object': False, 'eval_accumulation_steps': None, 'include_for_metrics': [], 'batch_eval_metrics': False, 'save_only_model': False, 'save_strategy': 'steps', 'save_steps': 100, 'save_on_each_node': False, 'save_total_limit': None, 'enable_jit_checkpoint': False, 'push_to_hub': False, 'hub_token': '<HUB_TOKEN>', 'hub_private_repo': None, 'hub_model_id': 'w-ahmad/6L-mlp-linear-3L', 'hub_strategy': 'every_save', 'hub_always_push': False, 'hub_revision': None, 'load_best_model_at_end': False, 'metric_for_best_model': None, 'greater_is_better': None, 'ignore_data_skip': False, 'restore_callback_states_from_checkpoint': False, 'full_determinism': False, 'seed': 42, 'data_seed': 42, 'use_cpu': False, 'accelerator_config': {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}, 'parallelism_config': None, 'dataloader_drop_last': False, 'dataloader_num_workers': 0, 'dataloader_pin_memory': True, 'dataloader_persistent_workers': False, 'dataloader_prefetch_factor': None, 'dataloader_multiprocessing_context': None, 'dataloader_in_order': True, 'remove_unused_columns': False, 'label_names': None, 'train_sampling_strategy': 'random', 'length_column_name': 'length', 'ddp_find_unused_parameters': None, 'ddp_bucket_cap_mb': None, 'ddp_broadcast_buffers': None, 'ddp_static_graph': None, 'ddp_backend': None, 'ddp_timeout': 1800, 'fsdp': None, 'fsdp_config': None, 'deepspeed': None, 'debug': [], 'skip_memory_metrics': True, 'do_train': False, 'do_eval': True, 'do_predict': False, 'resume_from_checkpoint': None, 'local_rank': -1}
22
  2026-08-19 17:12:16,122 INFO MainThread:3920475 [wandb_config.py:__setitem__():155] [no run ID] config set model/num_parameters = 1016704 - <bound method Run._config_callback of <wandb.sdk.wandb_run.Run object at 0x14f219615650>>
23
  2026-08-19 17:12:16,122 INFO MainThread:3920475 [wandb_run.py:_config_callback():1346] config_cb model/num_parameters 1016704 None
24
+ 2026-08-19 17:14:57,366 INFO MainThread:3920475 [wandb_run.py:_finish():2383] finishing run deepnevro-deepnevro/research-ultimate-checking/2nil4rqn
25
+ 2026-08-19 17:14:57,367 INFO MainThread:3920475 [wandb_run.py:_atexit_cleanup():2588] got exitcode: 0
26
+ 2026-08-19 17:14:57,367 INFO MainThread:3920475 [wandb_run.py:_restore():2570] restore
27
+ 2026-08-19 17:14:57,367 INFO MainThread:3920475 [wandb_run.py:_restore():2576] restore done
28
+ 2026-08-19 17:14:58,281 INFO MainThread:3920475 [wandb_run.py:_footer_sync_info():3993] logging synced files
zain/Activation/wandb/run-20260819_171215-2nil4rqn/run-2nil4rqn.wandb CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:eabb4a030bca62aea6f627cbe7cd12ff2a9480104eafd775ab756749b9220aa4
3
- size 196608
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a7b5da609b9fccdc3e7dd6fde9d7d49ec5a9b66eb9fe3dc139ddf206bfab728d
3
+ size 236528