| /home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available. |
| warnings.warn('Grouped GEMM not available.') |
| wandb: [wandb.login()] Loaded credentials for https://api.wandb.ai from /home/henry/.netrc. |
| wandb: Currently logged in as: hbfreed to https://api.wandb.ai. Use `wandb login |
| wandb: Tracking run with wandb version 0.28.0 |
| wandb: Run data is saved locally in outputs/healed/grid_math/glean_keep75_s1225/wandb/run-20260716_142818-9x4iij2d |
| wandb: Run `wandb offline` to turn off syncing. |
| wandb: Syncing run glean-math-keep75-s1225 |
| wandb: βοΈ View project at https://wandb.ai/hbfreed/glean-grid |
| wandb: π View run at https://wandb.ai/hbfreed/glean-grid/runs/9x4iij2d |
|
Loading checkpoint shards: 0%| | 0/3 [00:00<?, ?it/s]
Loading checkpoint shards: 33%|ββββ | 1/3 [00:01<00:03, 1.53s/it]
Loading checkpoint shards: 67%|βββββββ | 2/3 [00:03<00:01, 1.59s/it]
Loading checkpoint shards: 100%|ββββββββββ| 3/3 [00:03<00:00, 1.04it/s]
Loading checkpoint shards: 100%|ββββββββββ| 3/3 [00:03<00:00, 1.13s/it] |
| resumed student weights from outputs/healed/grid_math/glean_keep75_s1225/step0100 (fresh optimizer, step counter at 0) |
| 12115 cached top-128 chat trajectories / 6,476,634 unique tokens | 53 steps/epoch | 150 total steps | student params 5.31B | teacher overlap=False |
| restored optimizer/scheduler state from step 100; rebuilt 260 paged buffers |
| {"step": 101, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.016109059700369834, "tokens": 120000, "cumulative_loss_tokens": 12120000, "grad_norm": 0.2041015625, "lr": 3e-05, "finish_rate": 0.798, "comp_len": 538.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 63.3, "frames": {"chat": 223}, "mem_gb": 21.95} |
| {"step": 102, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.016362527111552966, "tokens": 120000, "cumulative_loss_tokens": 12240000, "grad_norm": 0.21875, "lr": 3e-05, "finish_rate": 0.772, "comp_len": 582.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.6, "frames": {"chat": 206}, "mem_gb": 22.1} |
| {"step": 103, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.013859456993810212, "tokens": 120000, "cumulative_loss_tokens": 12360000, "grad_norm": 0.228515625, "lr": 3e-05, "finish_rate": 0.784, "comp_len": 563.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.6, "frames": {"chat": 213}, "mem_gb": 22.02} |
| {"step": 104, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.018646482431787688, "tokens": 120000, "cumulative_loss_tokens": 12480000, "grad_norm": 0.22265625, "lr": 3e-05, "finish_rate": 0.843, "comp_len": 538.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.9, "frames": {"chat": 223}, "mem_gb": 21.96} |
| {"step": 105, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.015217417630545484, "tokens": 120000, "cumulative_loss_tokens": 12600000, "grad_norm": 0.201171875, "lr": 3e-05, "finish_rate": 0.828, "comp_len": 528.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.6, "frames": {"chat": 227}, "mem_gb": 22.07} |
| {"step": 106, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.018008409859096478, "tokens": 120000, "cumulative_loss_tokens": 12720000, "grad_norm": 0.255859375, "lr": 3e-05, "finish_rate": 0.889, "comp_len": 474.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.6, "frames": {"chat": 253}, "mem_gb": 22.09} |
| {"step": 107, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012912440449073135, "tokens": 120000, "cumulative_loss_tokens": 12840000, "grad_norm": 0.1865234375, "lr": 3e-05, "finish_rate": 0.792, "comp_len": 555.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.6, "frames": {"chat": 216}, "mem_gb": 22.1} |
| {"step": 108, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010693734145350754, "tokens": 120000, "cumulative_loss_tokens": 12960000, "grad_norm": 0.1748046875, "lr": 3e-05, "finish_rate": 0.766, "comp_len": 585.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.3, "frames": {"chat": 205}, "mem_gb": 22.07} |
| {"step": 109, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.015156030047448197, "tokens": 120000, "cumulative_loss_tokens": 13080000, "grad_norm": 0.2138671875, "lr": 3e-05, "finish_rate": 0.729, "comp_len": 579.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.3, "frames": {"chat": 207}, "mem_gb": 22.16} |
| {"step": 110, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.013672215391354015, "tokens": 120000, "cumulative_loss_tokens": 13200000, "grad_norm": 0.1982421875, "lr": 3e-05, "finish_rate": 0.814, "comp_len": 558.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.5, "frames": {"chat": 215}, "mem_gb": 22.08} |
| The attention mask is not set and cannot be inferred from input because pad token is same as eos token. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results. |
| [eval step 110] sample: 'To solve this problem, we need to analyze the spiral pattern of numbers from 1 to 49 arranged on a square grid and identify the four shaded squares that lie on the same diagonal as the number 7. We th' |
| {"step": 111, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011061214675944453, "tokens": 120000, "cumulative_loss_tokens": 13320000, "grad_norm": 0.2060546875, "lr": 3e-05, "finish_rate": 0.86, "comp_len": 526.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.2, "frames": {"chat": 228}, "mem_gb": 22.1} |
| {"step": 112, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012901123281742912, "tokens": 120000, "cumulative_loss_tokens": 13440000, "grad_norm": 0.1982421875, "lr": 3e-05, "finish_rate": 0.747, "comp_len": 543.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.9, "frames": {"chat": 221}, "mem_gb": 22.14} |
| {"step": 113, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009574843731914492, "tokens": 120000, "cumulative_loss_tokens": 13560000, "grad_norm": 0.15625, "lr": 3e-05, "finish_rate": 0.882, "comp_len": 472.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.2, "frames": {"chat": 254}, "mem_gb": 21.93} |
| {"step": 114, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011826626441131036, "tokens": 120000, "cumulative_loss_tokens": 13680000, "grad_norm": 0.2392578125, "lr": 3e-05, "finish_rate": 0.843, "comp_len": 571.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.7, "frames": {"chat": 210}, "mem_gb": 22.06} |
| {"step": 115, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010666260581545066, "tokens": 120000, "cumulative_loss_tokens": 13800000, "grad_norm": 0.166015625, "lr": 3e-05, "finish_rate": 0.827, "comp_len": 531.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.6, "frames": {"chat": 226}, "mem_gb": 22.02} |
| {"step": 116, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01181182961029117, "tokens": 120000, "cumulative_loss_tokens": 13920000, "grad_norm": 0.193359375, "lr": 3e-05, "finish_rate": 0.802, "comp_len": 566.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.3, "frames": {"chat": 212}, "mem_gb": 22.09} |
| {"step": 117, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.014291961643429628, "tokens": 120000, "cumulative_loss_tokens": 14040000, "grad_norm": 0.20703125, "lr": 3e-05, "finish_rate": 0.754, "comp_len": 568.7, "t_data_s": 0.2, "t_rollout_s": 0.0, "t_step_s": 47.3, "frames": {"chat": 211}, "mem_gb": 22.02} |
| {"step": 118, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011632958819546426, "tokens": 120000, "cumulative_loss_tokens": 14160000, "grad_norm": 0.1640625, "lr": 3e-05, "finish_rate": 0.776, "comp_len": 612.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 42.9, "frames": {"chat": 196}, "mem_gb": 22.07} |
| {"step": 119, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011126866567115454, "tokens": 120000, "cumulative_loss_tokens": 14280000, "grad_norm": 0.2119140625, "lr": 3e-05, "finish_rate": 0.811, "comp_len": 566.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.0, "frames": {"chat": 212}, "mem_gb": 22.09} |
| {"step": 120, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010413994966812121, "tokens": 120000, "cumulative_loss_tokens": 14400000, "grad_norm": 0.16015625, "lr": 3e-05, "finish_rate": 0.877, "comp_len": 491.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.0, "frames": {"chat": 244}, "mem_gb": 22.0} |
| [eval step 120] sample: "To solve this problem, we need to understand the structure of the spiral pattern and identify the numbers on the same diagonal as the number 7. Let's break down the problem step-by-step:\n\n1. **Underst" |
| {"step": 121, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010520029380288906, "tokens": 120000, "cumulative_loss_tokens": 14520000, "grad_norm": 0.166015625, "lr": 3e-05, "finish_rate": 0.838, "comp_len": 540.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.8, "frames": {"chat": 222}, "mem_gb": 22.05} |
| {"step": 122, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011724793996859807, "tokens": 120000, "cumulative_loss_tokens": 14640000, "grad_norm": 0.1884765625, "lr": 3e-05, "finish_rate": 0.78, "comp_len": 550.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.6, "frames": {"chat": 218}, "mem_gb": 22.09} |
| {"step": 123, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011794728558894713, "tokens": 120000, "cumulative_loss_tokens": 14760000, "grad_norm": 0.1982421875, "lr": 3e-05, "finish_rate": 0.913, "comp_len": 476.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.8, "frames": {"chat": 252}, "mem_gb": 21.97} |
| {"step": 124, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012576708099556466, "tokens": 120000, "cumulative_loss_tokens": 14880000, "grad_norm": 0.19140625, "lr": 3e-05, "finish_rate": 0.728, "comp_len": 594.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.5, "frames": {"chat": 202}, "mem_gb": 22.14} |
| {"step": 125, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012525027039841128, "tokens": 120000, "cumulative_loss_tokens": 15000000, "grad_norm": 0.44140625, "lr": 3e-05, "finish_rate": 0.835, "comp_len": 506.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.6, "frames": {"chat": 237}, "mem_gb": 22.1} |
| {"step": 126, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011847056540916674, "tokens": 120000, "cumulative_loss_tokens": 15120000, "grad_norm": 0.166015625, "lr": 3e-05, "finish_rate": 0.868, "comp_len": 512.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.5, "frames": {"chat": 234}, "mem_gb": 22.08} |
| {"step": 127, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010580201309620558, "tokens": 120000, "cumulative_loss_tokens": 15240000, "grad_norm": 0.193359375, "lr": 3e-05, "finish_rate": 0.809, "comp_len": 558.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.7, "frames": {"chat": 215}, "mem_gb": 22.1} |
| {"step": 128, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010746679509648433, "tokens": 120000, "cumulative_loss_tokens": 15360000, "grad_norm": 0.1796875, "lr": 3e-05, "finish_rate": 0.876, "comp_len": 512.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.7, "frames": {"chat": 234}, "mem_gb": 22.03} |
| {"step": 129, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010033142010107016, "tokens": 120000, "cumulative_loss_tokens": 15480000, "grad_norm": 0.1884765625, "lr": 3e-05, "finish_rate": 0.801, "comp_len": 555.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.8, "frames": {"chat": 216}, "mem_gb": 22.08} |
| {"step": 130, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01059437686605379, "tokens": 120000, "cumulative_loss_tokens": 15600000, "grad_norm": 0.1787109375, "lr": 3e-05, "finish_rate": 0.805, "comp_len": 571.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.3, "frames": {"chat": 210}, "mem_gb": 22.05} |
| [eval step 130] sample: 'To solve this problem, we need to arrange the numbers from 1 to 49 in a spiral pattern on a square grid starting from the center. We then identify the four shaded squares that lie on the same diagonal' |
| {"step": 131, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011416423269287528, "tokens": 120000, "cumulative_loss_tokens": 15720000, "grad_norm": 0.169921875, "lr": 3e-05, "finish_rate": 0.719, "comp_len": 603.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.1, "frames": {"chat": 199}, "mem_gb": 22.09} |
| {"step": 132, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011045920558762736, "tokens": 120000, "cumulative_loss_tokens": 15840000, "grad_norm": 0.1865234375, "lr": 3e-05, "finish_rate": 0.824, "comp_len": 571.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.4, "frames": {"chat": 210}, "mem_gb": 22.11} |
| {"step": 133, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010189920482278103, "tokens": 120000, "cumulative_loss_tokens": 15960000, "grad_norm": 0.1689453125, "lr": 3e-05, "finish_rate": 0.902, "comp_len": 533.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.7, "frames": {"chat": 225}, "mem_gb": 22.05} |
| {"step": 134, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011348977557197213, "tokens": 120000, "cumulative_loss_tokens": 16080000, "grad_norm": 0.1728515625, "lr": 3e-05, "finish_rate": 0.913, "comp_len": 474.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.6, "frames": {"chat": 253}, "mem_gb": 21.95} |
| {"step": 135, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010840521045137818, "tokens": 120000, "cumulative_loss_tokens": 16200000, "grad_norm": 0.1513671875, "lr": 3e-05, "finish_rate": 0.903, "comp_len": 485.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.1, "frames": {"chat": 247}, "mem_gb": 22.07} |
| {"step": 136, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01431039827777228, "tokens": 120000, "cumulative_loss_tokens": 16320000, "grad_norm": 0.2021484375, "lr": 3e-05, "finish_rate": 0.836, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.4, "frames": {"chat": 238}, "mem_gb": 22.07} |
| {"step": 137, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012493417616401954, "tokens": 120000, "cumulative_loss_tokens": 16440000, "grad_norm": 0.16796875, "lr": 3e-05, "finish_rate": 0.86, "comp_len": 510.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.6, "frames": {"chat": 235}, "mem_gb": 22.09} |
| {"step": 138, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011775391096648916, "tokens": 120000, "cumulative_loss_tokens": 16560000, "grad_norm": 0.1640625, "lr": 3e-05, "finish_rate": 0.805, "comp_len": 558.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.3, "frames": {"chat": 215}, "mem_gb": 22.06} |
| {"step": 139, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011272327020104665, "tokens": 120000, "cumulative_loss_tokens": 16680000, "grad_norm": 0.177734375, "lr": 3e-05, "finish_rate": 0.925, "comp_len": 447.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.0, "frames": {"chat": 268}, "mem_gb": 22.06} |
| {"step": 140, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011195119415794033, "tokens": 120000, "cumulative_loss_tokens": 16800000, "grad_norm": 0.1689453125, "lr": 3e-05, "finish_rate": 0.825, "comp_len": 526.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.1, "frames": {"chat": 228}, "mem_gb": 22.09} |
| [eval step 140] sample: 'To solve this problem, we need to arrange the numbers from 1 to 49 in a spiral pattern on a square grid and identify the four shaded squares that lie on the same diagonal as the number 7. Then, we wil' |
| {"step": 141, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01146084621019351, "tokens": 120000, "cumulative_loss_tokens": 16920000, "grad_norm": 0.1513671875, "lr": 3e-05, "finish_rate": 0.881, "comp_len": 476.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.6, "frames": {"chat": 252}, "mem_gb": 22.03} |
| {"step": 142, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01266345674659824, "tokens": 120000, "cumulative_loss_tokens": 17040000, "grad_norm": 0.1875, "lr": 3e-05, "finish_rate": 0.821, "comp_len": 538.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.1, "frames": {"chat": 223}, "mem_gb": 22.11} |
| {"step": 143, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012728927766492901, "tokens": 120000, "cumulative_loss_tokens": 17160000, "grad_norm": 0.1689453125, "lr": 3e-05, "finish_rate": 0.805, "comp_len": 531.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.6, "frames": {"chat": 226}, "mem_gb": 22.09} |
| {"step": 144, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.015653314464636303, "tokens": 120000, "cumulative_loss_tokens": 17280000, "grad_norm": 0.2421875, "lr": 3e-05, "finish_rate": 0.731, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.1, "frames": {"chat": 208}, "mem_gb": 22.14} |
| {"step": 145, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.00990355064412579, "tokens": 120000, "cumulative_loss_tokens": 17400000, "grad_norm": 0.154296875, "lr": 3e-05, "finish_rate": 0.883, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.6, "frames": {"chat": 240}, "mem_gb": 22.03} |
| {"step": 146, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011352584307268262, "tokens": 120000, "cumulative_loss_tokens": 17520000, "grad_norm": 0.171875, "lr": 3e-05, "finish_rate": 0.842, "comp_len": 540.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.9, "frames": {"chat": 222}, "mem_gb": 22.02} |
| {"step": 147, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009270905922904301, "tokens": 120000, "cumulative_loss_tokens": 17640000, "grad_norm": 0.1396484375, "lr": 3e-05, "finish_rate": 0.881, "comp_len": 508.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.6, "frames": {"chat": 236}, "mem_gb": 22.09} |
| {"step": 148, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010650625815118353, "tokens": 120000, "cumulative_loss_tokens": 17760000, "grad_norm": 0.181640625, "lr": 3e-05, "finish_rate": 0.834, "comp_len": 553.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.3, "frames": {"chat": 217}, "mem_gb": 22.06} |
| {"step": 149, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010312878351037702, "tokens": 120000, "cumulative_loss_tokens": 17880000, "grad_norm": 0.1611328125, "lr": 3e-05, "finish_rate": 0.921, "comp_len": 476.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.2, "frames": {"chat": 252}, "mem_gb": 21.97} |
| {"step": 150, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009814306650865667, "tokens": 120000, "cumulative_loss_tokens": 18000000, "grad_norm": 0.14453125, "lr": 3e-05, "finish_rate": 0.847, "comp_len": 540.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.7, "frames": {"chat": 222}, "mem_gb": 22.08} |
| [eval step 150] sample: 'To solve this problem, we need to arrange the numbers from 1 to 49 in a spiral pattern on a square grid and identify the four numbers that lie on the same diagonal as the number 7. We then need to det' |
| checkpoint snapshot queued -> outputs/healed/grid_math/glean_keep75_s1225/step0150 |
| wandb: updating run metadata |
| wandb: uploading output.log; uploading wandb-summary.json; uploading config.yaml |
| wandb: |
| wandb: Run history: |
| wandb: comp_len β
βββ
ββββββββ
βββββ
β
βββββββββ
βββββββ
β
ββ
ββ
β
|
| wandb: cumulative_loss_tokens βββββββββββββββββββββ
β
β
β
β
βββββββββββββββ |
| wandb: epoch ββββββββββββββββββββββββββββββββββββββββ |
| wandb: finish_rate ββββ
β
ββββββ
β
βββββ
βββββββββββββ
βββββββ
ββ
β
|
| wandb: forward_topk_kl βββββ
βββ
βββββββ
βββββββββββββββ
ββββββββββ |
| wandb: grad_norm ββββββββββββββββββββββββββββββββββββββββ |
| wandb: lr ββββββββββββββββββββββββββββββββββββββββ |
| wandb: mem_gb βββββ
ββ
βββββ
ββ
ββ
ββββββββ
ββ
ββ
β
ββ
ββββββββ
β |
| wandb: step βββββββββββββββββββββ
β
β
β
β
βββββββββββββββ |
| wandb: t_data_s ββββββββββββββββββββββββββββββββββββββββ |
| wandb: +3 ... |
| wandb: |
| wandb: Run summary: |
| wandb: comp_len 540.5 |
| wandb: cumulative_loss_tokens 18000000 |
| wandb: epoch 2 |
| wandb: finish_rate 0.847 |
| wandb: forward_topk_kl 0.00981 |
| wandb: grad_norm 0.14453 |
| wandb: lr 3e-05 |
| wandb: mem_gb 22.08 |
| wandb: step 150 |
| wandb: t_data_s 0 |
| wandb: +4 ... |
| wandb: |
| wandb: π View run glean-math-keep75-s1225 at: https://wandb.ai/hbfreed/glean-grid/runs/9x4iij2d |
| wandb: βοΈ View project at: https://wandb.ai/hbfreed/glean-grid |
| wandb: Synced 5 W&B file(s), 0 media file(s), 0 artifact file(s) and 0 other file(s) |
| wandb: Find logs at: outputs/healed/grid_math/glean_keep75_s1225/wandb/run-20260716_142818-9x4iij2d/logs |
| { |
| "correct": 910, |
| "accuracy": 0.6899166034874905, |
| "finished": 1315, |
| "finish_rate": 0.9969673995451099, |
| "mean_completion_tokens": 113.55724033358605 |
| } |
| saved item-level results -> outputs/evals/grid_math/glean_keep75_s1225_step100_chat.json |
| { |
| "correct": 909, |
| "accuracy": 0.689158453373768, |
| "finished": 1316, |
| "finish_rate": 0.9977255496588324, |
| "mean_completion_tokens": 114.47536012130402 |
| } |
| saved item-level results -> outputs/evals/grid_math/glean_keep75_s1225_step150_chat.json |
|
|