variable-reap-archive / healed /grid_math /glean_keep75_s1225.console.log
hbfreed's picture
Add files using upload-large-folder tool
d6fae6a verified
Raw
History Blame Contribute Delete
22.4 kB
/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
warnings.warn('Grouped GEMM not available.')
wandb: [wandb.login()] Loaded credentials for https://api.wandb.ai from /home/henry/.netrc.
wandb: Currently logged in as: hbfreed to https://api.wandb.ai. Use `wandb login --relogin` to force relogin
wandb: Tracking run with wandb version 0.28.0
wandb: Run data is saved locally in outputs/healed/grid_math/glean_keep75_s1225/wandb/run-20260716_142818-9x4iij2d
wandb: Run `wandb offline` to turn off syncing.
wandb: Syncing run glean-math-keep75-s1225
wandb: ⭐️ View project at https://wandb.ai/hbfreed/glean-grid
wandb: πŸš€ View run at https://wandb.ai/hbfreed/glean-grid/runs/9x4iij2d
Loading checkpoint shards: 0%| | 0/3 [00:00<?, ?it/s] Loading checkpoint shards: 33%|β–ˆβ–ˆβ–ˆβ–Ž | 1/3 [00:01<00:03, 1.53s/it] Loading checkpoint shards: 67%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 2/3 [00:03<00:01, 1.59s/it] Loading checkpoint shards: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 3/3 [00:03<00:00, 1.04it/s] Loading checkpoint shards: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 3/3 [00:03<00:00, 1.13s/it]
resumed student weights from outputs/healed/grid_math/glean_keep75_s1225/step0100 (fresh optimizer, step counter at 0)
12115 cached top-128 chat trajectories / 6,476,634 unique tokens | 53 steps/epoch | 150 total steps | student params 5.31B | teacher overlap=False
restored optimizer/scheduler state from step 100; rebuilt 260 paged buffers
{"step": 101, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.016109059700369834, "tokens": 120000, "cumulative_loss_tokens": 12120000, "grad_norm": 0.2041015625, "lr": 3e-05, "finish_rate": 0.798, "comp_len": 538.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 63.3, "frames": {"chat": 223}, "mem_gb": 21.95}
{"step": 102, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.016362527111552966, "tokens": 120000, "cumulative_loss_tokens": 12240000, "grad_norm": 0.21875, "lr": 3e-05, "finish_rate": 0.772, "comp_len": 582.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.6, "frames": {"chat": 206}, "mem_gb": 22.1}
{"step": 103, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.013859456993810212, "tokens": 120000, "cumulative_loss_tokens": 12360000, "grad_norm": 0.228515625, "lr": 3e-05, "finish_rate": 0.784, "comp_len": 563.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.6, "frames": {"chat": 213}, "mem_gb": 22.02}
{"step": 104, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.018646482431787688, "tokens": 120000, "cumulative_loss_tokens": 12480000, "grad_norm": 0.22265625, "lr": 3e-05, "finish_rate": 0.843, "comp_len": 538.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.9, "frames": {"chat": 223}, "mem_gb": 21.96}
{"step": 105, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.015217417630545484, "tokens": 120000, "cumulative_loss_tokens": 12600000, "grad_norm": 0.201171875, "lr": 3e-05, "finish_rate": 0.828, "comp_len": 528.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.6, "frames": {"chat": 227}, "mem_gb": 22.07}
{"step": 106, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.018008409859096478, "tokens": 120000, "cumulative_loss_tokens": 12720000, "grad_norm": 0.255859375, "lr": 3e-05, "finish_rate": 0.889, "comp_len": 474.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.6, "frames": {"chat": 253}, "mem_gb": 22.09}
{"step": 107, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012912440449073135, "tokens": 120000, "cumulative_loss_tokens": 12840000, "grad_norm": 0.1865234375, "lr": 3e-05, "finish_rate": 0.792, "comp_len": 555.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.6, "frames": {"chat": 216}, "mem_gb": 22.1}
{"step": 108, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010693734145350754, "tokens": 120000, "cumulative_loss_tokens": 12960000, "grad_norm": 0.1748046875, "lr": 3e-05, "finish_rate": 0.766, "comp_len": 585.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.3, "frames": {"chat": 205}, "mem_gb": 22.07}
{"step": 109, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.015156030047448197, "tokens": 120000, "cumulative_loss_tokens": 13080000, "grad_norm": 0.2138671875, "lr": 3e-05, "finish_rate": 0.729, "comp_len": 579.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.3, "frames": {"chat": 207}, "mem_gb": 22.16}
{"step": 110, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.013672215391354015, "tokens": 120000, "cumulative_loss_tokens": 13200000, "grad_norm": 0.1982421875, "lr": 3e-05, "finish_rate": 0.814, "comp_len": 558.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.5, "frames": {"chat": 215}, "mem_gb": 22.08}
The attention mask is not set and cannot be inferred from input because pad token is same as eos token. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.
[eval step 110] sample: 'To solve this problem, we need to analyze the spiral pattern of numbers from 1 to 49 arranged on a square grid and identify the four shaded squares that lie on the same diagonal as the number 7. We th'
{"step": 111, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011061214675944453, "tokens": 120000, "cumulative_loss_tokens": 13320000, "grad_norm": 0.2060546875, "lr": 3e-05, "finish_rate": 0.86, "comp_len": 526.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.2, "frames": {"chat": 228}, "mem_gb": 22.1}
{"step": 112, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012901123281742912, "tokens": 120000, "cumulative_loss_tokens": 13440000, "grad_norm": 0.1982421875, "lr": 3e-05, "finish_rate": 0.747, "comp_len": 543.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.9, "frames": {"chat": 221}, "mem_gb": 22.14}
{"step": 113, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009574843731914492, "tokens": 120000, "cumulative_loss_tokens": 13560000, "grad_norm": 0.15625, "lr": 3e-05, "finish_rate": 0.882, "comp_len": 472.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.2, "frames": {"chat": 254}, "mem_gb": 21.93}
{"step": 114, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011826626441131036, "tokens": 120000, "cumulative_loss_tokens": 13680000, "grad_norm": 0.2392578125, "lr": 3e-05, "finish_rate": 0.843, "comp_len": 571.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.7, "frames": {"chat": 210}, "mem_gb": 22.06}
{"step": 115, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010666260581545066, "tokens": 120000, "cumulative_loss_tokens": 13800000, "grad_norm": 0.166015625, "lr": 3e-05, "finish_rate": 0.827, "comp_len": 531.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.6, "frames": {"chat": 226}, "mem_gb": 22.02}
{"step": 116, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01181182961029117, "tokens": 120000, "cumulative_loss_tokens": 13920000, "grad_norm": 0.193359375, "lr": 3e-05, "finish_rate": 0.802, "comp_len": 566.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.3, "frames": {"chat": 212}, "mem_gb": 22.09}
{"step": 117, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.014291961643429628, "tokens": 120000, "cumulative_loss_tokens": 14040000, "grad_norm": 0.20703125, "lr": 3e-05, "finish_rate": 0.754, "comp_len": 568.7, "t_data_s": 0.2, "t_rollout_s": 0.0, "t_step_s": 47.3, "frames": {"chat": 211}, "mem_gb": 22.02}
{"step": 118, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011632958819546426, "tokens": 120000, "cumulative_loss_tokens": 14160000, "grad_norm": 0.1640625, "lr": 3e-05, "finish_rate": 0.776, "comp_len": 612.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 42.9, "frames": {"chat": 196}, "mem_gb": 22.07}
{"step": 119, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011126866567115454, "tokens": 120000, "cumulative_loss_tokens": 14280000, "grad_norm": 0.2119140625, "lr": 3e-05, "finish_rate": 0.811, "comp_len": 566.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.0, "frames": {"chat": 212}, "mem_gb": 22.09}
{"step": 120, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010413994966812121, "tokens": 120000, "cumulative_loss_tokens": 14400000, "grad_norm": 0.16015625, "lr": 3e-05, "finish_rate": 0.877, "comp_len": 491.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.0, "frames": {"chat": 244}, "mem_gb": 22.0}
[eval step 120] sample: "To solve this problem, we need to understand the structure of the spiral pattern and identify the numbers on the same diagonal as the number 7. Let's break down the problem step-by-step:\n\n1. **Underst"
{"step": 121, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010520029380288906, "tokens": 120000, "cumulative_loss_tokens": 14520000, "grad_norm": 0.166015625, "lr": 3e-05, "finish_rate": 0.838, "comp_len": 540.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.8, "frames": {"chat": 222}, "mem_gb": 22.05}
{"step": 122, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011724793996859807, "tokens": 120000, "cumulative_loss_tokens": 14640000, "grad_norm": 0.1884765625, "lr": 3e-05, "finish_rate": 0.78, "comp_len": 550.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.6, "frames": {"chat": 218}, "mem_gb": 22.09}
{"step": 123, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011794728558894713, "tokens": 120000, "cumulative_loss_tokens": 14760000, "grad_norm": 0.1982421875, "lr": 3e-05, "finish_rate": 0.913, "comp_len": 476.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.8, "frames": {"chat": 252}, "mem_gb": 21.97}
{"step": 124, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012576708099556466, "tokens": 120000, "cumulative_loss_tokens": 14880000, "grad_norm": 0.19140625, "lr": 3e-05, "finish_rate": 0.728, "comp_len": 594.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.5, "frames": {"chat": 202}, "mem_gb": 22.14}
{"step": 125, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012525027039841128, "tokens": 120000, "cumulative_loss_tokens": 15000000, "grad_norm": 0.44140625, "lr": 3e-05, "finish_rate": 0.835, "comp_len": 506.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.6, "frames": {"chat": 237}, "mem_gb": 22.1}
{"step": 126, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011847056540916674, "tokens": 120000, "cumulative_loss_tokens": 15120000, "grad_norm": 0.166015625, "lr": 3e-05, "finish_rate": 0.868, "comp_len": 512.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.5, "frames": {"chat": 234}, "mem_gb": 22.08}
{"step": 127, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010580201309620558, "tokens": 120000, "cumulative_loss_tokens": 15240000, "grad_norm": 0.193359375, "lr": 3e-05, "finish_rate": 0.809, "comp_len": 558.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.7, "frames": {"chat": 215}, "mem_gb": 22.1}
{"step": 128, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010746679509648433, "tokens": 120000, "cumulative_loss_tokens": 15360000, "grad_norm": 0.1796875, "lr": 3e-05, "finish_rate": 0.876, "comp_len": 512.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.7, "frames": {"chat": 234}, "mem_gb": 22.03}
{"step": 129, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010033142010107016, "tokens": 120000, "cumulative_loss_tokens": 15480000, "grad_norm": 0.1884765625, "lr": 3e-05, "finish_rate": 0.801, "comp_len": 555.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.8, "frames": {"chat": 216}, "mem_gb": 22.08}
{"step": 130, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01059437686605379, "tokens": 120000, "cumulative_loss_tokens": 15600000, "grad_norm": 0.1787109375, "lr": 3e-05, "finish_rate": 0.805, "comp_len": 571.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.3, "frames": {"chat": 210}, "mem_gb": 22.05}
[eval step 130] sample: 'To solve this problem, we need to arrange the numbers from 1 to 49 in a spiral pattern on a square grid starting from the center. We then identify the four shaded squares that lie on the same diagonal'
{"step": 131, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011416423269287528, "tokens": 120000, "cumulative_loss_tokens": 15720000, "grad_norm": 0.169921875, "lr": 3e-05, "finish_rate": 0.719, "comp_len": 603.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.1, "frames": {"chat": 199}, "mem_gb": 22.09}
{"step": 132, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011045920558762736, "tokens": 120000, "cumulative_loss_tokens": 15840000, "grad_norm": 0.1865234375, "lr": 3e-05, "finish_rate": 0.824, "comp_len": 571.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.4, "frames": {"chat": 210}, "mem_gb": 22.11}
{"step": 133, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010189920482278103, "tokens": 120000, "cumulative_loss_tokens": 15960000, "grad_norm": 0.1689453125, "lr": 3e-05, "finish_rate": 0.902, "comp_len": 533.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.7, "frames": {"chat": 225}, "mem_gb": 22.05}
{"step": 134, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011348977557197213, "tokens": 120000, "cumulative_loss_tokens": 16080000, "grad_norm": 0.1728515625, "lr": 3e-05, "finish_rate": 0.913, "comp_len": 474.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.6, "frames": {"chat": 253}, "mem_gb": 21.95}
{"step": 135, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010840521045137818, "tokens": 120000, "cumulative_loss_tokens": 16200000, "grad_norm": 0.1513671875, "lr": 3e-05, "finish_rate": 0.903, "comp_len": 485.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.1, "frames": {"chat": 247}, "mem_gb": 22.07}
{"step": 136, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01431039827777228, "tokens": 120000, "cumulative_loss_tokens": 16320000, "grad_norm": 0.2021484375, "lr": 3e-05, "finish_rate": 0.836, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.4, "frames": {"chat": 238}, "mem_gb": 22.07}
{"step": 137, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012493417616401954, "tokens": 120000, "cumulative_loss_tokens": 16440000, "grad_norm": 0.16796875, "lr": 3e-05, "finish_rate": 0.86, "comp_len": 510.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.6, "frames": {"chat": 235}, "mem_gb": 22.09}
{"step": 138, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011775391096648916, "tokens": 120000, "cumulative_loss_tokens": 16560000, "grad_norm": 0.1640625, "lr": 3e-05, "finish_rate": 0.805, "comp_len": 558.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.3, "frames": {"chat": 215}, "mem_gb": 22.06}
{"step": 139, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011272327020104665, "tokens": 120000, "cumulative_loss_tokens": 16680000, "grad_norm": 0.177734375, "lr": 3e-05, "finish_rate": 0.925, "comp_len": 447.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.0, "frames": {"chat": 268}, "mem_gb": 22.06}
{"step": 140, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011195119415794033, "tokens": 120000, "cumulative_loss_tokens": 16800000, "grad_norm": 0.1689453125, "lr": 3e-05, "finish_rate": 0.825, "comp_len": 526.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.1, "frames": {"chat": 228}, "mem_gb": 22.09}
[eval step 140] sample: 'To solve this problem, we need to arrange the numbers from 1 to 49 in a spiral pattern on a square grid and identify the four shaded squares that lie on the same diagonal as the number 7. Then, we wil'
{"step": 141, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01146084621019351, "tokens": 120000, "cumulative_loss_tokens": 16920000, "grad_norm": 0.1513671875, "lr": 3e-05, "finish_rate": 0.881, "comp_len": 476.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.6, "frames": {"chat": 252}, "mem_gb": 22.03}
{"step": 142, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01266345674659824, "tokens": 120000, "cumulative_loss_tokens": 17040000, "grad_norm": 0.1875, "lr": 3e-05, "finish_rate": 0.821, "comp_len": 538.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.1, "frames": {"chat": 223}, "mem_gb": 22.11}
{"step": 143, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012728927766492901, "tokens": 120000, "cumulative_loss_tokens": 17160000, "grad_norm": 0.1689453125, "lr": 3e-05, "finish_rate": 0.805, "comp_len": 531.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.6, "frames": {"chat": 226}, "mem_gb": 22.09}
{"step": 144, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.015653314464636303, "tokens": 120000, "cumulative_loss_tokens": 17280000, "grad_norm": 0.2421875, "lr": 3e-05, "finish_rate": 0.731, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.1, "frames": {"chat": 208}, "mem_gb": 22.14}
{"step": 145, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.00990355064412579, "tokens": 120000, "cumulative_loss_tokens": 17400000, "grad_norm": 0.154296875, "lr": 3e-05, "finish_rate": 0.883, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.6, "frames": {"chat": 240}, "mem_gb": 22.03}
{"step": 146, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011352584307268262, "tokens": 120000, "cumulative_loss_tokens": 17520000, "grad_norm": 0.171875, "lr": 3e-05, "finish_rate": 0.842, "comp_len": 540.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.9, "frames": {"chat": 222}, "mem_gb": 22.02}
{"step": 147, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009270905922904301, "tokens": 120000, "cumulative_loss_tokens": 17640000, "grad_norm": 0.1396484375, "lr": 3e-05, "finish_rate": 0.881, "comp_len": 508.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.6, "frames": {"chat": 236}, "mem_gb": 22.09}
{"step": 148, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010650625815118353, "tokens": 120000, "cumulative_loss_tokens": 17760000, "grad_norm": 0.181640625, "lr": 3e-05, "finish_rate": 0.834, "comp_len": 553.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.3, "frames": {"chat": 217}, "mem_gb": 22.06}
{"step": 149, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010312878351037702, "tokens": 120000, "cumulative_loss_tokens": 17880000, "grad_norm": 0.1611328125, "lr": 3e-05, "finish_rate": 0.921, "comp_len": 476.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.2, "frames": {"chat": 252}, "mem_gb": 21.97}
{"step": 150, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009814306650865667, "tokens": 120000, "cumulative_loss_tokens": 18000000, "grad_norm": 0.14453125, "lr": 3e-05, "finish_rate": 0.847, "comp_len": 540.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.7, "frames": {"chat": 222}, "mem_gb": 22.08}
[eval step 150] sample: 'To solve this problem, we need to arrange the numbers from 1 to 49 in a spiral pattern on a square grid and identify the four numbers that lie on the same diagonal as the number 7. We then need to det'
checkpoint snapshot queued -> outputs/healed/grid_math/glean_keep75_s1225/step0150
wandb: updating run metadata
wandb: uploading output.log; uploading wandb-summary.json; uploading config.yaml
wandb:
wandb: Run history:
wandb: comp_len β–…β–‡β–†β–…β–„β–†β–‡β–‡β–†β–„β–†β–…β–†β–†β–ˆβ–ƒβ–…β–…β–‚β–‡β–„β–†β–„β–†β–†β–†β–…β–‚β–ƒβ–ƒβ–†β–β–‚β–…β–…β–ƒβ–…β–„β–…β–…
wandb: cumulative_loss_tokens β–β–β–β–β–‚β–‚β–‚β–‚β–‚β–‚β–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–„β–„β–„β–„β–„β–…β–…β–…β–…β–…β–†β–†β–†β–†β–†β–†β–‡β–‡β–‡β–‡β–‡β–‡β–ˆβ–ˆβ–ˆ
wandb: epoch β–β–β–β–β–β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
wandb: finish_rate β–ƒβ–ƒβ–ƒβ–…β–…β–ƒβ–‚β–β–„β–‚β–…β–…β–„β–‚β–ƒβ–†β–…β–ƒβ–ˆβ–β–†β–„β–†β–„β–„β–„β–‡β–ˆβ–‡β–…β–„β–ˆβ–„β–†β–„β–β–…β–†β–…β–…
wandb: forward_topk_kl β–†β–†β–„β–ˆβ–…β–„β–‚β–…β–„β–‚β–β–ƒβ–‚β–ƒβ–…β–‚β–‚β–ƒβ–ƒβ–ƒβ–ƒβ–‚β–‚β–‚β–‚β–‚β–‚β–ƒβ–‚β–…β–ƒβ–‚β–‚β–ƒβ–„β–β–ƒβ–β–‚β–
wandb: grad_norm β–‚β–ƒβ–ƒβ–ƒβ–‚β–‚β–‚β–ƒβ–‚β–ƒβ–β–‚β–‚β–‚β–ƒβ–‚β–‚β–‚β–‚β–ˆβ–‚β–‚β–‚β–‚β–‚β–‚β–‚β–β–‚β–‚β–‚β–‚β–β–‚β–‚β–β–‚β–β–‚β–
wandb: lr ▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁
wandb: mem_gb β–‚β–†β–„β–‚β–…β–†β–…β–ˆβ–†β–†β–β–…β–„β–…β–†β–…β–†β–‚β–‡β–†β–†β–„β–†β–…β–†β–…β–‚β–…β–…β–†β–…β–†β–„β–†β–†β–„β–„β–†β–…β–†
wandb: step β–β–β–β–β–‚β–‚β–‚β–‚β–‚β–‚β–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–„β–„β–„β–„β–„β–…β–…β–…β–…β–…β–†β–†β–†β–†β–†β–†β–‡β–‡β–‡β–‡β–‡β–‡β–ˆβ–ˆβ–ˆ
wandb: t_data_s β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–ˆβ–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–
wandb: +3 ...
wandb:
wandb: Run summary:
wandb: comp_len 540.5
wandb: cumulative_loss_tokens 18000000
wandb: epoch 2
wandb: finish_rate 0.847
wandb: forward_topk_kl 0.00981
wandb: grad_norm 0.14453
wandb: lr 3e-05
wandb: mem_gb 22.08
wandb: step 150
wandb: t_data_s 0
wandb: +4 ...
wandb:
wandb: πŸš€ View run glean-math-keep75-s1225 at: https://wandb.ai/hbfreed/glean-grid/runs/9x4iij2d
wandb: ⭐️ View project at: https://wandb.ai/hbfreed/glean-grid
wandb: Synced 5 W&B file(s), 0 media file(s), 0 artifact file(s) and 0 other file(s)
wandb: Find logs at: outputs/healed/grid_math/glean_keep75_s1225/wandb/run-20260716_142818-9x4iij2d/logs
{
"correct": 910,
"accuracy": 0.6899166034874905,
"finished": 1315,
"finish_rate": 0.9969673995451099,
"mean_completion_tokens": 113.55724033358605
}
saved item-level results -> outputs/evals/grid_math/glean_keep75_s1225_step100_chat.json
{
"correct": 909,
"accuracy": 0.689158453373768,
"finished": 1316,
"finish_rate": 0.9977255496588324,
"mean_completion_tokens": 114.47536012130402
}
saved item-level results -> outputs/evals/grid_math/glean_keep75_s1225_step150_chat.json