variable-reap-archive / healed /grid_math /glean_keep75_s1226.console.log
hbfreed's picture
Add files using upload-large-folder tool
d6fae6a verified
Raw
History Blame Contribute Delete
22.3 kB
/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
warnings.warn('Grouped GEMM not available.')
wandb: [wandb.login()] Loaded credentials for https://api.wandb.ai from /home/henry/.netrc.
wandb: Currently logged in as: hbfreed to https://api.wandb.ai. Use `wandb login --relogin` to force relogin
wandb: Tracking run with wandb version 0.28.0
wandb: Run data is saved locally in outputs/healed/grid_math/glean_keep75_s1226/wandb/run-20260716_142818-kru5sldj
wandb: Run `wandb offline` to turn off syncing.
wandb: Syncing run glean-math-keep75-s1226
wandb: ⭐️ View project at https://wandb.ai/hbfreed/glean-grid
wandb: πŸš€ View run at https://wandb.ai/hbfreed/glean-grid/runs/kru5sldj
Loading checkpoint shards: 0%| | 0/3 [00:00<?, ?it/s] Loading checkpoint shards: 33%|β–ˆβ–ˆβ–ˆβ–Ž | 1/3 [00:01<00:02, 1.35s/it] Loading checkpoint shards: 67%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 2/3 [00:02<00:01, 1.51s/it] Loading checkpoint shards: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 3/3 [00:03<00:00, 1.09it/s] Loading checkpoint shards: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 3/3 [00:03<00:00, 1.06s/it]
resumed student weights from outputs/healed/grid_math/glean_keep75_s1226/step0100 (fresh optimizer, step counter at 0)
12115 cached top-128 chat trajectories / 6,476,634 unique tokens | 53 steps/epoch | 150 total steps | student params 5.31B | teacher overlap=False
restored optimizer/scheduler state from step 100; rebuilt 260 paged buffers
{"step": 101, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.017711289564648177, "tokens": 120000, "cumulative_loss_tokens": 12120000, "grad_norm": 0.22265625, "lr": 3e-05, "finish_rate": 0.858, "comp_len": 517.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 62.0, "frames": {"chat": 232}, "mem_gb": 21.89}
{"step": 102, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.017982081608990362, "tokens": 120000, "cumulative_loss_tokens": 12240000, "grad_norm": 0.2578125, "lr": 3e-05, "finish_rate": 0.832, "comp_len": 545.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.6, "frames": {"chat": 220}, "mem_gb": 22.09}
{"step": 103, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.0159690285191716, "tokens": 120000, "cumulative_loss_tokens": 12360000, "grad_norm": 0.1982421875, "lr": 3e-05, "finish_rate": 0.776, "comp_len": 571.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.5, "frames": {"chat": 210}, "mem_gb": 22.14}
{"step": 104, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.014078767126984894, "tokens": 120000, "cumulative_loss_tokens": 12480000, "grad_norm": 0.1845703125, "lr": 3e-05, "finish_rate": 0.81, "comp_len": 531.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.1, "frames": {"chat": 226}, "mem_gb": 22.06}
{"step": 105, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.015146795610602325, "tokens": 120000, "cumulative_loss_tokens": 12600000, "grad_norm": 0.248046875, "lr": 3e-05, "finish_rate": 0.741, "comp_len": 566.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 43.3, "frames": {"chat": 212}, "mem_gb": 22.09}
{"step": 106, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.014050660304430251, "tokens": 120000, "cumulative_loss_tokens": 12720000, "grad_norm": 0.201171875, "lr": 3e-05, "finish_rate": 0.839, "comp_len": 508.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.2, "frames": {"chat": 236}, "mem_gb": 22.1}
{"step": 107, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012824523648739948, "tokens": 120000, "cumulative_loss_tokens": 12840000, "grad_norm": 0.193359375, "lr": 3e-05, "finish_rate": 0.928, "comp_len": 454.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.7, "frames": {"chat": 264}, "mem_gb": 21.97}
{"step": 108, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.017827476510805233, "tokens": 120000, "cumulative_loss_tokens": 12960000, "grad_norm": 0.2421875, "lr": 3e-05, "finish_rate": 0.834, "comp_len": 524.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.3, "frames": {"chat": 229}, "mem_gb": 22.07}
{"step": 109, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01010443297145733, "tokens": 120000, "cumulative_loss_tokens": 13080000, "grad_norm": 0.16015625, "lr": 3e-05, "finish_rate": 0.903, "comp_len": 465.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.8, "frames": {"chat": 258}, "mem_gb": 21.95}
{"step": 110, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.013931514149834403, "tokens": 120000, "cumulative_loss_tokens": 13200000, "grad_norm": 0.19140625, "lr": 3e-05, "finish_rate": 0.755, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.0, "frames": {"chat": 208}, "mem_gb": 22.11}
The attention mask is not set and cannot be inferred from input because pad token is same as eos token. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.
[eval step 110] sample: 'To solve this problem, we need to understand the geometric properties involved. When the midpoints of the sides of a triangle are connected, the segments joining these midpoints form a smaller triangl'
{"step": 111, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010035461849397204, "tokens": 120000, "cumulative_loss_tokens": 13320000, "grad_norm": 0.16015625, "lr": 3e-05, "finish_rate": 0.88, "comp_len": 481.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 45.9, "frames": {"chat": 249}, "mem_gb": 22.02}
{"step": 112, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01025480254034434, "tokens": 120000, "cumulative_loss_tokens": 13440000, "grad_norm": 0.1796875, "lr": 3e-05, "finish_rate": 0.845, "comp_len": 545.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 42.8, "frames": {"chat": 220}, "mem_gb": 22.09}
{"step": 113, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010523743329072991, "tokens": 120000, "cumulative_loss_tokens": 13560000, "grad_norm": 0.1708984375, "lr": 3e-05, "finish_rate": 0.834, "comp_len": 538.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 43.4, "frames": {"chat": 223}, "mem_gb": 22.08}
{"step": 114, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01153164215181023, "tokens": 120000, "cumulative_loss_tokens": 13680000, "grad_norm": 0.17578125, "lr": 3e-05, "finish_rate": 0.833, "comp_len": 543.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 43.8, "frames": {"chat": 221}, "mem_gb": 22.09}
{"step": 115, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009408667080748516, "tokens": 120000, "cumulative_loss_tokens": 13800000, "grad_norm": 0.1552734375, "lr": 3e-05, "finish_rate": 0.9, "comp_len": 521.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.4, "frames": {"chat": 230}, "mem_gb": 21.99}
{"step": 116, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011410523822395286, "tokens": 120000, "cumulative_loss_tokens": 13920000, "grad_norm": 0.1630859375, "lr": 3e-05, "finish_rate": 0.776, "comp_len": 560.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 44.1, "frames": {"chat": 214}, "mem_gb": 22.07}
{"step": 117, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01616842130373698, "tokens": 120000, "cumulative_loss_tokens": 14040000, "grad_norm": 0.224609375, "lr": 3e-05, "finish_rate": 0.766, "comp_len": 560.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 46.1, "frames": {"chat": 214}, "mem_gb": 22.08}
{"step": 118, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.013691102667020944, "tokens": 120000, "cumulative_loss_tokens": 14160000, "grad_norm": 0.19921875, "lr": 3e-05, "finish_rate": 0.786, "comp_len": 571.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.8, "frames": {"chat": 210}, "mem_gb": 22.13}
{"step": 119, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.013348402527160942, "tokens": 120000, "cumulative_loss_tokens": 14280000, "grad_norm": 0.1884765625, "lr": 3e-05, "finish_rate": 0.776, "comp_len": 560.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 51.0, "frames": {"chat": 214}, "mem_gb": 22.09}
{"step": 120, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011879650606094704, "tokens": 120000, "cumulative_loss_tokens": 14400000, "grad_norm": 0.2001953125, "lr": 3e-05, "finish_rate": 0.791, "comp_len": 558.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 52.1, "frames": {"chat": 215}, "mem_gb": 22.05}
[eval step 120] sample: 'To solve this problem, we need to understand the geometric properties involved when the midpoints of the sides of a triangle are connected by segments. This process creates a new triangle, known as th'
{"step": 121, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.013821079181631406, "tokens": 120000, "cumulative_loss_tokens": 14520000, "grad_norm": 0.1845703125, "lr": 3e-05, "finish_rate": 0.721, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 51.3, "frames": {"chat": 208}, "mem_gb": 22.09}
{"step": 122, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010628624554680815, "tokens": 120000, "cumulative_loss_tokens": 14640000, "grad_norm": 0.1708984375, "lr": 3e-05, "finish_rate": 0.789, "comp_len": 550.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.7, "frames": {"chat": 218}, "mem_gb": 21.97}
{"step": 123, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010161227444757242, "tokens": 120000, "cumulative_loss_tokens": 14760000, "grad_norm": 0.1640625, "lr": 3e-05, "finish_rate": 0.876, "comp_len": 515.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 50.1, "frames": {"chat": 233}, "mem_gb": 21.99}
{"step": 124, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009278946231456938, "tokens": 120000, "cumulative_loss_tokens": 14880000, "grad_norm": 0.16015625, "lr": 3e-05, "finish_rate": 0.861, "comp_len": 519.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.3, "frames": {"chat": 231}, "mem_gb": 22.03}
{"step": 125, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011464818606327754, "tokens": 120000, "cumulative_loss_tokens": 15000000, "grad_norm": 0.1611328125, "lr": 3e-05, "finish_rate": 0.868, "comp_len": 510.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 51.0, "frames": {"chat": 235}, "mem_gb": 22.22}
{"step": 126, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010276864761835895, "tokens": 120000, "cumulative_loss_tokens": 15120000, "grad_norm": 0.1484375, "lr": 3e-05, "finish_rate": 0.843, "comp_len": 555.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.4, "frames": {"chat": 216}, "mem_gb": 22.08}
{"step": 127, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009667191798787098, "tokens": 120000, "cumulative_loss_tokens": 15240000, "grad_norm": 0.162109375, "lr": 3e-05, "finish_rate": 0.831, "comp_len": 506.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 50.3, "frames": {"chat": 237}, "mem_gb": 22.1}
{"step": 128, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.014880005238496233, "tokens": 120000, "cumulative_loss_tokens": 15360000, "grad_norm": 0.25, "lr": 3e-05, "finish_rate": 0.734, "comp_len": 591.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.5, "frames": {"chat": 203}, "mem_gb": 22.1}
{"step": 129, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010798121559165885, "tokens": 120000, "cumulative_loss_tokens": 15480000, "grad_norm": 0.166015625, "lr": 3e-05, "finish_rate": 0.873, "comp_len": 508.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 50.5, "frames": {"chat": 236}, "mem_gb": 22.13}
{"step": 130, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009982232672628015, "tokens": 120000, "cumulative_loss_tokens": 15600000, "grad_norm": 0.1513671875, "lr": 3e-05, "finish_rate": 0.734, "comp_len": 560.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.7, "frames": {"chat": 214}, "mem_gb": 22.1}
[eval step 130] sample: 'To solve this problem, we need to understand the geometric properties involved when the midpoints of the sides of a triangle are connected by segments. This process creates a new triangle, known as th'
{"step": 131, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01276522857361318, "tokens": 120000, "cumulative_loss_tokens": 15720000, "grad_norm": 0.1826171875, "lr": 3e-05, "finish_rate": 0.78, "comp_len": 574.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 50.5, "frames": {"chat": 209}, "mem_gb": 22.09}
{"step": 132, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01355358442418122, "tokens": 120000, "cumulative_loss_tokens": 15840000, "grad_norm": 0.2216796875, "lr": 3e-05, "finish_rate": 0.906, "comp_len": 468.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 53.2, "frames": {"chat": 256}, "mem_gb": 22.1}
{"step": 133, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.00979031468022537, "tokens": 120000, "cumulative_loss_tokens": 15960000, "grad_norm": 0.1591796875, "lr": 3e-05, "finish_rate": 0.878, "comp_len": 521.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 50.9, "frames": {"chat": 230}, "mem_gb": 21.96}
{"step": 134, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012052156168699731, "tokens": 120000, "cumulative_loss_tokens": 16080000, "grad_norm": 0.193359375, "lr": 3e-05, "finish_rate": 0.822, "comp_len": 521.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 51.4, "frames": {"chat": 230}, "mem_gb": 22.15}
{"step": 135, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011693784629782506, "tokens": 120000, "cumulative_loss_tokens": 16200000, "grad_norm": 0.1640625, "lr": 3e-05, "finish_rate": 0.881, "comp_len": 528.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 50.3, "frames": {"chat": 227}, "mem_gb": 22.05}
{"step": 136, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011806925121663758, "tokens": 120000, "cumulative_loss_tokens": 16320000, "grad_norm": 0.150390625, "lr": 3e-05, "finish_rate": 0.755, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.8, "frames": {"chat": 208}, "mem_gb": 22.11}
{"step": 137, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011330220682158445, "tokens": 120000, "cumulative_loss_tokens": 16440000, "grad_norm": 0.166015625, "lr": 3e-05, "finish_rate": 0.699, "comp_len": 582.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 50.6, "frames": {"chat": 206}, "mem_gb": 22.12}
{"step": 138, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.010682711551792455, "tokens": 120000, "cumulative_loss_tokens": 16560000, "grad_norm": 0.15625, "lr": 3e-05, "finish_rate": 0.82, "comp_len": 526.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 50.0, "frames": {"chat": 228}, "mem_gb": 22.0}
{"step": 139, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011466035542547858, "tokens": 120000, "cumulative_loss_tokens": 16680000, "grad_norm": 0.15234375, "lr": 3e-05, "finish_rate": 0.835, "comp_len": 535.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 50.1, "frames": {"chat": 224}, "mem_gb": 22.09}
{"step": 140, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.012643014844014155, "tokens": 120000, "cumulative_loss_tokens": 16800000, "grad_norm": 0.236328125, "lr": 3e-05, "finish_rate": 0.66, "comp_len": 600.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.2, "frames": {"chat": 200}, "mem_gb": 22.13}
[eval step 140] sample: 'To solve this problem, we need to understand the geometric properties involved when the midpoints of the sides of a triangle are connected by segments. This process creates a new triangle, known as th'
{"step": 141, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01080410843101951, "tokens": 120000, "cumulative_loss_tokens": 16920000, "grad_norm": 0.1689453125, "lr": 3e-05, "finish_rate": 0.714, "comp_len": 612.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 47.3, "frames": {"chat": 196}, "mem_gb": 22.11}
{"step": 142, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009731826097378507, "tokens": 120000, "cumulative_loss_tokens": 17040000, "grad_norm": 0.140625, "lr": 3e-05, "finish_rate": 0.834, "comp_len": 538.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.8, "frames": {"chat": 223}, "mem_gb": 22.09}
{"step": 143, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011139519218046916, "tokens": 120000, "cumulative_loss_tokens": 17160000, "grad_norm": 0.2119140625, "lr": 3e-05, "finish_rate": 0.869, "comp_len": 563.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 48.4, "frames": {"chat": 213}, "mem_gb": 21.98}
{"step": 144, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009973561418584237, "tokens": 120000, "cumulative_loss_tokens": 17280000, "grad_norm": 0.146484375, "lr": 3e-05, "finish_rate": 0.879, "comp_len": 517.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 50.8, "frames": {"chat": 232}, "mem_gb": 22.02}
{"step": 145, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009199376441648928, "tokens": 120000, "cumulative_loss_tokens": 17400000, "grad_norm": 0.16015625, "lr": 3e-05, "finish_rate": 0.861, "comp_len": 538.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.9, "frames": {"chat": 223}, "mem_gb": 22.02}
{"step": 146, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.0111230931228473, "tokens": 120000, "cumulative_loss_tokens": 17520000, "grad_norm": 0.1611328125, "lr": 3e-05, "finish_rate": 0.85, "comp_len": 515.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 50.5, "frames": {"chat": 233}, "mem_gb": 22.11}
{"step": 147, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.011176060675464882, "tokens": 120000, "cumulative_loss_tokens": 17640000, "grad_norm": 0.1474609375, "lr": 3e-05, "finish_rate": 0.816, "comp_len": 553.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.2, "frames": {"chat": 217}, "mem_gb": 22.11}
{"step": 148, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.01491896515111827, "tokens": 120000, "cumulative_loss_tokens": 17760000, "grad_norm": 0.185546875, "lr": 3e-05, "finish_rate": 0.752, "comp_len": 594.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 49.8, "frames": {"chat": 202}, "mem_gb": 22.17}
{"step": 149, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.00973438565802838, "tokens": 120000, "cumulative_loss_tokens": 17880000, "grad_norm": 0.1484375, "lr": 3e-05, "finish_rate": 0.858, "comp_len": 474.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 52.5, "frames": {"chat": 253}, "mem_gb": 22.03}
{"step": 150, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.009605388223500147, "tokens": 120000, "cumulative_loss_tokens": 18000000, "grad_norm": 0.1552734375, "lr": 3e-05, "finish_rate": 0.879, "comp_len": 519.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 50.2, "frames": {"chat": 231}, "mem_gb": 22.03}
[eval step 150] sample: 'To solve this problem, we need to understand the geometric properties involved when the midpoints of the sides of a triangle are connected by segments. This process creates a new triangle, known as th'
checkpoint snapshot queued -> outputs/healed/grid_math/glean_keep75_s1226/step0150
wandb: updating run metadata
wandb: uploading output.log
wandb:
wandb: Run history:
wandb: comp_len β–ƒβ–…β–„β–†β–ƒβ–β–†β–‚β–…β–„β–„β–†β–†β–†β–†β–†β–…β–ƒβ–„β–ƒβ–ƒβ–‡β–ƒβ–†β–†β–„β–„β–„β–†β–‡β–„β–‡β–ˆβ–„β–†β–„β–ƒβ–…β–‡β–„
wandb: cumulative_loss_tokens β–β–β–β–β–‚β–‚β–‚β–‚β–‚β–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–„β–„β–„β–„β–„β–…β–…β–…β–…β–…β–†β–†β–†β–†β–†β–†β–‡β–‡β–‡β–‡β–‡β–‡β–ˆβ–ˆβ–ˆ
wandb: epoch β–β–β–β–β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
wandb: finish_rate β–†β–…β–„β–…β–ƒβ–ˆβ–†β–‡β–ƒβ–†β–†β–‡β–„β–„β–„β–„β–ƒβ–„β–‡β–†β–†β–…β–ƒβ–‡β–ƒβ–‡β–‡β–…β–‡β–ƒβ–…β–†β–β–‚β–†β–†β–†β–…β–ƒβ–‡
wandb: forward_topk_kl β–ˆβ–ˆβ–†β–…β–†β–„β–ˆβ–‚β–…β–‚β–‚β–ƒβ–ƒβ–…β–„β–…β–‚β–‚β–β–ƒβ–β–†β–‚β–‚β–„β–β–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–„β–‚β–β–ƒβ–β–ƒβ–ƒβ–†β–
wandb: grad_norm β–†β–ˆβ–„β–„β–‡β–„β–‡β–‚β–„β–‚β–ƒβ–‚β–‚β–†β–…β–…β–„β–ƒβ–‚β–‚β–β–‚β–ˆβ–ƒβ–‚β–†β–‚β–„β–‚β–‚β–‚β–‡β–ƒβ–β–…β–‚β–‚β–β–„β–‚
wandb: lr ▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁
wandb: mem_gb β–β–…β–†β–…β–…β–ƒβ–…β–‚β–†β–„β–…β–…β–ƒβ–…β–…β–…β–„β–…β–ƒβ–ƒβ–ˆβ–…β–…β–…β–†β–…β–‚β–‡β–„β–†β–ƒβ–…β–†β–†β–…β–„β–„β–†β–†β–„
wandb: step β–β–β–β–β–‚β–‚β–‚β–‚β–‚β–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–„β–„β–„β–„β–„β–…β–…β–…β–…β–…β–…β–†β–†β–†β–†β–†β–‡β–‡β–‡β–‡β–‡β–‡β–ˆβ–ˆβ–ˆ
wandb: t_data_s ▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁
wandb: +3 ...
wandb:
wandb: Run summary:
wandb: comp_len 519.5
wandb: cumulative_loss_tokens 18000000
wandb: epoch 2
wandb: finish_rate 0.879
wandb: forward_topk_kl 0.00961
wandb: grad_norm 0.15527
wandb: lr 3e-05
wandb: mem_gb 22.03
wandb: step 150
wandb: t_data_s 0
wandb: +4 ...
wandb:
wandb: πŸš€ View run glean-math-keep75-s1226 at: https://wandb.ai/hbfreed/glean-grid/runs/kru5sldj
wandb: ⭐️ View project at: https://wandb.ai/hbfreed/glean-grid
wandb: Synced 5 W&B file(s), 0 media file(s), 0 artifact file(s) and 0 other file(s)
wandb: Find logs at: outputs/healed/grid_math/glean_keep75_s1226/wandb/run-20260716_142818-kru5sldj/logs
{
"correct": 915,
"accuracy": 0.6937073540561031,
"finished": 1314,
"finish_rate": 0.9962092494313874,
"mean_completion_tokens": 114.16527672479151
}
saved item-level results -> outputs/evals/grid_math/glean_keep75_s1226_step100_chat.json
{
"correct": 925,
"accuracy": 0.7012888551933283,
"finished": 1314,
"finish_rate": 0.9962092494313874,
"mean_completion_tokens": 115.14329037149355
}
saved item-level results -> outputs/evals/grid_math/glean_keep75_s1226_step150_chat.json