variable-reap-archive / healed /grid_math /glean_keep25_s1224.console.log
hbfreed's picture
Add files using upload-large-folder tool
d6fae6a verified
Raw
History Blame Contribute Delete
56.7 kB
/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
warnings.warn('Grouped GEMM not available.')
wandb: [wandb.login()] Loaded credentials for https://api.wandb.ai from /home/henry/.netrc.
wandb: Currently logged in as: hbfreed to https://api.wandb.ai. Use `wandb login --relogin` to force relogin
wandb: setting up run 97sk7igo
wandb: Tracking run with wandb version 0.28.0
wandb: Run data is saved locally in outputs/healed/grid_math/glean_keep25_s1224/wandb/run-20260716_040623-97sk7igo
wandb: Run `wandb offline` to turn off syncing.
wandb: Syncing run glean-math-keep25-s1224
wandb: ⭐️ View project at https://wandb.ai/hbfreed/glean-grid
wandb: πŸš€ View run at https://wandb.ai/hbfreed/glean-grid/runs/97sk7igo
12115 cached top-128 chat trajectories / 6,476,634 unique tokens | 53 steps/epoch | 150 total steps | student params 2.09B | teacher overlap=False
{"step": 1, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 1.4345772517378133, "tokens": 120000, "cumulative_loss_tokens": 120000, "grad_norm": 114.5, "lr": 6e-06, "finish_rate": 0.907, "comp_len": 508.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.6, "frames": {"chat": 236}, "mem_gb": 9.77}
The attention mask is not set and cannot be inferred from input because pad token is same as eos token. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.
[eval step 1] sample: '\nThe value of $p$ is the sum of $a$ and $b$. Find the value of $a + b$.\n\n\n### The value of $p$ is the sum of $a$ and $b$. Find the value of $a + b$.\n\n\nLet the value of'
{"step": 2, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 1.401542197600007, "tokens": 120000, "cumulative_loss_tokens": 240000, "grad_norm": 101.0, "lr": 9e-06, "finish_rate": 0.781, "comp_len": 558.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 31.7, "frames": {"chat": 215}, "mem_gb": 10.01}
{"step": 3, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 1.1813887482275565, "tokens": 120000, "cumulative_loss_tokens": 360000, "grad_norm": 63.75, "lr": 1.2e-05, "finish_rate": 0.825, "comp_len": 553.0, "t_data_s": 0.1, "t_rollout_s": 0.0, "t_step_s": 32.9, "frames": {"chat": 217}, "mem_gb": 9.88}
{"step": 4, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.9318435715690255, "tokens": 120000, "cumulative_loss_tokens": 480000, "grad_norm": 12.875, "lr": 1.5e-05, "finish_rate": 0.8, "comp_len": 585.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 33.8, "frames": {"chat": 205}, "mem_gb": 9.94}
{"step": 5, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.7287501879028976, "tokens": 120000, "cumulative_loss_tokens": 600000, "grad_norm": 8.4375, "lr": 1.8e-05, "finish_rate": 0.834, "comp_len": 524.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.4, "frames": {"chat": 229}, "mem_gb": 9.92}
{"step": 6, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.7777328455592195, "tokens": 120000, "cumulative_loss_tokens": 720000, "grad_norm": 23.625, "lr": 2.1e-05, "finish_rate": 0.812, "comp_len": 538.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.5, "frames": {"chat": 223}, "mem_gb": 9.98}
{"step": 7, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.5918175142496824, "tokens": 120000, "cumulative_loss_tokens": 840000, "grad_norm": 4.65625, "lr": 2.4e-05, "finish_rate": 0.708, "comp_len": 594.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.7, "frames": {"chat": 202}, "mem_gb": 10.02}
{"step": 8, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.5632855169591804, "tokens": 120000, "cumulative_loss_tokens": 960000, "grad_norm": 3.0, "lr": 2.7000000000000002e-05, "finish_rate": 0.77, "comp_len": 574.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.5, "frames": {"chat": 209}, "mem_gb": 9.99}
{"step": 9, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.40953814367316665, "tokens": 120000, "cumulative_loss_tokens": 1080000, "grad_norm": 2.03125, "lr": 3e-05, "finish_rate": 0.885, "comp_len": 528.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.3, "frames": {"chat": 227}, "mem_gb": 9.97}
{"step": 10, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.38932390790109833, "tokens": 120000, "cumulative_loss_tokens": 1200000, "grad_norm": 1.5703125, "lr": 3e-05, "finish_rate": 0.848, "comp_len": 521.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.4, "frames": {"chat": 230}, "mem_gb": 10.04}
[eval step 10] sample: 'To solve the problem, we need to determine the values of \\(a\\), \\(b\\), and \\(p\\) given the equations:\n\n1. \\(a + b = k\\)\n2. \\(k + m = p\\)\n3. \\(p + a = r\\)\n4. \\(b'
{"step": 11, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.3458594021844367, "tokens": 120000, "cumulative_loss_tokens": 1320000, "grad_norm": 1.2265625, "lr": 3e-05, "finish_rate": 0.879, "comp_len": 519.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.9, "frames": {"chat": 231}, "mem_gb": 9.89}
{"step": 12, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.3208039687448492, "tokens": 120000, "cumulative_loss_tokens": 1440000, "grad_norm": 1.0390625, "lr": 3e-05, "finish_rate": 0.882, "comp_len": 489.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.5, "frames": {"chat": 245}, "mem_gb": 9.97}
{"step": 13, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.2981381397678206, "tokens": 120000, "cumulative_loss_tokens": 1560000, "grad_norm": 0.921875, "lr": 3e-05, "finish_rate": 0.81, "comp_len": 571.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.2, "frames": {"chat": 210}, "mem_gb": 9.98}
{"step": 14, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.3454915365646283, "tokens": 120000, "cumulative_loss_tokens": 1680000, "grad_norm": 0.98828125, "lr": 3e-05, "finish_rate": 0.758, "comp_len": 568.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.3, "frames": {"chat": 211}, "mem_gb": 9.97}
{"step": 15, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.26953837820465365, "tokens": 120000, "cumulative_loss_tokens": 1800000, "grad_norm": 0.80859375, "lr": 3e-05, "finish_rate": 0.819, "comp_len": 543.0, "t_data_s": 0.1, "t_rollout_s": 0.0, "t_step_s": 36.1, "frames": {"chat": 221}, "mem_gb": 10.03}
{"step": 16, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.2559446799742058, "tokens": 120000, "cumulative_loss_tokens": 1920000, "grad_norm": 0.8671875, "lr": 3e-05, "finish_rate": 0.912, "comp_len": 480.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.2, "frames": {"chat": 250}, "mem_gb": 9.84}
{"step": 17, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.2738502737318476, "tokens": 120000, "cumulative_loss_tokens": 2040000, "grad_norm": 0.8046875, "lr": 3e-05, "finish_rate": 0.79, "comp_len": 524.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 39.8, "frames": {"chat": 229}, "mem_gb": 10.02}
{"step": 18, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.2143578802034259, "tokens": 120000, "cumulative_loss_tokens": 2160000, "grad_norm": 0.65625, "lr": 3e-05, "finish_rate": 0.888, "comp_len": 480.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 39.3, "frames": {"chat": 250}, "mem_gb": 9.99}
{"step": 19, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.24581084873105088, "tokens": 120000, "cumulative_loss_tokens": 2280000, "grad_norm": 0.73828125, "lr": 3e-05, "finish_rate": 0.844, "comp_len": 519.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 39.2, "frames": {"chat": 231}, "mem_gb": 9.87}
{"step": 20, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.22843340362496675, "tokens": 120000, "cumulative_loss_tokens": 2400000, "grad_norm": 0.8125, "lr": 3e-05, "finish_rate": 0.844, "comp_len": 535.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.6, "frames": {"chat": 224}, "mem_gb": 9.91}
[eval step 20] sample: 'To solve the problem, we need to determine the values of \\(a\\), \\(b\\), \\(k\\), \\(p\\), and \\(m\\) given the equations:\n\n1. \\(a + b = k\\)\n2. \\(k + m = p\\)\n3. \\(p + a'
{"step": 21, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.21071218295594057, "tokens": 120000, "cumulative_loss_tokens": 2520000, "grad_norm": 0.64453125, "lr": 3e-05, "finish_rate": 0.802, "comp_len": 566.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 33.2, "frames": {"chat": 212}, "mem_gb": 9.95}
{"step": 22, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.1929007501606519, "tokens": 120000, "cumulative_loss_tokens": 2640000, "grad_norm": 0.640625, "lr": 3e-05, "finish_rate": 0.87, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.5, "frames": {"chat": 238}, "mem_gb": 9.9}
{"step": 23, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.2146618340227753, "tokens": 120000, "cumulative_loss_tokens": 2760000, "grad_norm": 0.68359375, "lr": 3e-05, "finish_rate": 0.903, "comp_len": 466.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.9, "frames": {"chat": 257}, "mem_gb": 9.78}
{"step": 24, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.1882299357444669, "tokens": 120000, "cumulative_loss_tokens": 2880000, "grad_norm": 0.59375, "lr": 3e-05, "finish_rate": 0.868, "comp_len": 528.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.2, "frames": {"chat": 227}, "mem_gb": 9.97}
{"step": 25, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.21730685276426376, "tokens": 120000, "cumulative_loss_tokens": 3000000, "grad_norm": 0.65234375, "lr": 3e-05, "finish_rate": 0.838, "comp_len": 526.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.2, "frames": {"chat": 228}, "mem_gb": 10.0}
{"step": 26, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.21057547978740185, "tokens": 120000, "cumulative_loss_tokens": 3120000, "grad_norm": 0.56640625, "lr": 3e-05, "finish_rate": 0.803, "comp_len": 515.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.9, "frames": {"chat": 233}, "mem_gb": 9.99}
{"step": 27, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.1898427692937975, "tokens": 120000, "cumulative_loss_tokens": 3240000, "grad_norm": 0.6015625, "lr": 3e-05, "finish_rate": 0.863, "comp_len": 515.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.8, "frames": {"chat": 233}, "mem_gb": 9.99}
{"step": 28, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.25671816173580786, "tokens": 120000, "cumulative_loss_tokens": 3360000, "grad_norm": 0.65625, "lr": 3e-05, "finish_rate": 0.731, "comp_len": 609.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.6, "frames": {"chat": 197}, "mem_gb": 10.08}
{"step": 29, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.22320921461253115, "tokens": 120000, "cumulative_loss_tokens": 3480000, "grad_norm": 0.62109375, "lr": 3e-05, "finish_rate": 0.862, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.8, "frames": {"chat": 239}, "mem_gb": 9.84}
{"step": 30, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.21965685539674012, "tokens": 120000, "cumulative_loss_tokens": 3600000, "grad_norm": 0.7578125, "lr": 3e-05, "finish_rate": 0.83, "comp_len": 535.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.2, "frames": {"chat": 224}, "mem_gb": 9.88}
[eval step 30] sample: 'To solve this problem, we need to determine the values of \\(a\\), \\(b\\), \\(m\\), and \\(p\\) given the equations:\n\n\\[\n\\begin{align*}\na + b &= k \\\\\nk + m &= p \\\\\np + a &= r \\\\\nb'
{"step": 31, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.18031089373938738, "tokens": 120000, "cumulative_loss_tokens": 3720000, "grad_norm": 0.62109375, "lr": 3e-05, "finish_rate": 0.788, "comp_len": 553.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.7, "frames": {"chat": 217}, "mem_gb": 10.0}
{"step": 32, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.17237651355092723, "tokens": 120000, "cumulative_loss_tokens": 3840000, "grad_norm": 0.53515625, "lr": 3e-05, "finish_rate": 0.876, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.0, "frames": {"chat": 241}, "mem_gb": 10.0}
{"step": 33, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.17687927930454414, "tokens": 120000, "cumulative_loss_tokens": 3960000, "grad_norm": 0.546875, "lr": 3e-05, "finish_rate": 0.835, "comp_len": 550.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.1, "frames": {"chat": 218}, "mem_gb": 9.97}
{"step": 34, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.17552895108697314, "tokens": 120000, "cumulative_loss_tokens": 4080000, "grad_norm": 0.49609375, "lr": 3e-05, "finish_rate": 0.767, "comp_len": 582.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.1, "frames": {"chat": 206}, "mem_gb": 9.98}
{"step": 35, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.18765739145452778, "tokens": 120000, "cumulative_loss_tokens": 4200000, "grad_norm": 0.578125, "lr": 3e-05, "finish_rate": 0.845, "comp_len": 517.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 39.0, "frames": {"chat": 232}, "mem_gb": 10.03}
{"step": 36, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.19498228631795694, "tokens": 120000, "cumulative_loss_tokens": 4320000, "grad_norm": 0.51953125, "lr": 3e-05, "finish_rate": 0.771, "comp_len": 550.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 39.9, "frames": {"chat": 218}, "mem_gb": 10.04}
{"step": 37, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.19024058150469014, "tokens": 120000, "cumulative_loss_tokens": 4440000, "grad_norm": 0.474609375, "lr": 3e-05, "finish_rate": 0.779, "comp_len": 563.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.6, "frames": {"chat": 213}, "mem_gb": 10.0}
{"step": 38, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.19190885730143636, "tokens": 120000, "cumulative_loss_tokens": 4560000, "grad_norm": 0.53125, "lr": 3e-05, "finish_rate": 0.887, "comp_len": 483.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 39.8, "frames": {"chat": 248}, "mem_gb": 9.97}
{"step": 39, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.1669260192029178, "tokens": 120000, "cumulative_loss_tokens": 4680000, "grad_norm": 0.5390625, "lr": 3e-05, "finish_rate": 0.803, "comp_len": 550.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.7, "frames": {"chat": 218}, "mem_gb": 10.04}
{"step": 40, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.15053241867311298, "tokens": 120000, "cumulative_loss_tokens": 4800000, "grad_norm": 0.44921875, "lr": 3e-05, "finish_rate": 0.851, "comp_len": 543.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.2, "frames": {"chat": 221}, "mem_gb": 9.99}
[eval step 40] sample: 'To solve the problem, we need to determine the values of \\(a\\), \\(b\\), \\(m\\), \\(p\\), and \\(r\\) given the equations:\n\n\\[\na + b = k\n\\]\n\\[\nk + m = p\n\\]\n\\[\np + a ='
{"step": 41, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.15927721466608347, "tokens": 120000, "cumulative_loss_tokens": 4920000, "grad_norm": 0.46484375, "lr": 3e-05, "finish_rate": 0.894, "comp_len": 508.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.2, "frames": {"chat": 236}, "mem_gb": 9.92}
{"step": 42, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.16593128025457263, "tokens": 120000, "cumulative_loss_tokens": 5040000, "grad_norm": 0.50390625, "lr": 3e-05, "finish_rate": 0.858, "comp_len": 487.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 40.2, "frames": {"chat": 246}, "mem_gb": 9.84}
{"step": 43, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.16362612133746346, "tokens": 120000, "cumulative_loss_tokens": 5160000, "grad_norm": 0.46875, "lr": 3e-05, "finish_rate": 0.838, "comp_len": 512.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 39.4, "frames": {"chat": 234}, "mem_gb": 10.1}
{"step": 44, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.13735881909814973, "tokens": 120000, "cumulative_loss_tokens": 5280000, "grad_norm": 0.466796875, "lr": 3e-05, "finish_rate": 0.748, "comp_len": 594.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.1, "frames": {"chat": 202}, "mem_gb": 9.98}
{"step": 45, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.15262007395885885, "tokens": 120000, "cumulative_loss_tokens": 5400000, "grad_norm": 0.44140625, "lr": 3e-05, "finish_rate": 0.811, "comp_len": 553.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.7, "frames": {"chat": 217}, "mem_gb": 9.99}
{"step": 46, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.15682636960850407, "tokens": 120000, "cumulative_loss_tokens": 5520000, "grad_norm": 0.46484375, "lr": 3e-05, "finish_rate": 0.866, "comp_len": 535.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.0, "frames": {"chat": 224}, "mem_gb": 9.99}
{"step": 47, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.17931298827622086, "tokens": 120000, "cumulative_loss_tokens": 5640000, "grad_norm": 0.51171875, "lr": 3e-05, "finish_rate": 0.753, "comp_len": 558.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.8, "frames": {"chat": 215}, "mem_gb": 10.01}
{"step": 48, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.1443968278159077, "tokens": 120000, "cumulative_loss_tokens": 5760000, "grad_norm": 0.451171875, "lr": 3e-05, "finish_rate": 0.884, "comp_len": 463.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.6, "frames": {"chat": 259}, "mem_gb": 9.92}
{"step": 49, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.1630018111831819, "tokens": 120000, "cumulative_loss_tokens": 5880000, "grad_norm": 0.4765625, "lr": 3e-05, "finish_rate": 0.829, "comp_len": 571.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.6, "frames": {"chat": 210}, "mem_gb": 9.99}
{"step": 50, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.1874668618524447, "tokens": 120000, "cumulative_loss_tokens": 6000000, "grad_norm": 0.50390625, "lr": 3e-05, "finish_rate": 0.77, "comp_len": 563.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.3, "frames": {"chat": 213}, "mem_gb": 10.05}
[eval step 50] sample: 'To solve the problem, we need to determine the values of \\(a\\), \\(b\\), \\(m\\), \\(p\\), and \\(k\\) given the equations:\n\n\\[\n\\begin{align*}\na + b &= k \\\\\nk + m &= p \\\\\np + a &='
checkpoint snapshot queued -> outputs/healed/grid_math/glean_keep25_s1224/step0050
{"step": 51, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.15269769354338447, "tokens": 120000, "cumulative_loss_tokens": 6120000, "grad_norm": 0.453125, "lr": 3e-05, "finish_rate": 0.815, "comp_len": 540.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 32.0, "frames": {"chat": 222}, "mem_gb": 9.95}
{"step": 52, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.13562167353965343, "tokens": 120000, "cumulative_loss_tokens": 6240000, "grad_norm": 0.431640625, "lr": 3e-05, "finish_rate": 0.889, "comp_len": 510.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.1, "frames": {"chat": 235}, "mem_gb": 10.01}
{"step": 53, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.1533853217214346, "tokens": 120000, "cumulative_loss_tokens": 6360000, "grad_norm": 0.423828125, "lr": 3e-05, "finish_rate": 0.798, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.0, "frames": {"chat": 208}, "mem_gb": 9.96}
{"step": 54, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.1419351307667171, "tokens": 120000, "cumulative_loss_tokens": 6480000, "grad_norm": 0.458984375, "lr": 3e-05, "finish_rate": 0.733, "comp_len": 628.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.6, "frames": {"chat": 191}, "mem_gb": 10.0}
{"step": 55, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.11651230888419474, "tokens": 120000, "cumulative_loss_tokens": 6600000, "grad_norm": 0.39453125, "lr": 3e-05, "finish_rate": 0.845, "comp_len": 547.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.9, "frames": {"chat": 219}, "mem_gb": 10.0}
{"step": 56, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.11036905342560882, "tokens": 120000, "cumulative_loss_tokens": 6720000, "grad_norm": 0.3984375, "lr": 3e-05, "finish_rate": 0.778, "comp_len": 579.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.5, "frames": {"chat": 207}, "mem_gb": 10.0}
{"step": 57, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.15549946540420254, "tokens": 120000, "cumulative_loss_tokens": 6840000, "grad_norm": 0.5, "lr": 3e-05, "finish_rate": 0.755, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.4, "frames": {"chat": 208}, "mem_gb": 9.96}
{"step": 58, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10701905750250444, "tokens": 120000, "cumulative_loss_tokens": 6960000, "grad_norm": 0.353515625, "lr": 3e-05, "finish_rate": 0.799, "comp_len": 547.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.8, "frames": {"chat": 219}, "mem_gb": 10.0}
{"step": 59, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.1119194057648691, "tokens": 120000, "cumulative_loss_tokens": 7080000, "grad_norm": 0.40234375, "lr": 3e-05, "finish_rate": 0.915, "comp_len": 487.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 39.1, "frames": {"chat": 246}, "mem_gb": 9.87}
{"step": 60, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.14874430537996813, "tokens": 120000, "cumulative_loss_tokens": 7200000, "grad_norm": 0.443359375, "lr": 3e-05, "finish_rate": 0.704, "comp_len": 582.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.6, "frames": {"chat": 206}, "mem_gb": 10.02}
[eval step 60] sample: 'To solve the problem, we need to determine the values of \\(a\\), \\(b\\), \\(m\\), \\(p\\), and \\(k\\) given the equations:\n\n\\[\n\\begin{align*}\na + b &= k \\\\\nk + m &= p \\\\\np + a &='
{"step": 61, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.11569037632470329, "tokens": 120000, "cumulative_loss_tokens": 7320000, "grad_norm": 0.419921875, "lr": 3e-05, "finish_rate": 0.876, "comp_len": 515.0, "t_data_s": 0.1, "t_rollout_s": 0.0, "t_step_s": 34.3, "frames": {"chat": 233}, "mem_gb": 10.0}
{"step": 62, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.12268456889608254, "tokens": 120000, "cumulative_loss_tokens": 7440000, "grad_norm": 0.421875, "lr": 3e-05, "finish_rate": 0.847, "comp_len": 524.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.1, "frames": {"chat": 229}, "mem_gb": 9.87}
{"step": 63, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.09885425716893126, "tokens": 120000, "cumulative_loss_tokens": 7560000, "grad_norm": 0.3515625, "lr": 3e-05, "finish_rate": 0.864, "comp_len": 508.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.6, "frames": {"chat": 236}, "mem_gb": 9.9}
{"step": 64, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.12324718012257169, "tokens": 120000, "cumulative_loss_tokens": 7680000, "grad_norm": 0.369140625, "lr": 3e-05, "finish_rate": 0.87, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.2, "frames": {"chat": 239}, "mem_gb": 9.79}
{"step": 65, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10336422332112367, "tokens": 120000, "cumulative_loss_tokens": 7800000, "grad_norm": 0.326171875, "lr": 3e-05, "finish_rate": 0.867, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.0, "frames": {"chat": 241}, "mem_gb": 9.91}
{"step": 66, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.11680785041383157, "tokens": 120000, "cumulative_loss_tokens": 7920000, "grad_norm": 0.36328125, "lr": 3e-05, "finish_rate": 0.863, "comp_len": 531.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.0, "frames": {"chat": 226}, "mem_gb": 9.87}
{"step": 67, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10119038897647212, "tokens": 120000, "cumulative_loss_tokens": 8040000, "grad_norm": 0.6171875, "lr": 3e-05, "finish_rate": 0.893, "comp_len": 512.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.9, "frames": {"chat": 234}, "mem_gb": 10.0}
{"step": 68, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10523584497369205, "tokens": 120000, "cumulative_loss_tokens": 8160000, "grad_norm": 0.361328125, "lr": 3e-05, "finish_rate": 0.914, "comp_len": 466.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.9, "frames": {"chat": 257}, "mem_gb": 9.99}
{"step": 69, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.14938643554880593, "tokens": 120000, "cumulative_loss_tokens": 8280000, "grad_norm": 0.4140625, "lr": 3e-05, "finish_rate": 0.76, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 39.8, "frames": {"chat": 208}, "mem_gb": 10.05}
{"step": 70, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.12613096526097506, "tokens": 120000, "cumulative_loss_tokens": 8400000, "grad_norm": 0.408203125, "lr": 3e-05, "finish_rate": 0.763, "comp_len": 568.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.0, "frames": {"chat": 211}, "mem_gb": 10.02}
[eval step 70] sample: "To solve the problem, we need to determine the values of \\(a\\), \\(b\\), \\(m\\), \\(p\\), and \\(r\\) such that the given equations hold true. Let's break down the problem into manageable steps:\n\n1. **Unders"
{"step": 71, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.1367481416762496, "tokens": 120000, "cumulative_loss_tokens": 8520000, "grad_norm": 0.466796875, "lr": 3e-05, "finish_rate": 0.806, "comp_len": 528.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.7, "frames": {"chat": 227}, "mem_gb": 10.0}
{"step": 72, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.12594295182231194, "tokens": 120000, "cumulative_loss_tokens": 8640000, "grad_norm": 0.408203125, "lr": 3e-05, "finish_rate": 0.796, "comp_len": 568.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.8, "frames": {"chat": 211}, "mem_gb": 9.98}
{"step": 73, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10491307358372336, "tokens": 120000, "cumulative_loss_tokens": 8760000, "grad_norm": 0.373046875, "lr": 3e-05, "finish_rate": 0.861, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.5, "frames": {"chat": 238}, "mem_gb": 10.0}
{"step": 74, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10930286093257989, "tokens": 120000, "cumulative_loss_tokens": 8880000, "grad_norm": 0.361328125, "lr": 3e-05, "finish_rate": 0.835, "comp_len": 506.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 40.0, "frames": {"chat": 237}, "mem_gb": 10.04}
{"step": 75, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.12709836459675183, "tokens": 120000, "cumulative_loss_tokens": 9000000, "grad_norm": 0.37109375, "lr": 3e-05, "finish_rate": 0.721, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.2, "frames": {"chat": 208}, "mem_gb": 10.04}
{"step": 76, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.11188181798982745, "tokens": 120000, "cumulative_loss_tokens": 9120000, "grad_norm": 0.34765625, "lr": 3e-05, "finish_rate": 0.801, "comp_len": 543.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.8, "frames": {"chat": 221}, "mem_gb": 10.12}
{"step": 77, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.11114264213865002, "tokens": 120000, "cumulative_loss_tokens": 9240000, "grad_norm": 0.369140625, "lr": 3e-05, "finish_rate": 0.853, "comp_len": 517.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.9, "frames": {"chat": 232}, "mem_gb": 9.96}
{"step": 78, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10996809244149675, "tokens": 120000, "cumulative_loss_tokens": 9360000, "grad_norm": 0.337890625, "lr": 3e-05, "finish_rate": 0.764, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.0, "frames": {"chat": 208}, "mem_gb": 9.99}
{"step": 79, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10375189399402589, "tokens": 120000, "cumulative_loss_tokens": 9480000, "grad_norm": 0.609375, "lr": 3e-05, "finish_rate": 0.837, "comp_len": 528.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.1, "frames": {"chat": 227}, "mem_gb": 9.91}
{"step": 80, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.11647727334791173, "tokens": 120000, "cumulative_loss_tokens": 9600000, "grad_norm": 0.392578125, "lr": 3e-05, "finish_rate": 0.824, "comp_len": 543.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.9, "frames": {"chat": 221}, "mem_gb": 9.94}
[eval step 80] sample: "To solve this problem, we need to determine the values of \\(a\\), \\(b\\), \\(m\\), \\(p\\), and \\(k\\) such that the given equations hold true. Let's break down the problem step-by-step:\n\n1. **Define Variabl"
{"step": 81, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.09754915808814889, "tokens": 120000, "cumulative_loss_tokens": 9720000, "grad_norm": 0.318359375, "lr": 3e-05, "finish_rate": 0.815, "comp_len": 517.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.5, "frames": {"chat": 232}, "mem_gb": 10.01}
{"step": 82, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10922176310662181, "tokens": 120000, "cumulative_loss_tokens": 9840000, "grad_norm": 0.359375, "lr": 3e-05, "finish_rate": 0.822, "comp_len": 547.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.6, "frames": {"chat": 219}, "mem_gb": 10.01}
{"step": 83, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.11440609233131012, "tokens": 120000, "cumulative_loss_tokens": 9960000, "grad_norm": 0.357421875, "lr": 3e-05, "finish_rate": 0.713, "comp_len": 615.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.7, "frames": {"chat": 195}, "mem_gb": 10.1}
{"step": 84, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.11024705316449206, "tokens": 120000, "cumulative_loss_tokens": 10080000, "grad_norm": 0.35546875, "lr": 3e-05, "finish_rate": 0.833, "comp_len": 555.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.0, "frames": {"chat": 216}, "mem_gb": 10.0}
{"step": 85, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.11196330596281526, "tokens": 120000, "cumulative_loss_tokens": 10200000, "grad_norm": 0.400390625, "lr": 3e-05, "finish_rate": 0.788, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.9, "frames": {"chat": 208}, "mem_gb": 9.89}
{"step": 86, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10395217997139941, "tokens": 120000, "cumulative_loss_tokens": 10320000, "grad_norm": 0.375, "lr": 3e-05, "finish_rate": 0.919, "comp_len": 510.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.8, "frames": {"chat": 235}, "mem_gb": 9.88}
{"step": 87, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.11349097539822882, "tokens": 120000, "cumulative_loss_tokens": 10440000, "grad_norm": 0.435546875, "lr": 3e-05, "finish_rate": 0.853, "comp_len": 533.3, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.9, "frames": {"chat": 225}, "mem_gb": 9.99}
{"step": 88, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.13914468128886073, "tokens": 120000, "cumulative_loss_tokens": 10560000, "grad_norm": 0.384765625, "lr": 3e-05, "finish_rate": 0.77, "comp_len": 563.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.8, "frames": {"chat": 213}, "mem_gb": 10.08}
{"step": 89, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.09369528885387506, "tokens": 120000, "cumulative_loss_tokens": 10680000, "grad_norm": 0.3515625, "lr": 3e-05, "finish_rate": 0.922, "comp_len": 466.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.3, "frames": {"chat": 257}, "mem_gb": 9.76}
{"step": 90, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.12627653089668603, "tokens": 120000, "cumulative_loss_tokens": 10800000, "grad_norm": 0.38671875, "lr": 3e-05, "finish_rate": 0.792, "comp_len": 566.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.3, "frames": {"chat": 212}, "mem_gb": 10.03}
[eval step 90] sample: "To solve this problem, we need to determine the values of \\(a\\), \\(b\\), \\(m\\), \\(p\\), and \\(k\\) such that the given equations hold true. Let's break down the problem step-by-step:\n\n1. **Understand the"
{"step": 91, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10226257944550986, "tokens": 120000, "cumulative_loss_tokens": 10920000, "grad_norm": 0.36328125, "lr": 3e-05, "finish_rate": 0.833, "comp_len": 543.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 33.9, "frames": {"chat": 221}, "mem_gb": 10.0}
{"step": 92, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.09869740050801386, "tokens": 120000, "cumulative_loss_tokens": 11040000, "grad_norm": 0.326171875, "lr": 3e-05, "finish_rate": 0.868, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.4, "frames": {"chat": 242}, "mem_gb": 10.0}
{"step": 93, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.09069403918914808, "tokens": 120000, "cumulative_loss_tokens": 11160000, "grad_norm": 0.3125, "lr": 3e-05, "finish_rate": 0.836, "comp_len": 545.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.9, "frames": {"chat": 220}, "mem_gb": 9.96}
{"step": 94, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.09289198616895204, "tokens": 120000, "cumulative_loss_tokens": 11280000, "grad_norm": 0.33984375, "lr": 3e-05, "finish_rate": 0.896, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.9, "frames": {"chat": 240}, "mem_gb": 9.86}
{"step": 95, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10285916468321035, "tokens": 120000, "cumulative_loss_tokens": 11400000, "grad_norm": 1.140625, "lr": 3e-05, "finish_rate": 0.728, "comp_len": 582.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.8, "frames": {"chat": 206}, "mem_gb": 9.99}
{"step": 96, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.11189375975215808, "tokens": 120000, "cumulative_loss_tokens": 11520000, "grad_norm": 0.36328125, "lr": 3e-05, "finish_rate": 0.867, "comp_len": 531.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.1, "frames": {"chat": 226}, "mem_gb": 10.0}
{"step": 97, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.14419980257482579, "tokens": 120000, "cumulative_loss_tokens": 11640000, "grad_norm": 0.478515625, "lr": 3e-05, "finish_rate": 0.877, "comp_len": 491.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.6, "frames": {"chat": 244}, "mem_gb": 9.79}
{"step": 98, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.1254452561319495, "tokens": 120000, "cumulative_loss_tokens": 11760000, "grad_norm": 0.412109375, "lr": 3e-05, "finish_rate": 0.804, "comp_len": 535.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.3, "frames": {"chat": 224}, "mem_gb": 10.01}
{"step": 99, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.09918995484355837, "tokens": 120000, "cumulative_loss_tokens": 11880000, "grad_norm": 0.349609375, "lr": 3e-05, "finish_rate": 0.923, "comp_len": 442.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 39.2, "frames": {"chat": 271}, "mem_gb": 9.73}
{"step": 100, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10652838590508328, "tokens": 120000, "cumulative_loss_tokens": 12000000, "grad_norm": 0.361328125, "lr": 3e-05, "finish_rate": 0.856, "comp_len": 508.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.0, "frames": {"chat": 236}, "mem_gb": 10.01}
[eval step 100] sample: "To solve the problem, we need to determine the values of \\(a\\), \\(b\\), \\(m\\), \\(p\\), and \\(r\\) such that the given equations hold true. Let's break down the problem step-by-step:\n\n1. **Understand the "
checkpoint snapshot queued -> outputs/healed/grid_math/glean_keep25_s1224/step0100
{"step": 101, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10529813103003738, "tokens": 120000, "cumulative_loss_tokens": 12120000, "grad_norm": 0.369140625, "lr": 3e-05, "finish_rate": 0.841, "comp_len": 517.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.2, "frames": {"chat": 232}, "mem_gb": 9.88}
{"step": 102, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.09428196538443057, "tokens": 120000, "cumulative_loss_tokens": 12240000, "grad_norm": 0.33984375, "lr": 3e-05, "finish_rate": 0.79, "comp_len": 571.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.3, "frames": {"chat": 210}, "mem_gb": 9.94}
{"step": 103, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.09178287912054608, "tokens": 120000, "cumulative_loss_tokens": 12360000, "grad_norm": 0.33984375, "lr": 3e-05, "finish_rate": 0.811, "comp_len": 553.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.0, "frames": {"chat": 217}, "mem_gb": 9.9}
{"step": 104, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.1119527683553286, "tokens": 120000, "cumulative_loss_tokens": 12480000, "grad_norm": 0.375, "lr": 3e-05, "finish_rate": 0.839, "comp_len": 535.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.4, "frames": {"chat": 224}, "mem_gb": 10.02}
{"step": 105, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.1298952497580399, "tokens": 120000, "cumulative_loss_tokens": 12600000, "grad_norm": 0.388671875, "lr": 3e-05, "finish_rate": 0.749, "comp_len": 591.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.6, "frames": {"chat": 203}, "mem_gb": 9.87}
{"step": 106, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.10422293208353221, "tokens": 120000, "cumulative_loss_tokens": 12720000, "grad_norm": 0.345703125, "lr": 3e-05, "finish_rate": 0.887, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.1, "frames": {"chat": 239}, "mem_gb": 9.97}
{"step": 107, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07617318357517942, "tokens": 120000, "cumulative_loss_tokens": 12840000, "grad_norm": 0.31640625, "lr": 3e-05, "finish_rate": 0.902, "comp_len": 472.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.5, "frames": {"chat": 254}, "mem_gb": 9.88}
{"step": 108, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07282634121378263, "tokens": 120000, "cumulative_loss_tokens": 12960000, "grad_norm": 0.291015625, "lr": 3e-05, "finish_rate": 0.876, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.8, "frames": {"chat": 241}, "mem_gb": 9.98}
{"step": 109, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.108472445377366, "tokens": 120000, "cumulative_loss_tokens": 13080000, "grad_norm": 0.34375, "lr": 3e-05, "finish_rate": 0.746, "comp_len": 563.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.2, "frames": {"chat": 213}, "mem_gb": 10.01}
{"step": 110, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.08850192172794293, "tokens": 120000, "cumulative_loss_tokens": 13200000, "grad_norm": 0.30859375, "lr": 3e-05, "finish_rate": 0.864, "comp_len": 543.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.6, "frames": {"chat": 221}, "mem_gb": 10.05}
[eval step 110] sample: 'To solve the problem, we need to determine the values of \\(a\\), \\(b\\), \\(m\\), \\(r\\), and \\(p\\) such that each letter represents a non-zero digit and satisfies the given equations:\n\n\\[\n\\begin{align*}\na'
{"step": 111, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.09740933078067998, "tokens": 120000, "cumulative_loss_tokens": 13320000, "grad_norm": 0.341796875, "lr": 3e-05, "finish_rate": 0.745, "comp_len": 612.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 30.5, "frames": {"chat": 196}, "mem_gb": 10.01}
{"step": 112, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07808580766382317, "tokens": 120000, "cumulative_loss_tokens": 13440000, "grad_norm": 0.294921875, "lr": 3e-05, "finish_rate": 0.926, "comp_len": 444.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.3, "frames": {"chat": 270}, "mem_gb": 9.82}
{"step": 113, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07404585122984524, "tokens": 120000, "cumulative_loss_tokens": 13560000, "grad_norm": 0.294921875, "lr": 3e-05, "finish_rate": 0.815, "comp_len": 555.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 32.9, "frames": {"chat": 216}, "mem_gb": 9.99}
{"step": 114, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.08850063937480251, "tokens": 120000, "cumulative_loss_tokens": 13680000, "grad_norm": 0.30078125, "lr": 3e-05, "finish_rate": 0.775, "comp_len": 600.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 31.3, "frames": {"chat": 200}, "mem_gb": 9.96}
{"step": 115, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07580333673724284, "tokens": 120000, "cumulative_loss_tokens": 13800000, "grad_norm": 0.30078125, "lr": 3e-05, "finish_rate": 0.767, "comp_len": 582.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 32.2, "frames": {"chat": 206}, "mem_gb": 9.91}
{"step": 116, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.0730737258046555, "tokens": 120000, "cumulative_loss_tokens": 13920000, "grad_norm": 0.279296875, "lr": 3e-05, "finish_rate": 0.902, "comp_len": 512.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 33.1, "frames": {"chat": 234}, "mem_gb": 9.95}
{"step": 117, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.08553545155295482, "tokens": 120000, "cumulative_loss_tokens": 14040000, "grad_norm": 0.318359375, "lr": 3e-05, "finish_rate": 0.823, "comp_len": 558.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 31.9, "frames": {"chat": 215}, "mem_gb": 9.96}
{"step": 118, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07238297607673642, "tokens": 120000, "cumulative_loss_tokens": 14160000, "grad_norm": 0.291015625, "lr": 3e-05, "finish_rate": 0.922, "comp_len": 470.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 33.9, "frames": {"chat": 255}, "mem_gb": 9.94}
{"step": 119, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07544196644419184, "tokens": 120000, "cumulative_loss_tokens": 14280000, "grad_norm": 0.29296875, "lr": 3e-05, "finish_rate": 0.892, "comp_len": 480.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.0, "frames": {"chat": 250}, "mem_gb": 9.82}
{"step": 120, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07670339237060397, "tokens": 120000, "cumulative_loss_tokens": 14400000, "grad_norm": 0.291015625, "lr": 3e-05, "finish_rate": 0.884, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 33.6, "frames": {"chat": 242}, "mem_gb": 9.99}
[eval step 120] sample: 'To solve the problem, we need to determine the values of \\(a\\), \\(b\\), \\(m\\), \\(r\\), and \\(p\\) such that each letter represents a non-zero digit and satisfies the given equations:\n\n\\[\n\\begin{align*}\na'
{"step": 121, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.0937881386723059, "tokens": 120000, "cumulative_loss_tokens": 14520000, "grad_norm": 0.314453125, "lr": 3e-05, "finish_rate": 0.729, "comp_len": 603.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 32.5, "frames": {"chat": 199}, "mem_gb": 10.0}
{"step": 122, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.11198460068336377, "tokens": 120000, "cumulative_loss_tokens": 14640000, "grad_norm": 0.375, "lr": 3e-05, "finish_rate": 0.784, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.5, "frames": {"chat": 208}, "mem_gb": 10.04}
{"step": 123, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07519433575168562, "tokens": 120000, "cumulative_loss_tokens": 14760000, "grad_norm": 0.31640625, "lr": 3e-05, "finish_rate": 0.764, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 32.8, "frames": {"chat": 208}, "mem_gb": 9.97}
{"step": 124, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.09269439431062589, "tokens": 120000, "cumulative_loss_tokens": 14880000, "grad_norm": 0.314453125, "lr": 3e-05, "finish_rate": 0.732, "comp_len": 574.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.5, "frames": {"chat": 209}, "mem_gb": 10.12}
{"step": 125, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.06901245591950914, "tokens": 120000, "cumulative_loss_tokens": 15000000, "grad_norm": 0.287109375, "lr": 3e-05, "finish_rate": 0.855, "comp_len": 510.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.6, "frames": {"chat": 235}, "mem_gb": 9.96}
{"step": 126, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07502137167186787, "tokens": 120000, "cumulative_loss_tokens": 15120000, "grad_norm": 0.294921875, "lr": 3e-05, "finish_rate": 0.74, "comp_len": 588.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.1, "frames": {"chat": 204}, "mem_gb": 9.95}
{"step": 127, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.09983784483978525, "tokens": 120000, "cumulative_loss_tokens": 15240000, "grad_norm": 0.3359375, "lr": 3e-05, "finish_rate": 0.745, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.7, "frames": {"chat": 208}, "mem_gb": 10.01}
{"step": 128, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07167404105997023, "tokens": 120000, "cumulative_loss_tokens": 15360000, "grad_norm": 0.271484375, "lr": 3e-05, "finish_rate": 0.825, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.4, "frames": {"chat": 240}, "mem_gb": 10.0}
{"step": 129, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07390924430101489, "tokens": 120000, "cumulative_loss_tokens": 15480000, "grad_norm": 0.3046875, "lr": 3e-05, "finish_rate": 0.89, "comp_len": 487.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.4, "frames": {"chat": 246}, "mem_gb": 9.99}
{"step": 130, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.08214766127246742, "tokens": 120000, "cumulative_loss_tokens": 15600000, "grad_norm": 0.318359375, "lr": 3e-05, "finish_rate": 0.909, "comp_len": 493.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.0, "frames": {"chat": 243}, "mem_gb": 9.81}
[eval step 130] sample: 'To solve the problem, we need to determine the values of \\(a\\), \\(b\\), \\(m\\), \\(r\\), and \\(p\\) such that each letter represents a non-zero digit and satisfies the given equations:\n\n\\[\n\\begin{align*}\na'
{"step": 131, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.09511748868422583, "tokens": 120000, "cumulative_loss_tokens": 15720000, "grad_norm": 0.32421875, "lr": 3e-05, "finish_rate": 0.745, "comp_len": 576.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 32.2, "frames": {"chat": 208}, "mem_gb": 10.01}
{"step": 132, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.09182360006729141, "tokens": 120000, "cumulative_loss_tokens": 15840000, "grad_norm": 0.31640625, "lr": 3e-05, "finish_rate": 0.817, "comp_len": 547.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.4, "frames": {"chat": 219}, "mem_gb": 10.0}
{"step": 133, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.10137045447360724, "tokens": 120000, "cumulative_loss_tokens": 15960000, "grad_norm": 0.333984375, "lr": 3e-05, "finish_rate": 0.782, "comp_len": 568.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.6, "frames": {"chat": 211}, "mem_gb": 10.01}
{"step": 134, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.08328102445462719, "tokens": 120000, "cumulative_loss_tokens": 16080000, "grad_norm": 0.32421875, "lr": 3e-05, "finish_rate": 0.862, "comp_len": 517.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.4, "frames": {"chat": 232}, "mem_gb": 9.97}
{"step": 135, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.10078765981163208, "tokens": 120000, "cumulative_loss_tokens": 16200000, "grad_norm": 0.33984375, "lr": 3e-05, "finish_rate": 0.804, "comp_len": 560.7, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.4, "frames": {"chat": 214}, "mem_gb": 10.01}
{"step": 136, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07711706876040747, "tokens": 120000, "cumulative_loss_tokens": 16320000, "grad_norm": 0.287109375, "lr": 3e-05, "finish_rate": 0.819, "comp_len": 531.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.1, "frames": {"chat": 226}, "mem_gb": 9.9}
{"step": 137, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07596234880803773, "tokens": 120000, "cumulative_loss_tokens": 16440000, "grad_norm": 0.3046875, "lr": 3e-05, "finish_rate": 0.819, "comp_len": 571.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.6, "frames": {"chat": 210}, "mem_gb": 10.01}
{"step": 138, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.06859002932261986, "tokens": 120000, "cumulative_loss_tokens": 16560000, "grad_norm": 0.29296875, "lr": 3e-05, "finish_rate": 0.858, "comp_len": 550.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.6, "frames": {"chat": 218}, "mem_gb": 9.83}
{"step": 139, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.06959025802219597, "tokens": 120000, "cumulative_loss_tokens": 16680000, "grad_norm": 0.27734375, "lr": 3e-05, "finish_rate": 0.858, "comp_len": 515.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.3, "frames": {"chat": 233}, "mem_gb": 9.99}
{"step": 140, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.09598405006804193, "tokens": 120000, "cumulative_loss_tokens": 16800000, "grad_norm": 0.3125, "lr": 3e-05, "finish_rate": 0.786, "comp_len": 558.1, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.3, "frames": {"chat": 215}, "mem_gb": 10.01}
[eval step 140] sample: "To solve the problem, we need to determine the values of \\(a\\), \\(b\\), \\(m\\), \\(r\\), and \\(p\\) such that the given equations hold true. Let's break down the problem step-by-step and use Python with Sy"
{"step": 141, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.08247211161882927, "tokens": 120000, "cumulative_loss_tokens": 16920000, "grad_norm": 0.310546875, "lr": 3e-05, "finish_rate": 0.845, "comp_len": 515.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.0, "frames": {"chat": 233}, "mem_gb": 9.99}
{"step": 142, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07299311864568542, "tokens": 120000, "cumulative_loss_tokens": 17040000, "grad_norm": 0.279296875, "lr": 3e-05, "finish_rate": 0.766, "comp_len": 574.2, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 37.7, "frames": {"chat": 209}, "mem_gb": 9.94}
{"step": 143, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.06838633713191375, "tokens": 120000, "cumulative_loss_tokens": 17160000, "grad_norm": 0.263671875, "lr": 3e-05, "finish_rate": 0.908, "comp_len": 458.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 40.7, "frames": {"chat": 262}, "mem_gb": 9.87}
{"step": 144, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07117203206044312, "tokens": 120000, "cumulative_loss_tokens": 17280000, "grad_norm": 0.271484375, "lr": 3e-05, "finish_rate": 0.9, "comp_len": 481.9, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 39.9, "frames": {"chat": 249}, "mem_gb": 9.96}
{"step": 145, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.09463274955069646, "tokens": 120000, "cumulative_loss_tokens": 17400000, "grad_norm": 0.88671875, "lr": 3e-05, "finish_rate": 0.819, "comp_len": 528.6, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 39.3, "frames": {"chat": 227}, "mem_gb": 10.0}
{"step": 146, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07090324180225531, "tokens": 120000, "cumulative_loss_tokens": 17520000, "grad_norm": 0.27734375, "lr": 3e-05, "finish_rate": 0.814, "comp_len": 543.0, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 38.3, "frames": {"chat": 221}, "mem_gb": 9.99}
{"step": 147, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07361622673333623, "tokens": 120000, "cumulative_loss_tokens": 17640000, "grad_norm": 0.287109375, "lr": 3e-05, "finish_rate": 0.859, "comp_len": 512.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.0, "frames": {"chat": 234}, "mem_gb": 10.01}
{"step": 148, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.06774209337647383, "tokens": 120000, "cumulative_loss_tokens": 17760000, "grad_norm": 0.302734375, "lr": 3e-05, "finish_rate": 0.817, "comp_len": 563.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 34.8, "frames": {"chat": 213}, "mem_gb": 9.96}
{"step": 149, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.06980254276961399, "tokens": 120000, "cumulative_loss_tokens": 17880000, "grad_norm": 0.30078125, "lr": 3e-05, "finish_rate": 0.836, "comp_len": 563.4, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 35.0, "frames": {"chat": 213}, "mem_gb": 9.89}
{"step": 150, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.07166271255910396, "tokens": 120000, "cumulative_loss_tokens": 18000000, "grad_norm": 0.296875, "lr": 3e-05, "finish_rate": 0.906, "comp_len": 512.8, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 36.5, "frames": {"chat": 234}, "mem_gb": 9.92}
[eval step 150] sample: 'To solve the given system of equations involving the digits \\(a\\), \\(b\\), \\(k\\), \\(m\\), and \\(p\\), we need to ensure that each digit is a non-zero digit (i.e., \\(1 \\leq a, b, k, m, p \\leq'
checkpoint snapshot queued -> outputs/healed/grid_math/glean_keep25_s1224/step0150
wandb: updating run metadata
wandb: uploading output.log; uploading wandb-summary.json; uploading config.yaml
wandb:
wandb: Run history:
wandb: comp_len β–‡β–ƒβ–ƒβ–†β–„β–ƒβ–β–ƒβ–‡β–ƒβ–†β–…β–ƒβ–ƒβ–…β–ˆβ–‚β–ƒβ–β–…β–†β–„β–ƒβ–‡β–‚β–ƒβ–†β–…β–ƒβ–…β–β–‚β–‡β–†β–†β–‚β–…β–…β–†β–ƒ
wandb: cumulative_loss_tokens β–β–β–β–β–β–‚β–‚β–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–„β–„β–„β–„β–„β–„β–„β–…β–…β–…β–…β–…β–…β–†β–†β–†β–†β–†β–†β–‡β–‡β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
wandb: epoch β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–…β–…β–…β–…β–…β–…β–…β–…β–…β–…β–…β–…β–…β–…β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
wandb: finish_rate β–„β–β–ƒβ–†β–‡β–…β–ˆβ–„β–…β–…β–‡β–‚β–†β–ƒβ–…β–‡β–…β–„β–ƒβ–‚β–†β–ˆβ–ƒβ–…β–β–…β–…β–ˆβ–„β–…β–†β–‚β–†β–‚β–‚β–‚β–ƒβ–…β–‡β–…
wandb: forward_topk_kl β–ˆβ–†β–…β–‚β–‚β–‚β–‚β–‚β–‚β–‚β–‚β–‚β–β–‚β–‚β–‚β–β–‚β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–
wandb: grad_norm β–ˆβ–‚β–‚β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–
wandb: lr β–β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
wandb: mem_gb β–†β–…β–…β–…β–ƒβ–…β–‡β–‚β–…β–†β–…β–…β–β–ˆβ–…β–…β–ƒβ–†β–…β–…β–…β–ˆβ–‚β–…β–…β–‚β–ƒβ–‚β–…β–‚β–β–β–ˆβ–„β–…β–…β–…β–β–„β–ƒ
wandb: step β–β–β–β–β–‚β–‚β–‚β–‚β–‚β–‚β–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–„β–„β–…β–…β–…β–…β–…β–†β–†β–†β–†β–†β–‡β–‡β–‡β–‡β–‡β–‡β–ˆβ–ˆβ–ˆβ–ˆ
wandb: t_data_s β–β–β–β–β–β–β–β–β–β–β–β–β–ˆβ–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–
wandb: +3 ...
wandb:
wandb: Run summary:
wandb: comp_len 512.8
wandb: cumulative_loss_tokens 18000000
wandb: epoch 2
wandb: finish_rate 0.906
wandb: forward_topk_kl 0.07166
wandb: grad_norm 0.29688
wandb: lr 3e-05
wandb: mem_gb 9.92
wandb: step 150
wandb: t_data_s 0
wandb: +4 ...
wandb:
wandb: πŸš€ View run glean-math-keep25-s1224 at: https://wandb.ai/hbfreed/glean-grid/runs/97sk7igo
wandb: ⭐️ View project at: https://wandb.ai/hbfreed/glean-grid
wandb: Synced 5 W&B file(s), 0 media file(s), 0 artifact file(s) and 0 other file(s)
wandb: Find logs at: outputs/healed/grid_math/glean_keep25_s1224/wandb/run-20260716_040623-97sk7igo/logs
{
"correct": 545,
"accuracy": 0.4131918119787718,
"finished": 1283,
"finish_rate": 0.9727065959059894,
"mean_completion_tokens": 202.013646702047
}
saved item-level results -> outputs/evals/grid_math/glean_keep25_s1224_step100_chat.json
{
"correct": 569,
"accuracy": 0.4313874147081122,
"finished": 1277,
"finish_rate": 0.9681576952236542,
"mean_completion_tokens": 176.21683093252463
}
saved item-level results -> outputs/evals/grid_math/glean_keep25_s1224_step150_chat.json