variable-reap-archive / healed /warmup_keep50.log
hbfreed's picture
Add files using upload-large-folder tool
a2966b3 verified
Raw
History Blame Contribute Delete
72 kB
wandb: [wandb.login()] Loaded credentials for https://api.wandb.ai from /home/henry/.netrc.
wandb: Currently logged in as: hbfreed to https://api.wandb.ai. Use `wandb login --relogin` to force relogin
wandb: Tracking run with wandb version 0.28.0
wandb: Run data is saved locally in outputs/healed/keep50_offpolicy_warmup_s1224/wandb/run-20260801_230924-9td6b5cn
wandb: Run `wandb offline` to turn off syncing.
wandb: Syncing run offpolicy-warmup-keep50-s1224
wandb: ⭐️ View project at https://wandb.ai/hbfreed/glean-heal
wandb: πŸš€ View run at https://wandb.ai/hbfreed/glean-heal/runs/9td6b5cn
Loading checkpoint shards: 0%| | 0/2 [00:00<?, ?it/s] Loading checkpoint shards: 50%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 1/2 [00:00<00:00, 5.87it/s] Loading checkpoint shards: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 2/2 [00:00<00:00, 8.05it/s]
12115 cached top-128 chat trajectories / 6,476,634 unique tokens | 53 steps/epoch | 150 total steps | student params 3.70B | teacher overlap=False
{"step": 1, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.21866693885562322, "tokens": 120000, "cumulative_loss_tokens": 120000, "grad_norm": 0.703125, "lr": 6e-06, "finish_rate": 0.907, "comp_len": 508.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.337, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.4, "frames": {"chat": 236}, "mem_gb": 9.77, "mem_gb_teacher": 9.77}
The attention mask is not set and cannot be inferred from input because pad token is same as eos token. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.
[eval step 1] sample: 'To solve the given system of equations:\n\\[\n\\begin{align*}\na + b &= k, \\\\\nk + m &= p, \\\\\np + a &= r, \\\\\nb + m + r &= 18,\n\\end{align*}\n\\]\nwe need to determine the values'
{"step": 2, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.27999618121907116, "tokens": 120000, "cumulative_loss_tokens": 240000, "grad_norm": 0.80078125, "lr": 9e-06, "finish_rate": 0.781, "comp_len": 558.1, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.474, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.1, "frames": {"chat": 215}, "mem_gb": 9.96, "mem_gb_teacher": 9.96}
{"step": 3, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.3080657049433639, "tokens": 120000, "cumulative_loss_tokens": 360000, "grad_norm": 0.89453125, "lr": 1.2e-05, "finish_rate": 0.825, "comp_len": 553.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.376, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.3, "frames": {"chat": 217}, "mem_gb": 9.83, "mem_gb_teacher": 9.83}
{"step": 4, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.27968517109975216, "tokens": 120000, "cumulative_loss_tokens": 480000, "grad_norm": 0.80078125, "lr": 1.5e-05, "finish_rate": 0.8, "comp_len": 585.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.366, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.6, "frames": {"chat": 205}, "mem_gb": 9.89, "mem_gb_teacher": 9.89}
{"step": 5, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.2137425679458926, "tokens": 120000, "cumulative_loss_tokens": 600000, "grad_norm": 0.62890625, "lr": 1.8e-05, "finish_rate": 0.834, "comp_len": 524.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.318, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.8, "frames": {"chat": 229}, "mem_gb": 9.86, "mem_gb_teacher": 9.86}
{"step": 6, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.34967619865822297, "tokens": 120000, "cumulative_loss_tokens": 720000, "grad_norm": 1.3125, "lr": 2.1e-05, "finish_rate": 0.812, "comp_len": 538.1, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.336, "t_data_s": 0.1, "t_rollout_s": 0.0, "t_step_s": 25.4, "frames": {"chat": 223}, "mem_gb": 9.93, "mem_gb_teacher": 9.93}
{"step": 7, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.23986013823635877, "tokens": 120000, "cumulative_loss_tokens": 840000, "grad_norm": 0.5546875, "lr": 2.4e-05, "finish_rate": 0.708, "comp_len": 594.1, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.42, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.6, "frames": {"chat": 202}, "mem_gb": 9.97, "mem_gb_teacher": 9.97}
{"step": 8, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.3001840955584, "tokens": 120000, "cumulative_loss_tokens": 960000, "grad_norm": 0.78515625, "lr": 2.7000000000000002e-05, "finish_rate": 0.77, "comp_len": 574.2, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.453, "t_data_s": 0.1, "t_rollout_s": 0.0, "t_step_s": 24.9, "frames": {"chat": 209}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 9, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.21187542014177888, "tokens": 120000, "cumulative_loss_tokens": 1080000, "grad_norm": 0.57421875, "lr": 3e-05, "finish_rate": 0.885, "comp_len": 528.6, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.397, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.7, "frames": {"chat": 227}, "mem_gb": 9.92, "mem_gb_teacher": 9.92}
{"step": 10, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.21670206268125525, "tokens": 120000, "cumulative_loss_tokens": 1200000, "grad_norm": 0.6328125, "lr": 3e-05, "finish_rate": 0.848, "comp_len": 521.7, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.546, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.0, "frames": {"chat": 230}, "mem_gb": 9.99, "mem_gb_teacher": 9.99}
[eval step 10] sample: "To solve this problem, we need to find the values of \\(a\\), \\(b\\), \\(k\\), \\(m\\), and \\(p\\) that satisfy the given equations. Let's break down the problem step-by-step:\n\n1. **Understand the Equations:*"
{"step": 11, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.21374244357372324, "tokens": 120000, "cumulative_loss_tokens": 1320000, "grad_norm": 0.52734375, "lr": 3e-05, "finish_rate": 0.879, "comp_len": 519.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.426, "t_data_s": 0.1, "t_rollout_s": 0.0, "t_step_s": 26.0, "frames": {"chat": 231}, "mem_gb": 9.84, "mem_gb_teacher": 9.84}
{"step": 12, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.23642759951651096, "tokens": 120000, "cumulative_loss_tokens": 1440000, "grad_norm": 0.625, "lr": 3e-05, "finish_rate": 0.882, "comp_len": 489.8, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.246, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.3, "frames": {"chat": 245}, "mem_gb": 9.92, "mem_gb_teacher": 9.92}
{"step": 13, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.2684279385884603, "tokens": 120000, "cumulative_loss_tokens": 1560000, "grad_norm": 0.73828125, "lr": 3e-05, "finish_rate": 0.81, "comp_len": 571.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.375, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.7, "frames": {"chat": 210}, "mem_gb": 9.92, "mem_gb_teacher": 9.92}
{"step": 14, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.2720188724226008, "tokens": 120000, "cumulative_loss_tokens": 1680000, "grad_norm": 0.69140625, "lr": 3e-05, "finish_rate": 0.758, "comp_len": 568.7, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.312, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.3, "frames": {"chat": 211}, "mem_gb": 9.92, "mem_gb_teacher": 9.92}
{"step": 15, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.2650659183566769, "tokens": 120000, "cumulative_loss_tokens": 1800000, "grad_norm": 0.56640625, "lr": 3e-05, "finish_rate": 0.819, "comp_len": 543.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.413, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.1, "frames": {"chat": 221}, "mem_gb": 9.98, "mem_gb_teacher": 9.98}
{"step": 16, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.22639379921114694, "tokens": 120000, "cumulative_loss_tokens": 1920000, "grad_norm": 0.609375, "lr": 3e-05, "finish_rate": 0.912, "comp_len": 480.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.366, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.1, "frames": {"chat": 250}, "mem_gb": 9.79, "mem_gb_teacher": 9.79}
{"step": 17, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.2573891951084137, "tokens": 120000, "cumulative_loss_tokens": 2040000, "grad_norm": 0.796875, "lr": 3e-05, "finish_rate": 0.79, "comp_len": 524.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.402, "t_data_s": 0.1, "t_rollout_s": 0.0, "t_step_s": 26.9, "frames": {"chat": 229}, "mem_gb": 9.96, "mem_gb_teacher": 9.96}
{"step": 18, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.22218937695200244, "tokens": 120000, "cumulative_loss_tokens": 2160000, "grad_norm": 0.77734375, "lr": 3e-05, "finish_rate": 0.888, "comp_len": 480.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.479, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.7, "frames": {"chat": 250}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 19, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.33289732446968556, "tokens": 120000, "cumulative_loss_tokens": 2280000, "grad_norm": 2.546875, "lr": 3e-05, "finish_rate": 0.844, "comp_len": 519.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.363, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.6, "frames": {"chat": 231}, "mem_gb": 9.81, "mem_gb_teacher": 9.81}
{"step": 20, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.327336692000553, "tokens": 120000, "cumulative_loss_tokens": 2400000, "grad_norm": 1.3984375, "lr": 3e-05, "finish_rate": 0.844, "comp_len": 535.7, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.433, "t_data_s": 0.1, "t_rollout_s": 0.0, "t_step_s": 25.4, "frames": {"chat": 224}, "mem_gb": 9.86, "mem_gb_teacher": 9.86}
[eval step 20] sample: '```python\ndef f(a, b, c, d, e, f, g, h):\n ax + by + cz + ey + fx + gy + hz = 0\n ```\n\nWe need to find the values of \\(a\\), \\(b\\), \\(c'
{"step": 21, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.33612352709385257, "tokens": 120000, "cumulative_loss_tokens": 2520000, "grad_norm": 1.640625, "lr": 3e-05, "finish_rate": 0.802, "comp_len": 566.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.329, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.9, "frames": {"chat": 212}, "mem_gb": 9.9, "mem_gb_teacher": 9.9}
{"step": 22, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.2883151387684047, "tokens": 120000, "cumulative_loss_tokens": 2640000, "grad_norm": 1.3984375, "lr": 3e-05, "finish_rate": 0.87, "comp_len": 504.2, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.376, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.9, "frames": {"chat": 238}, "mem_gb": 9.85, "mem_gb_teacher": 9.85}
{"step": 23, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.25279294178610046, "tokens": 120000, "cumulative_loss_tokens": 2760000, "grad_norm": 1.3828125, "lr": 3e-05, "finish_rate": 0.903, "comp_len": 466.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.445, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.9, "frames": {"chat": 257}, "mem_gb": 9.73, "mem_gb_teacher": 9.73}
{"step": 24, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.2615377183983723, "tokens": 120000, "cumulative_loss_tokens": 2880000, "grad_norm": 1.5625, "lr": 3e-05, "finish_rate": 0.868, "comp_len": 528.6, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.405, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.7, "frames": {"chat": 227}, "mem_gb": 9.92, "mem_gb_teacher": 9.92}
{"step": 25, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.32016109869256615, "tokens": 120000, "cumulative_loss_tokens": 3000000, "grad_norm": 2.546875, "lr": 3e-05, "finish_rate": 0.838, "comp_len": 526.3, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.359, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.8, "frames": {"chat": 228}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 26, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.34886745701755084, "tokens": 120000, "cumulative_loss_tokens": 3120000, "grad_norm": 2.0625, "lr": 3e-05, "finish_rate": 0.803, "comp_len": 515.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.542, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.1, "frames": {"chat": 233}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 27, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.3173991178593288, "tokens": 120000, "cumulative_loss_tokens": 3240000, "grad_norm": 2.671875, "lr": 3e-05, "finish_rate": 0.863, "comp_len": 515.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.351, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.3, "frames": {"chat": 233}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 28, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.3661775833528489, "tokens": 120000, "cumulative_loss_tokens": 3360000, "grad_norm": 2.984375, "lr": 3e-05, "finish_rate": 0.731, "comp_len": 609.1, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.388, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.3, "frames": {"chat": 197}, "mem_gb": 10.03, "mem_gb_teacher": 10.03}
{"step": 29, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.36429248579877116, "tokens": 120000, "cumulative_loss_tokens": 3480000, "grad_norm": 5.625, "lr": 3e-05, "finish_rate": 0.862, "comp_len": 502.1, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.384, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.0, "frames": {"chat": 239}, "mem_gb": 9.79, "mem_gb_teacher": 9.79}
{"step": 30, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.3706551998923222, "tokens": 120000, "cumulative_loss_tokens": 3600000, "grad_norm": 4.34375, "lr": 3e-05, "finish_rate": 0.83, "comp_len": 535.7, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.302, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.1, "frames": {"chat": 224}, "mem_gb": 9.83, "mem_gb_teacher": 9.83}
[eval step 30] sample: 'To solve the problem, we need to determine the value of \\( p \\) given the equations:\n\n1. \\( a + b = k \\)\n2. \\( k + m = p \\)\n3. \\( p + a = r \\)\n4. \\( b + m + r ='
{"step": 31, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.35791126018886765, "tokens": 120000, "cumulative_loss_tokens": 3720000, "grad_norm": 2.03125, "lr": 3e-05, "finish_rate": 0.788, "comp_len": 553.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.388, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.3, "frames": {"chat": 217}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 32, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.3755841711225609, "tokens": 120000, "cumulative_loss_tokens": 3840000, "grad_norm": 2.3125, "lr": 3e-05, "finish_rate": 0.876, "comp_len": 497.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.259, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.2, "frames": {"chat": 241}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 33, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.3561015821622064, "tokens": 120000, "cumulative_loss_tokens": 3960000, "grad_norm": 2.359375, "lr": 3e-05, "finish_rate": 0.835, "comp_len": 550.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.407, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.9, "frames": {"chat": 218}, "mem_gb": 9.92, "mem_gb_teacher": 9.92}
{"step": 34, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.38959850126380724, "tokens": 120000, "cumulative_loss_tokens": 4080000, "grad_norm": 2.125, "lr": 3e-05, "finish_rate": 0.767, "comp_len": 582.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.439, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.9, "frames": {"chat": 206}, "mem_gb": 9.93, "mem_gb_teacher": 9.93}
{"step": 35, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.3970217802577963, "tokens": 120000, "cumulative_loss_tokens": 4200000, "grad_norm": 2.734375, "lr": 3e-05, "finish_rate": 0.845, "comp_len": 517.2, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.512, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.0, "frames": {"chat": 232}, "mem_gb": 9.97, "mem_gb_teacher": 9.97}
{"step": 36, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.4204790102675557, "tokens": 120000, "cumulative_loss_tokens": 4320000, "grad_norm": 1.796875, "lr": 3e-05, "finish_rate": 0.771, "comp_len": 550.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.432, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.6, "frames": {"chat": 218}, "mem_gb": 9.99, "mem_gb_teacher": 9.99}
{"step": 37, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.40933242611338694, "tokens": 120000, "cumulative_loss_tokens": 4440000, "grad_norm": 1.6015625, "lr": 3e-05, "finish_rate": 0.779, "comp_len": 563.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.428, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.9, "frames": {"chat": 213}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 38, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.3901712652951479, "tokens": 120000, "cumulative_loss_tokens": 4560000, "grad_norm": 2.640625, "lr": 3e-05, "finish_rate": 0.887, "comp_len": 483.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.527, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.5, "frames": {"chat": 248}, "mem_gb": 9.92, "mem_gb_teacher": 9.92}
{"step": 39, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.3685633979354054, "tokens": 120000, "cumulative_loss_tokens": 4680000, "grad_norm": 1.921875, "lr": 3e-05, "finish_rate": 0.803, "comp_len": 550.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.333, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.2, "frames": {"chat": 218}, "mem_gb": 9.99, "mem_gb_teacher": 9.99}
{"step": 40, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.39544522463083265, "tokens": 120000, "cumulative_loss_tokens": 4800000, "grad_norm": 2.3125, "lr": 3e-05, "finish_rate": 0.851, "comp_len": 543.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.401, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.7, "frames": {"chat": 221}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
[eval step 40] sample: "To solve the problem, we need to determine the value of \\(p\\) given the constraints. Let's break down the problem step-by-step:\n\n1. **Define Variables:**\n - Let \\(d\\) be the digit \\(0, 1, 2, \\ldots,"
{"step": 41, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.38834569306795796, "tokens": 120000, "cumulative_loss_tokens": 4920000, "grad_norm": 1.640625, "lr": 3e-05, "finish_rate": 0.894, "comp_len": 508.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.396, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.5, "frames": {"chat": 236}, "mem_gb": 9.87, "mem_gb_teacher": 9.87}
{"step": 42, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.40595541520963113, "tokens": 120000, "cumulative_loss_tokens": 5040000, "grad_norm": 2.390625, "lr": 3e-05, "finish_rate": 0.858, "comp_len": 487.8, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.338, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.6, "frames": {"chat": 246}, "mem_gb": 9.79, "mem_gb_teacher": 9.79}
{"step": 43, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.4013322958761205, "tokens": 120000, "cumulative_loss_tokens": 5160000, "grad_norm": 2.703125, "lr": 3e-05, "finish_rate": 0.838, "comp_len": 512.8, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.48, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.9, "frames": {"chat": 234}, "mem_gb": 10.05, "mem_gb_teacher": 10.05}
{"step": 44, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.41921593125065165, "tokens": 120000, "cumulative_loss_tokens": 5280000, "grad_norm": 2.984375, "lr": 3e-05, "finish_rate": 0.748, "comp_len": 594.1, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.474, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.7, "frames": {"chat": 202}, "mem_gb": 9.93, "mem_gb_teacher": 9.93}
{"step": 45, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.45117179917966327, "tokens": 120000, "cumulative_loss_tokens": 5400000, "grad_norm": 1.953125, "lr": 3e-05, "finish_rate": 0.811, "comp_len": 553.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.36, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.2, "frames": {"chat": 217}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 46, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.45246317840516564, "tokens": 120000, "cumulative_loss_tokens": 5520000, "grad_norm": 4.0625, "lr": 3e-05, "finish_rate": 0.866, "comp_len": 535.7, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.47, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.4, "frames": {"chat": 224}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 47, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.49074055876086153, "tokens": 120000, "cumulative_loss_tokens": 5640000, "grad_norm": 2.265625, "lr": 3e-05, "finish_rate": 0.753, "comp_len": 558.1, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.285, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.2, "frames": {"chat": 215}, "mem_gb": 9.96, "mem_gb_teacher": 9.96}
{"step": 48, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.502812573158741, "tokens": 120000, "cumulative_loss_tokens": 5760000, "grad_norm": 4.78125, "lr": 3e-05, "finish_rate": 0.884, "comp_len": 463.3, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.356, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.6, "frames": {"chat": 259}, "mem_gb": 9.87, "mem_gb_teacher": 9.87}
{"step": 49, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.4940076053115229, "tokens": 120000, "cumulative_loss_tokens": 5880000, "grad_norm": 2.875, "lr": 3e-05, "finish_rate": 0.829, "comp_len": 571.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.475, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.5, "frames": {"chat": 210}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 50, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.47275001460264127, "tokens": 120000, "cumulative_loss_tokens": 6000000, "grad_norm": 3.296875, "lr": 3e-05, "finish_rate": 0.77, "comp_len": 563.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.304, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.3, "frames": {"chat": 213}, "mem_gb": 9.99, "mem_gb_teacher": 9.99}
[eval step 50] sample: 'To solve the problem, we need to determine the values of \\(p\\), \\(k\\), \\(r\\), and \\(m\\).\n\nGiven:\n1. \\(a + b = k\\)\n2. \\(k + m = p\\)\n3. \\(p + a = r\\)\n'
checkpoint snapshot queued -> outputs/healed/keep50_offpolicy_warmup_s1224/step0050
{"step": 51, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.4224179199380179, "tokens": 120000, "cumulative_loss_tokens": 6120000, "grad_norm": 1.75, "lr": 3e-05, "finish_rate": 0.815, "comp_len": 540.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.453, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.1, "frames": {"chat": 222}, "mem_gb": 9.9, "mem_gb_teacher": 9.9}
{"step": 52, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.4378351664955417, "tokens": 120000, "cumulative_loss_tokens": 6240000, "grad_norm": 2.421875, "lr": 3e-05, "finish_rate": 0.889, "comp_len": 510.6, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.439, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.7, "frames": {"chat": 235}, "mem_gb": 9.96, "mem_gb_teacher": 9.96}
{"step": 53, "epoch": 0, "training_mode": "off-policy", "forward_topk_kl": 0.44709739099716145, "tokens": 120000, "cumulative_loss_tokens": 6360000, "grad_norm": 6.625, "lr": 3e-05, "finish_rate": 0.798, "comp_len": 576.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.417, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.4, "frames": {"chat": 208}, "mem_gb": 9.91, "mem_gb_teacher": 9.91}
{"step": 54, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.4153756284924845, "tokens": 120000, "cumulative_loss_tokens": 6480000, "grad_norm": 5.4375, "lr": 3e-05, "finish_rate": 0.733, "comp_len": 628.3, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.413, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 23.5, "frames": {"chat": 191}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 55, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.40591654521947107, "tokens": 120000, "cumulative_loss_tokens": 6600000, "grad_norm": 5.75, "lr": 3e-05, "finish_rate": 0.845, "comp_len": 547.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.523, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.1, "frames": {"chat": 219}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 56, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.43598461368133623, "tokens": 120000, "cumulative_loss_tokens": 6720000, "grad_norm": 7.0, "lr": 3e-05, "finish_rate": 0.778, "comp_len": 579.7, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.459, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.8, "frames": {"chat": 207}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 57, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.4643517450052003, "tokens": 120000, "cumulative_loss_tokens": 6840000, "grad_norm": 5.53125, "lr": 3e-05, "finish_rate": 0.755, "comp_len": 576.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.449, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.4, "frames": {"chat": 208}, "mem_gb": 9.91, "mem_gb_teacher": 9.91}
{"step": 58, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.41939549018144606, "tokens": 120000, "cumulative_loss_tokens": 6960000, "grad_norm": 26.125, "lr": 3e-05, "finish_rate": 0.799, "comp_len": 547.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.288, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.7, "frames": {"chat": 219}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 59, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.44048424391622343, "tokens": 120000, "cumulative_loss_tokens": 7080000, "grad_norm": 58.5, "lr": 3e-05, "finish_rate": 0.915, "comp_len": 487.8, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.329, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.8, "frames": {"chat": 246}, "mem_gb": 9.82, "mem_gb_teacher": 9.82}
{"step": 60, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.4594272028216471, "tokens": 120000, "cumulative_loss_tokens": 7200000, "grad_norm": 46.0, "lr": 3e-05, "finish_rate": 0.704, "comp_len": 582.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.437, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.3, "frames": {"chat": 206}, "mem_gb": 9.97, "mem_gb_teacher": 9.97}
[eval step 60] sample: "To solve this problem, we need to determine the values of \\(p\\), \\(k\\), and \\(r\\) based on the given equations. Here's the step-by-step approach:\n\n1. **Understand the Given Equations:**\n \\[\n \\begi"
{"step": 61, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.4421705829419196, "tokens": 120000, "cumulative_loss_tokens": 7320000, "grad_norm": 5.1875, "lr": 3e-05, "finish_rate": 0.876, "comp_len": 515.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.375, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.2, "frames": {"chat": 233}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 62, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.5114135434468587, "tokens": 120000, "cumulative_loss_tokens": 7440000, "grad_norm": 20.375, "lr": 3e-05, "finish_rate": 0.847, "comp_len": 524.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.363, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.4, "frames": {"chat": 229}, "mem_gb": 9.82, "mem_gb_teacher": 9.82}
{"step": 63, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.4925492258039614, "tokens": 120000, "cumulative_loss_tokens": 7560000, "grad_norm": 22.0, "lr": 3e-05, "finish_rate": 0.864, "comp_len": 508.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.285, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.7, "frames": {"chat": 236}, "mem_gb": 9.85, "mem_gb_teacher": 9.85}
{"step": 64, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.4923259206386904, "tokens": 120000, "cumulative_loss_tokens": 7680000, "grad_norm": 21.875, "lr": 3e-05, "finish_rate": 0.87, "comp_len": 502.1, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.37, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.0, "frames": {"chat": 239}, "mem_gb": 9.74, "mem_gb_teacher": 9.74}
{"step": 65, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.41652837800706427, "tokens": 120000, "cumulative_loss_tokens": 7800000, "grad_norm": 8.875, "lr": 3e-05, "finish_rate": 0.867, "comp_len": 497.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.349, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.8, "frames": {"chat": 241}, "mem_gb": 9.86, "mem_gb_teacher": 9.86}
{"step": 66, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3965861890381823, "tokens": 120000, "cumulative_loss_tokens": 7920000, "grad_norm": 3.484375, "lr": 3e-05, "finish_rate": 0.863, "comp_len": 531.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.44, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.1, "frames": {"chat": 226}, "mem_gb": 9.82, "mem_gb_teacher": 9.82}
{"step": 67, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.36972922986969353, "tokens": 120000, "cumulative_loss_tokens": 8040000, "grad_norm": 17.0, "lr": 3e-05, "finish_rate": 0.893, "comp_len": 512.8, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.32, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.0, "frames": {"chat": 234}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 68, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.374694859992216, "tokens": 120000, "cumulative_loss_tokens": 8160000, "grad_norm": 21.125, "lr": 3e-05, "finish_rate": 0.914, "comp_len": 466.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.401, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.2, "frames": {"chat": 257}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 69, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.4051088455612461, "tokens": 120000, "cumulative_loss_tokens": 8280000, "grad_norm": 7.3125, "lr": 3e-05, "finish_rate": 0.76, "comp_len": 576.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.523, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.1, "frames": {"chat": 208}, "mem_gb": 10.0, "mem_gb_teacher": 10.0}
{"step": 70, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.4079768711109956, "tokens": 120000, "cumulative_loss_tokens": 8400000, "grad_norm": 2.046875, "lr": 3e-05, "finish_rate": 0.763, "comp_len": 568.7, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.522, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.0, "frames": {"chat": 211}, "mem_gb": 9.97, "mem_gb_teacher": 9.97}
[eval step 70] sample: 'To determine the value of \\(p\\), we need to follow the steps outlined in the the problem:\n\n1. **Understand the Problem:**\n - \\(a + b = k\\)\n - \\(k + m = p\\)\n - \\(p + a = r\\)\n -'
{"step": 71, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3891906266813477, "tokens": 120000, "cumulative_loss_tokens": 8520000, "grad_norm": 3.484375, "lr": 3e-05, "finish_rate": 0.806, "comp_len": 528.6, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.407, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.1, "frames": {"chat": 227}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 72, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.4033234915149709, "tokens": 120000, "cumulative_loss_tokens": 8640000, "grad_norm": 5.03125, "lr": 3e-05, "finish_rate": 0.796, "comp_len": 568.7, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.434, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.2, "frames": {"chat": 211}, "mem_gb": 9.93, "mem_gb_teacher": 9.93}
{"step": 73, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3784339854914695, "tokens": 120000, "cumulative_loss_tokens": 8760000, "grad_norm": 4.0, "lr": 3e-05, "finish_rate": 0.861, "comp_len": 504.2, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.403, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.6, "frames": {"chat": 238}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 74, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.37151155275255443, "tokens": 120000, "cumulative_loss_tokens": 8880000, "grad_norm": 1.3515625, "lr": 3e-05, "finish_rate": 0.835, "comp_len": 506.3, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.517, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.8, "frames": {"chat": 237}, "mem_gb": 9.99, "mem_gb_teacher": 9.99}
{"step": 75, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.40700901261599115, "tokens": 120000, "cumulative_loss_tokens": 9000000, "grad_norm": 7.09375, "lr": 3e-05, "finish_rate": 0.721, "comp_len": 576.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.339, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.5, "frames": {"chat": 208}, "mem_gb": 9.99, "mem_gb_teacher": 9.99}
{"step": 76, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3612343382894993, "tokens": 120000, "cumulative_loss_tokens": 9120000, "grad_norm": 6.28125, "lr": 3e-05, "finish_rate": 0.801, "comp_len": 543.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.588, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.5, "frames": {"chat": 221}, "mem_gb": 10.07, "mem_gb_teacher": 10.07}
{"step": 77, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3821119113404304, "tokens": 120000, "cumulative_loss_tokens": 9240000, "grad_norm": 6.90625, "lr": 3e-05, "finish_rate": 0.853, "comp_len": 517.2, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.479, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.9, "frames": {"chat": 232}, "mem_gb": 9.91, "mem_gb_teacher": 9.91}
{"step": 78, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3988335593829552, "tokens": 120000, "cumulative_loss_tokens": 9360000, "grad_norm": 1.796875, "lr": 3e-05, "finish_rate": 0.764, "comp_len": 576.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.43, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.7, "frames": {"chat": 208}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 79, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.36860200558168194, "tokens": 120000, "cumulative_loss_tokens": 9480000, "grad_norm": 3.09375, "lr": 3e-05, "finish_rate": 0.837, "comp_len": 528.6, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.356, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.7, "frames": {"chat": 227}, "mem_gb": 9.86, "mem_gb_teacher": 9.86}
{"step": 80, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.39240696791845064, "tokens": 120000, "cumulative_loss_tokens": 9600000, "grad_norm": 1.9453125, "lr": 3e-05, "finish_rate": 0.824, "comp_len": 543.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.387, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.5, "frames": {"chat": 221}, "mem_gb": 9.89, "mem_gb_teacher": 9.89}
[eval step 80] sample: "To solve this problem, we need to determine the values of \\(p\\), \\(r\\), and \\(m\\) based on the given equations. Here's the step-by-step approach:\n\n1. **Understand the Given Equations:**\n \\[\n \\begi"
{"step": 81, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.38983683857706686, "tokens": 120000, "cumulative_loss_tokens": 9720000, "grad_norm": 1.796875, "lr": 3e-05, "finish_rate": 0.815, "comp_len": 517.2, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.457, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.8, "frames": {"chat": 232}, "mem_gb": 9.96, "mem_gb_teacher": 9.96}
{"step": 82, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3839597153416524, "tokens": 120000, "cumulative_loss_tokens": 9840000, "grad_norm": 1.390625, "lr": 3e-05, "finish_rate": 0.822, "comp_len": 547.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.423, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.4, "frames": {"chat": 219}, "mem_gb": 9.96, "mem_gb_teacher": 9.96}
{"step": 83, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.40005449274579685, "tokens": 120000, "cumulative_loss_tokens": 9960000, "grad_norm": 1.015625, "lr": 3e-05, "finish_rate": 0.713, "comp_len": 615.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.399, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.8, "frames": {"chat": 195}, "mem_gb": 10.05, "mem_gb_teacher": 10.05}
{"step": 84, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.37729036068630717, "tokens": 120000, "cumulative_loss_tokens": 10080000, "grad_norm": 2.421875, "lr": 3e-05, "finish_rate": 0.833, "comp_len": 555.6, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.503, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.9, "frames": {"chat": 216}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 85, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.39155043416718643, "tokens": 120000, "cumulative_loss_tokens": 10200000, "grad_norm": 1.640625, "lr": 3e-05, "finish_rate": 0.788, "comp_len": 576.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.37, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.9, "frames": {"chat": 208}, "mem_gb": 9.84, "mem_gb_teacher": 9.84}
{"step": 86, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3371589448125412, "tokens": 120000, "cumulative_loss_tokens": 10320000, "grad_norm": 1.625, "lr": 3e-05, "finish_rate": 0.919, "comp_len": 510.6, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.407, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.5, "frames": {"chat": 235}, "mem_gb": 9.83, "mem_gb_teacher": 9.83}
{"step": 87, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3228179054065297, "tokens": 120000, "cumulative_loss_tokens": 10440000, "grad_norm": 1.015625, "lr": 3e-05, "finish_rate": 0.853, "comp_len": 533.3, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.391, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.9, "frames": {"chat": 225}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 88, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3611420020165543, "tokens": 120000, "cumulative_loss_tokens": 10560000, "grad_norm": 2.078125, "lr": 3e-05, "finish_rate": 0.77, "comp_len": 563.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.483, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.3, "frames": {"chat": 213}, "mem_gb": 10.03, "mem_gb_teacher": 10.03}
{"step": 89, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3108938364227613, "tokens": 120000, "cumulative_loss_tokens": 10680000, "grad_norm": 1.25, "lr": 3e-05, "finish_rate": 0.922, "comp_len": 466.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.372, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.4, "frames": {"chat": 257}, "mem_gb": 9.71, "mem_gb_teacher": 9.71}
{"step": 90, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3567947539317111, "tokens": 120000, "cumulative_loss_tokens": 10800000, "grad_norm": 2.515625, "lr": 3e-05, "finish_rate": 0.792, "comp_len": 566.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.497, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.9, "frames": {"chat": 212}, "mem_gb": 9.98, "mem_gb_teacher": 9.98}
[eval step 90] sample: 'To determine the value of \\( p \\), we need to follow the steps outlined in thepy:\n\n1. **Understand the Problem:**\n - \\( p \\) is the value of \\( k \\) when \\( k + m = 18 \\).\n - \\( k \\) is the value'
{"step": 91, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.35618535781440636, "tokens": 120000, "cumulative_loss_tokens": 10920000, "grad_norm": 2.65625, "lr": 3e-05, "finish_rate": 0.833, "comp_len": 543.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.338, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.2, "frames": {"chat": 221}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 92, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3354658261674146, "tokens": 120000, "cumulative_loss_tokens": 11040000, "grad_norm": 2.390625, "lr": 3e-05, "finish_rate": 0.868, "comp_len": 495.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.442, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.0, "frames": {"chat": 242}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 93, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3363398057249685, "tokens": 120000, "cumulative_loss_tokens": 11160000, "grad_norm": 1.53125, "lr": 3e-05, "finish_rate": 0.836, "comp_len": 545.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.352, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.6, "frames": {"chat": 220}, "mem_gb": 9.91, "mem_gb_teacher": 9.91}
{"step": 94, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.30306674740935363, "tokens": 120000, "cumulative_loss_tokens": 11280000, "grad_norm": 0.7734375, "lr": 3e-05, "finish_rate": 0.896, "comp_len": 500.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.291, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.5, "frames": {"chat": 240}, "mem_gb": 9.81, "mem_gb_teacher": 9.81}
{"step": 95, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.29799054561704397, "tokens": 120000, "cumulative_loss_tokens": 11400000, "grad_norm": 1.40625, "lr": 3e-05, "finish_rate": 0.728, "comp_len": 582.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.359, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.8, "frames": {"chat": 206}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 96, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3275977665552249, "tokens": 120000, "cumulative_loss_tokens": 11520000, "grad_norm": 1.4921875, "lr": 3e-05, "finish_rate": 0.867, "comp_len": 531.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.402, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.6, "frames": {"chat": 226}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 97, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.34486291259291274, "tokens": 120000, "cumulative_loss_tokens": 11640000, "grad_norm": 1.25, "lr": 3e-05, "finish_rate": 0.877, "comp_len": 491.8, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.382, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.8, "frames": {"chat": 244}, "mem_gb": 9.74, "mem_gb_teacher": 9.74}
{"step": 98, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.31863660610305766, "tokens": 120000, "cumulative_loss_tokens": 11760000, "grad_norm": 1.09375, "lr": 3e-05, "finish_rate": 0.804, "comp_len": 535.7, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.406, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.9, "frames": {"chat": 224}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 99, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.313539165522034, "tokens": 120000, "cumulative_loss_tokens": 11880000, "grad_norm": 0.99609375, "lr": 3e-05, "finish_rate": 0.923, "comp_len": 442.8, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.319, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.9, "frames": {"chat": 271}, "mem_gb": 9.68, "mem_gb_teacher": 9.68}
{"step": 100, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.30199915543012323, "tokens": 120000, "cumulative_loss_tokens": 12000000, "grad_norm": 0.80078125, "lr": 3e-05, "finish_rate": 0.856, "comp_len": 508.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.362, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.5, "frames": {"chat": 236}, "mem_gb": 9.96, "mem_gb_teacher": 9.96}
[eval step 100] sample: "To solve the problem, we need to determine the values of \\(p\\) and \\(k\\) given the constraints. Let's break down the problem step-by-step:\n\n1. **Understand the Constraints:**\n - Each letter represen"
checkpoint snapshot queued -> outputs/healed/keep50_offpolicy_warmup_s1224/step0100
{"step": 101, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.30484250679599745, "tokens": 120000, "cumulative_loss_tokens": 12120000, "grad_norm": 0.75, "lr": 3e-05, "finish_rate": 0.841, "comp_len": 517.2, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.306, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.5, "frames": {"chat": 232}, "mem_gb": 9.83, "mem_gb_teacher": 9.83}
{"step": 102, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.30012307230867447, "tokens": 120000, "cumulative_loss_tokens": 12240000, "grad_norm": 1.1640625, "lr": 3e-05, "finish_rate": 0.79, "comp_len": 571.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.392, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.3, "frames": {"chat": 210}, "mem_gb": 9.89, "mem_gb_teacher": 9.89}
{"step": 103, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.2981501909478257, "tokens": 120000, "cumulative_loss_tokens": 12360000, "grad_norm": 1.2421875, "lr": 3e-05, "finish_rate": 0.811, "comp_len": 553.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.431, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.0, "frames": {"chat": 217}, "mem_gb": 9.85, "mem_gb_teacher": 9.85}
{"step": 104, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.29460206581093373, "tokens": 120000, "cumulative_loss_tokens": 12480000, "grad_norm": 0.83203125, "lr": 3e-05, "finish_rate": 0.839, "comp_len": 535.7, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.38, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.1, "frames": {"chat": 224}, "mem_gb": 9.97, "mem_gb_teacher": 9.97}
{"step": 105, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.3155513667286684, "tokens": 120000, "cumulative_loss_tokens": 12600000, "grad_norm": 0.83203125, "lr": 3e-05, "finish_rate": 0.749, "comp_len": 591.1, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.481, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.3, "frames": {"chat": 203}, "mem_gb": 9.82, "mem_gb_teacher": 9.82}
{"step": 106, "epoch": 1, "training_mode": "off-policy", "forward_topk_kl": 0.2852651763110111, "tokens": 120000, "cumulative_loss_tokens": 12720000, "grad_norm": 0.75, "lr": 3e-05, "finish_rate": 0.887, "comp_len": 502.1, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.326, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.8, "frames": {"chat": 239}, "mem_gb": 9.92, "mem_gb_teacher": 9.92}
{"step": 107, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2750922793724885, "tokens": 120000, "cumulative_loss_tokens": 12840000, "grad_norm": 0.93359375, "lr": 3e-05, "finish_rate": 0.902, "comp_len": 472.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.301, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.9, "frames": {"chat": 254}, "mem_gb": 9.83, "mem_gb_teacher": 9.83}
{"step": 108, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.27038282493477067, "tokens": 120000, "cumulative_loss_tokens": 12960000, "grad_norm": 0.91796875, "lr": 3e-05, "finish_rate": 0.876, "comp_len": 497.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.421, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.4, "frames": {"chat": 241}, "mem_gb": 9.93, "mem_gb_teacher": 9.93}
{"step": 109, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.30530283329064645, "tokens": 120000, "cumulative_loss_tokens": 13080000, "grad_norm": 0.8671875, "lr": 3e-05, "finish_rate": 0.746, "comp_len": 563.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.419, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.2, "frames": {"chat": 213}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 110, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2827595801195751, "tokens": 120000, "cumulative_loss_tokens": 13200000, "grad_norm": 0.78125, "lr": 3e-05, "finish_rate": 0.864, "comp_len": 543.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.641, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.6, "frames": {"chat": 221}, "mem_gb": 10.0, "mem_gb_teacher": 10.0}
[eval step 110] sample: "To solve the problem, we need to determine the digits \\(p\\), \\(a\\), and \\(b\\) that satisfy the given conditions. Let's break down the problem step-by-step:\n\n1. **Understand the Constraints:**\n - \\(p"
{"step": 111, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.3139219396378845, "tokens": 120000, "cumulative_loss_tokens": 13320000, "grad_norm": 0.9375, "lr": 3e-05, "finish_rate": 0.745, "comp_len": 612.2, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.334, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.6, "frames": {"chat": 196}, "mem_gb": 9.96, "mem_gb_teacher": 9.96}
{"step": 112, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.29290350563563405, "tokens": 120000, "cumulative_loss_tokens": 13440000, "grad_norm": 1.2109375, "lr": 3e-05, "finish_rate": 0.926, "comp_len": 444.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.427, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 28.5, "frames": {"chat": 270}, "mem_gb": 9.77, "mem_gb_teacher": 9.77}
{"step": 113, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2911341037095835, "tokens": 120000, "cumulative_loss_tokens": 13560000, "grad_norm": 1.2578125, "lr": 3e-05, "finish_rate": 0.815, "comp_len": 555.6, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.327, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.3, "frames": {"chat": 216}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 114, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.30757556294202804, "tokens": 120000, "cumulative_loss_tokens": 13680000, "grad_norm": 0.97265625, "lr": 3e-05, "finish_rate": 0.775, "comp_len": 600.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.314, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.2, "frames": {"chat": 200}, "mem_gb": 9.91, "mem_gb_teacher": 9.91}
{"step": 115, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.27265539040267467, "tokens": 120000, "cumulative_loss_tokens": 13800000, "grad_norm": 1.046875, "lr": 3e-05, "finish_rate": 0.767, "comp_len": 582.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.406, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.9, "frames": {"chat": 206}, "mem_gb": 9.86, "mem_gb_teacher": 9.86}
{"step": 116, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.25661204309028884, "tokens": 120000, "cumulative_loss_tokens": 13920000, "grad_norm": 1.0703125, "lr": 3e-05, "finish_rate": 0.902, "comp_len": 512.8, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.5, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.4, "frames": {"chat": 234}, "mem_gb": 9.9, "mem_gb_teacher": 9.9}
{"step": 117, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.28587267751296364, "tokens": 120000, "cumulative_loss_tokens": 14040000, "grad_norm": 1.078125, "lr": 3e-05, "finish_rate": 0.823, "comp_len": 558.1, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.325, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.5, "frames": {"chat": 215}, "mem_gb": 9.91, "mem_gb_teacher": 9.91}
{"step": 118, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.24581487802788615, "tokens": 120000, "cumulative_loss_tokens": 14160000, "grad_norm": 0.8515625, "lr": 3e-05, "finish_rate": 0.922, "comp_len": 470.6, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.447, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.0, "frames": {"chat": 255}, "mem_gb": 9.89, "mem_gb_teacher": 9.89}
{"step": 119, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2630173706655701, "tokens": 120000, "cumulative_loss_tokens": 14280000, "grad_norm": 1.0390625, "lr": 3e-05, "finish_rate": 0.892, "comp_len": 480.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.377, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.9, "frames": {"chat": 250}, "mem_gb": 9.77, "mem_gb_teacher": 9.77}
{"step": 120, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.26195900368392466, "tokens": 120000, "cumulative_loss_tokens": 14400000, "grad_norm": 0.85546875, "lr": 3e-05, "finish_rate": 0.884, "comp_len": 495.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.525, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.8, "frames": {"chat": 242}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
[eval step 120] sample: "To solve this problem, we need to determine the values of \\(p\\) and \\(k\\) based on the given equations. Let's break down the problem step-by-step:\n\n1. **Understand the Equations:**\n - \\(a + b = k\\)\n"
{"step": 121, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.27887726591676476, "tokens": 120000, "cumulative_loss_tokens": 14520000, "grad_norm": 0.98828125, "lr": 3e-05, "finish_rate": 0.729, "comp_len": 603.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.517, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.5, "frames": {"chat": 199}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 122, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.31573768441453576, "tokens": 120000, "cumulative_loss_tokens": 14640000, "grad_norm": 1.15625, "lr": 3e-05, "finish_rate": 0.784, "comp_len": 576.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.386, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.7, "frames": {"chat": 208}, "mem_gb": 9.99, "mem_gb_teacher": 9.99}
{"step": 123, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2738492341738194, "tokens": 120000, "cumulative_loss_tokens": 14760000, "grad_norm": 0.859375, "lr": 3e-05, "finish_rate": 0.764, "comp_len": 576.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.304, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.6, "frames": {"chat": 208}, "mem_gb": 9.92, "mem_gb_teacher": 9.92}
{"step": 124, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.3032253828023871, "tokens": 120000, "cumulative_loss_tokens": 14880000, "grad_norm": 1.21875, "lr": 3e-05, "finish_rate": 0.732, "comp_len": 574.2, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.473, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.8, "frames": {"chat": 209}, "mem_gb": 10.07, "mem_gb_teacher": 10.07}
{"step": 125, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.27386389810865125, "tokens": 120000, "cumulative_loss_tokens": 15000000, "grad_norm": 1.5859375, "lr": 3e-05, "finish_rate": 0.855, "comp_len": 510.6, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.321, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.6, "frames": {"chat": 235}, "mem_gb": 9.91, "mem_gb_teacher": 9.91}
{"step": 126, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2811740072357158, "tokens": 120000, "cumulative_loss_tokens": 15120000, "grad_norm": 1.4453125, "lr": 3e-05, "finish_rate": 0.74, "comp_len": 588.2, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.503, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.8, "frames": {"chat": 204}, "mem_gb": 9.9, "mem_gb_teacher": 9.9}
{"step": 127, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.32506899852765103, "tokens": 120000, "cumulative_loss_tokens": 15240000, "grad_norm": 2.453125, "lr": 3e-05, "finish_rate": 0.745, "comp_len": 576.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.498, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.8, "frames": {"chat": 208}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 128, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.27416869887411593, "tokens": 120000, "cumulative_loss_tokens": 15360000, "grad_norm": 1.375, "lr": 3e-05, "finish_rate": 0.825, "comp_len": 500.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.456, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.2, "frames": {"chat": 240}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 129, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.27367745394359033, "tokens": 120000, "cumulative_loss_tokens": 15480000, "grad_norm": 1.390625, "lr": 3e-05, "finish_rate": 0.89, "comp_len": 487.8, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.425, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.4, "frames": {"chat": 246}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 130, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.27279284177869556, "tokens": 120000, "cumulative_loss_tokens": 15600000, "grad_norm": 1.1171875, "lr": 3e-05, "finish_rate": 0.909, "comp_len": 493.8, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.299, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.1, "frames": {"chat": 243}, "mem_gb": 9.76, "mem_gb_teacher": 9.76}
[eval step 130] sample: "To solve the problem, we need to determine the values of \\(p\\), \\(r\\), and \\(k\\) given the constraints. Let's break down the problem step-by-step:\n\n1. **Understand the Constraints:**\n - \\(a + b = k\\"
{"step": 131, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2835182323958725, "tokens": 120000, "cumulative_loss_tokens": 15720000, "grad_norm": 0.953125, "lr": 3e-05, "finish_rate": 0.745, "comp_len": 576.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.449, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.8, "frames": {"chat": 208}, "mem_gb": 9.96, "mem_gb_teacher": 9.96}
{"step": 132, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2755484388658156, "tokens": 120000, "cumulative_loss_tokens": 15840000, "grad_norm": 0.8203125, "lr": 3e-05, "finish_rate": 0.817, "comp_len": 547.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.352, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.3, "frames": {"chat": 219}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 133, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.27247595444793504, "tokens": 120000, "cumulative_loss_tokens": 15960000, "grad_norm": 0.796875, "lr": 3e-05, "finish_rate": 0.782, "comp_len": 568.7, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.447, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.3, "frames": {"chat": 211}, "mem_gb": 9.96, "mem_gb_teacher": 9.96}
{"step": 134, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.252491025553147, "tokens": 120000, "cumulative_loss_tokens": 16080000, "grad_norm": 0.67578125, "lr": 3e-05, "finish_rate": 0.862, "comp_len": 517.2, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.283, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.0, "frames": {"chat": 232}, "mem_gb": 9.92, "mem_gb_teacher": 9.92}
{"step": 135, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2650780773670723, "tokens": 120000, "cumulative_loss_tokens": 16200000, "grad_norm": 0.71875, "lr": 3e-05, "finish_rate": 0.804, "comp_len": 560.7, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.403, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.3, "frames": {"chat": 214}, "mem_gb": 9.96, "mem_gb_teacher": 9.96}
{"step": 136, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2518176361516118, "tokens": 120000, "cumulative_loss_tokens": 16320000, "grad_norm": 0.890625, "lr": 3e-05, "finish_rate": 0.819, "comp_len": 531.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.395, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.8, "frames": {"chat": 226}, "mem_gb": 9.85, "mem_gb_teacher": 9.85}
{"step": 137, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.23503943474429348, "tokens": 120000, "cumulative_loss_tokens": 16440000, "grad_norm": 0.75, "lr": 3e-05, "finish_rate": 0.819, "comp_len": 571.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.404, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.4, "frames": {"chat": 210}, "mem_gb": 9.96, "mem_gb_teacher": 9.96}
{"step": 138, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2334536761138588, "tokens": 120000, "cumulative_loss_tokens": 16560000, "grad_norm": 0.6796875, "lr": 3e-05, "finish_rate": 0.858, "comp_len": 550.5, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.436, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.4, "frames": {"chat": 218}, "mem_gb": 9.78, "mem_gb_teacher": 9.78}
{"step": 139, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2328599476976941, "tokens": 120000, "cumulative_loss_tokens": 16680000, "grad_norm": 0.67578125, "lr": 3e-05, "finish_rate": 0.858, "comp_len": 515.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.354, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.5, "frames": {"chat": 233}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 140, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.26084051485359666, "tokens": 120000, "cumulative_loss_tokens": 16800000, "grad_norm": 0.70703125, "lr": 3e-05, "finish_rate": 0.786, "comp_len": 558.1, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.221, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.0, "frames": {"chat": 215}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
[eval step 140] sample: "To solve this problem, we need to determine the values of \\(p\\), \\(a\\), \\(b\\), \\(m\\), and \\(r\\) that satisfy the given equations. Let's break down the problem step-by-step:\n\n1. **Understand the Equati"
{"step": 141, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.26040056609710055, "tokens": 120000, "cumulative_loss_tokens": 16920000, "grad_norm": 0.83203125, "lr": 3e-05, "finish_rate": 0.845, "comp_len": 515.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.428, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.3, "frames": {"chat": 233}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 142, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2547849820467333, "tokens": 120000, "cumulative_loss_tokens": 17040000, "grad_norm": 0.91015625, "lr": 3e-05, "finish_rate": 0.766, "comp_len": 574.2, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.353, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.2, "frames": {"chat": 209}, "mem_gb": 9.89, "mem_gb_teacher": 9.89}
{"step": 143, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2485178017048786, "tokens": 120000, "cumulative_loss_tokens": 17160000, "grad_norm": 0.83203125, "lr": 3e-05, "finish_rate": 0.908, "comp_len": 458.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.308, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 27.5, "frames": {"chat": 262}, "mem_gb": 9.82, "mem_gb_teacher": 9.82}
{"step": 144, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.25363729545498886, "tokens": 120000, "cumulative_loss_tokens": 17280000, "grad_norm": 0.7578125, "lr": 3e-05, "finish_rate": 0.9, "comp_len": 481.9, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.241, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.7, "frames": {"chat": 249}, "mem_gb": 9.91, "mem_gb_teacher": 9.91}
{"step": 145, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.29562477170241375, "tokens": 120000, "cumulative_loss_tokens": 17400000, "grad_norm": 0.9453125, "lr": 3e-05, "finish_rate": 0.819, "comp_len": 528.6, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.364, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 26.4, "frames": {"chat": 227}, "mem_gb": 9.95, "mem_gb_teacher": 9.95}
{"step": 146, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.25183481702382365, "tokens": 120000, "cumulative_loss_tokens": 17520000, "grad_norm": 0.98828125, "lr": 3e-05, "finish_rate": 0.814, "comp_len": 543.0, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.356, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.5, "frames": {"chat": 221}, "mem_gb": 9.94, "mem_gb_teacher": 9.94}
{"step": 147, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.2600625433813781, "tokens": 120000, "cumulative_loss_tokens": 17640000, "grad_norm": 0.89453125, "lr": 3e-05, "finish_rate": 0.859, "comp_len": 512.8, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.473, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.4, "frames": {"chat": 234}, "mem_gb": 9.96, "mem_gb_teacher": 9.96}
{"step": 148, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.24899872875362636, "tokens": 120000, "cumulative_loss_tokens": 17760000, "grad_norm": 0.7265625, "lr": 3e-05, "finish_rate": 0.817, "comp_len": 563.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.29, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.5, "frames": {"chat": 213}, "mem_gb": 9.9, "mem_gb_teacher": 9.9}
{"step": 149, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.22898227033279836, "tokens": 120000, "cumulative_loss_tokens": 17880000, "grad_norm": 0.87890625, "lr": 3e-05, "finish_rate": 0.836, "comp_len": 563.4, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.455, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 24.7, "frames": {"chat": 213}, "mem_gb": 9.84, "mem_gb_teacher": 9.84}
{"step": 150, "epoch": 2, "training_mode": "off-policy", "forward_topk_kl": 0.23123409348068139, "tokens": 120000, "cumulative_loss_tokens": 18000000, "grad_norm": 0.82421875, "lr": 3e-05, "finish_rate": 0.906, "comp_len": 512.8, "dropped_truncated": 0, "gold_loss": null, "gold_lambda": null, "rep_ratio": 2.392, "t_data_s": 0.0, "t_rollout_s": 0.0, "t_step_s": 25.8, "frames": {"chat": 234}, "mem_gb": 9.87, "mem_gb_teacher": 9.87}
[eval step 150] sample: "To solve this problem, we need to determine the values of \\(p\\), \\(a\\), \\(b\\), \\(m\\), and \\(r\\) that satisfy the given equations. Let's break down the problem step-by-step:\n\n1. **Understand the Constr"
checkpoint snapshot queued -> outputs/healed/keep50_offpolicy_warmup_s1224/step0150
wandb: updating run metadata
wandb: uploading summary
wandb:
wandb: Run history:
wandb: comp_len β–†β–…β–…β–†β–„β–ˆβ–†β–‡β–†β–…β–‡β–„β–…β–„β–‚β–‡β–„β–†β–ƒβ–ƒβ–β–„β–„β–‡β–…β–ˆβ–‡β–„β–†β–ˆβ–„β–‡β–‡β–‡β–„β–‡β–†β–ƒβ–…β–†
wandb: cumulative_loss_tokens β–β–β–β–‚β–‚β–‚β–‚β–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–ƒβ–„β–„β–„β–„β–„β–„β–„β–…β–…β–…β–…β–†β–†β–†β–†β–‡β–‡β–‡β–‡β–‡β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
wandb: dropped_truncated ▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁
wandb: epoch β–β–β–β–β–β–β–β–β–β–β–…β–…β–…β–…β–…β–…β–…β–…β–…β–…β–…β–…β–…β–…β–…β–…β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
wandb: finish_rate β–β–†β–‡β–†β–„β–‚β–†β–‡β–†β–ƒβ–‡β–ƒβ–β–†β–‡β–ƒβ–ƒβ–…β–…β–†β–…β–…β–β–…β–ƒβ–‡β–†β–„β–†β–ƒβ–ƒβ–ˆβ–‡β–‚β–‚β–†β–ƒβ–ˆβ–…β–ˆ
wandb: forward_topk_kl β–ƒβ–ƒβ–β–‚β–β–…β–„β–†β–…β–…β–…β–†β–ˆβ–‡β–†β–‡β–‡β–ˆβ–…β–…β–…β–…β–„β–…β–ƒβ–„β–ƒβ–„β–ƒβ–ƒβ–ƒβ–ƒβ–‚β–ƒβ–‚β–‚β–‚β–ƒβ–‚β–
wandb: grad_norm β–β–β–β–β–β–β–β–β–β–‚β–β–β–β–β–β–β–‚β–‚β–‚β–‚β–ˆβ–‚β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–β–
wandb: lr β–β–ƒβ–…β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
wandb: mem_gb β–…β–†β–‡β–†β–‡β–ˆβ–†β–†β–‡β–ƒβ–†β–†β–†β–†β–„β–†β–†β–†β–†β–…β–ˆβ–„β–†β–β–…β–„β–†β–„β–ƒβ–†β–…β–†β–…β–†β–…β–†β–‡β–„β–†β–‡
wandb: mem_gb_teacher β–†β–‡β–†β–‚β–†β–†β–ƒβ–†β–†β–…β–ˆβ–…β–†β–†β–†β–β–ƒβ–†β–†β–‡β–…β–„β–†β–ˆβ–ƒβ–†β–†β–†β–ƒβ–…β–„β–†β–†β–…β–…β–†β–†β–…β–†β–†
wandb: +6 ...
wandb:
wandb: Run summary:
wandb: comp_len 512.8
wandb: cumulative_loss_tokens 18000000
wandb: dropped_truncated 0
wandb: epoch 2
wandb: finish_rate 0.906
wandb: forward_topk_kl 0.23123
wandb: grad_norm 0.82422
wandb: lr 3e-05
wandb: mem_gb 9.87
wandb: mem_gb_teacher 9.87
wandb: +7 ...
wandb:
wandb: πŸš€ View run offpolicy-warmup-keep50-s1224 at: https://wandb.ai/hbfreed/glean-heal/runs/9td6b5cn
wandb: ⭐️ View project at: https://wandb.ai/hbfreed/glean-heal
wandb: Synced 5 W&B file(s), 0 media file(s), 0 artifact file(s) and 0 other file(s)
wandb: Find logs at: outputs/healed/keep50_offpolicy_warmup_s1224/wandb/run-20260801_230924-9td6b5cn/logs