/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available. warnings.warn('Grouped GEMM not available.') wandb: [wandb.login()] Loaded credentials for https://api.wandb.ai from /home/henry/.netrc. wandb: Currently logged in as: hbfreed to https://api.wandb.ai. Use `wandb login --relogin` to force relogin wandb: setting up run vpov0urn wandb: Tracking run with wandb version 0.28.0 wandb: Run data is saved locally in outputs/healed/healing_breadth/glean_math_keep25_seed1224_long/wandb/run-20260715_063332-vpov0urn wandb: Run `wandb offline` to turn off syncing. wandb: Syncing run heal-glean-math-keep25-seed1224-long-step500 wandb: ⭐️ View project at https://wandb.ai/hbfreed/glean-heal wandb: 🚀 View run at https://wandb.ai/hbfreed/glean-heal/runs/vpov0urn Loading checkpoint shards: 0%| | 0/3 [00:00 63998 prompts Filter: 0%| | 0/63998 [00:00 63998 prompts WARNING: difficulty filter removed nothing — the slice likely carries difficulty=None (Dolci math sources do), so the teacher-competence guard is NOT in effect Map: 0%| | 0/63998 [00:00 main() File "/home/henry/Documents/PythonProjects/variable-reap/scripts/11_distill_on_policy.py", line 860, in main save_student(student, tokenizer, live) File "/home/henry/Documents/PythonProjects/variable-reap/scripts/11_distill_on_policy.py", line 974, in save_student write_student_snapshot(student, tokenizer, path, snapshot_student(student)) File "/home/henry/Documents/PythonProjects/variable-reap/scripts/11_distill_on_policy.py", line 961, in write_student_snapshot student.save_pretrained(path, state_dict=state_dict) File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/transformers/modeling_utils.py", line 4173, in save_pretrained safe_save_file(shard, os.path.join(save_directory, shard_file), metadata=metadata) File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/safetensors/torch.py", line 323, in save_file serialize_file( safetensors._safetensors_rust.SafetensorError: Error while serializing: I/O error: No space left on device (os error 28) /home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available. warnings.warn('Grouped GEMM not available.') wandb: [wandb.login()] Loaded credentials for https://api.wandb.ai from /home/henry/.netrc. wandb: Currently logged in as: hbfreed to https://api.wandb.ai. Use `wandb login --relogin` to force relogin wandb: setting up run i5fg3xzq wandb: Tracking run with wandb version 0.28.0 wandb: Run data is saved locally in outputs/healed/healing_breadth/glean_math_keep25_seed1224_long/wandb/run-20260715_071701-i5fg3xzq wandb: Run `wandb offline` to turn off syncing. wandb: Syncing run heal-glean-math-keep25-seed1224-long-step500 wandb: ⭐️ View project at https://wandb.ai/hbfreed/glean-heal wandb: 🚀 View run at https://wandb.ai/hbfreed/glean-heal/runs/i5fg3xzq Loading checkpoint shards: 0%| | 0/3 [00:00 63998 prompts difficulty <= 4 -> 63998 prompts WARNING: difficulty filter removed nothing — the slice likely carries difficulty=None (Dolci math sources do), so the teacher-competence guard is NOT in effect 63977 prompts | 940 steps/epoch | 500 total steps | student params 2.09B | teacher overlap=True restored optimizer/scheduler state from step 50; rebuilt 228 paged buffers {"step": 51, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.36993047905315957, "tokens": 120000, "cumulative_loss_tokens": 6120000, "grad_norm": 2.203125, "lr": 3e-05, "finish_rate": 0.025, "comp_len": 508.5, "t_data_s": 0.0, "t_rollout_s": 37.4, "t_step_s": 79.7, "t_refresh_s": 0.4, "mem_gb": 10.4} {"step": 52, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.3874280411116779, "tokens": 120000, "cumulative_loss_tokens": 6240000, "grad_norm": 1.765625, "lr": 3e-05, "finish_rate": 0.021, "comp_len": 508.5, "t_data_s": 0.0, "t_rollout_s": 37.0, "t_step_s": 72.9, "t_refresh_s": 0.3, "mem_gb": 10.38} {"step": 53, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.3746817006888489, "tokens": 120000, "cumulative_loss_tokens": 6360000, "grad_norm": 1.703125, "lr": 3e-05, "finish_rate": 0.046, "comp_len": 506.3, "t_data_s": 0.0, "t_rollout_s": 37.1, "t_step_s": 73.6, "t_refresh_s": 0.3, "mem_gb": 10.45} {"step": 54, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.35510130622684954, "tokens": 120000, "cumulative_loss_tokens": 6480000, "grad_norm": 2.140625, "lr": 3e-05, "finish_rate": 0.021, "comp_len": 508.5, "t_data_s": 0.0, "t_rollout_s": 37.7, "t_step_s": 74.6, "t_refresh_s": 0.3, "mem_gb": 10.49} {"step": 55, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.34786954234316947, "tokens": 120000, "cumulative_loss_tokens": 6600000, "grad_norm": 1.8359375, "lr": 3e-05, "finish_rate": 0.025, "comp_len": 508.5, "t_data_s": 0.0, "t_rollout_s": 36.9, "t_step_s": 72.5, "t_refresh_s": 0.3, "mem_gb": 10.37} {"step": 56, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.3465809899068127, "tokens": 120000, "cumulative_loss_tokens": 6720000, "grad_norm": 1.6796875, "lr": 3e-05, "finish_rate": 0.038, "comp_len": 508.5, "t_data_s": 0.0, "t_rollout_s": 36.9, "t_step_s": 72.6, "t_refresh_s": 0.3, "mem_gb": 10.43} {"step": 57, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.37013316290453074, "tokens": 120000, "cumulative_loss_tokens": 6840000, "grad_norm": 2.34375, "lr": 3e-05, "finish_rate": 0.059, "comp_len": 506.3, "t_data_s": 0.0, "t_rollout_s": 36.9, "t_step_s": 72.9, "t_refresh_s": 0.3, "mem_gb": 10.51} {"step": 58, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.3522711561133464, "tokens": 120000, "cumulative_loss_tokens": 6960000, "grad_norm": 1.6328125, "lr": 3e-05, "finish_rate": 0.104, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 35.5, "t_step_s": 72.8, "t_refresh_s": 0.3, "mem_gb": 10.47} {"step": 59, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.3397159261740744, "tokens": 120000, "cumulative_loss_tokens": 7080000, "grad_norm": 1.8203125, "lr": 3e-05, "finish_rate": 0.055, "comp_len": 506.3, "t_data_s": 0.0, "t_rollout_s": 36.8, "t_step_s": 72.5, "t_refresh_s": 0.3, "mem_gb": 10.37} {"step": 60, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.3382764852608244, "tokens": 120000, "cumulative_loss_tokens": 7200000, "grad_norm": 1.7265625, "lr": 3e-05, "finish_rate": 0.08, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 36.7, "t_step_s": 72.7, "t_refresh_s": 0.3, "mem_gb": 10.42} The attention mask is not set and cannot be inferred from input because pad token is same as eos token. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results. [eval step 60] sample: "To solve this problem, we need to consider the constraints and the total number of trees. Let's denote:\n- \\( m \\) as the number of maples.\n- \\( l \\) as the number of larks.\n\nGiven:\n1. The total number" {"step": 60, "gsm8k_n": 64, "gsm8k_quick_chat": 0.390625, "t_eval_s": 46.9} {"step": 61, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.31868271032124756, "tokens": 120000, "cumulative_loss_tokens": 7320000, "grad_norm": 1.5859375, "lr": 3e-05, "finish_rate": 0.059, "comp_len": 506.3, "t_data_s": 0.0, "t_rollout_s": 37.5, "t_step_s": 74.1, "t_refresh_s": 0.3, "mem_gb": 12.44} {"step": 62, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.3191785570286214, "tokens": 120000, "cumulative_loss_tokens": 7440000, "grad_norm": 1.5, "lr": 3e-05, "finish_rate": 0.063, "comp_len": 506.3, "t_data_s": 0.0, "t_rollout_s": 36.9, "t_step_s": 73.1, "t_refresh_s": 0.3, "mem_gb": 10.43} {"step": 63, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.3180096957164506, "tokens": 120000, "cumulative_loss_tokens": 7560000, "grad_norm": 1.5625, "lr": 3e-05, "finish_rate": 0.092, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 35.3, "t_step_s": 70.9, "t_refresh_s": 0.3, "mem_gb": 10.37} {"step": 64, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.288780878906697, "tokens": 120000, "cumulative_loss_tokens": 7680000, "grad_norm": 1.7578125, "lr": 3e-05, "finish_rate": 0.046, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 37.6, "t_step_s": 74.6, "t_refresh_s": 0.3, "mem_gb": 10.51} {"step": 65, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.29233540159290033, "tokens": 120000, "cumulative_loss_tokens": 7800000, "grad_norm": 1.4921875, "lr": 3e-05, "finish_rate": 0.084, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 36.2, "t_step_s": 71.9, "t_refresh_s": 0.3, "mem_gb": 10.5} {"step": 66, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2769135721528282, "tokens": 120000, "cumulative_loss_tokens": 7920000, "grad_norm": 1.3515625, "lr": 3e-05, "finish_rate": 0.105, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 35.6, "t_step_s": 71.3, "t_refresh_s": 0.3, "mem_gb": 10.53} {"step": 67, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.3056252569064498, "tokens": 120000, "cumulative_loss_tokens": 8040000, "grad_norm": 1.4375, "lr": 3e-05, "finish_rate": 0.181, "comp_len": 493.8, "t_data_s": 0.0, "t_rollout_s": 33.5, "t_step_s": 69.5, "t_refresh_s": 0.3, "mem_gb": 10.38} {"step": 68, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.30779550124146043, "tokens": 120000, "cumulative_loss_tokens": 8160000, "grad_norm": 1.5, "lr": 3e-05, "finish_rate": 0.129, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.9, "t_step_s": 71.2, "t_refresh_s": 0.3, "mem_gb": 10.42} {"step": 69, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2875990556370467, "tokens": 120000, "cumulative_loss_tokens": 8280000, "grad_norm": 1.4609375, "lr": 3e-05, "finish_rate": 0.132, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 33.8, "t_step_s": 69.7, "t_refresh_s": 0.3, "mem_gb": 10.38} {"step": 70, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.28901881557305653, "tokens": 120000, "cumulative_loss_tokens": 8400000, "grad_norm": 1.3828125, "lr": 3e-05, "finish_rate": 0.108, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 36.2, "t_step_s": 73.7, "t_refresh_s": 0.3, "mem_gb": 10.54} [eval step 70] sample: 'To solve this problem, we need to determine the maximum number of maples that can be planted along the alley given the constraints:\n\n1. The total number of trees is 75.\n2. There are no two maples betw' {"step": 70, "gsm8k_n": 64, "gsm8k_quick_chat": 0.28125, "t_eval_s": 44.7} {"step": 71, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.281591778382659, "tokens": 120000, "cumulative_loss_tokens": 8520000, "grad_norm": 1.53125, "lr": 3e-05, "finish_rate": 0.144, "comp_len": 493.8, "t_data_s": 0.0, "t_rollout_s": 33.2, "t_step_s": 69.5, "t_refresh_s": 0.3, "mem_gb": 12.44} {"step": 72, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.27427094464078544, "tokens": 120000, "cumulative_loss_tokens": 8640000, "grad_norm": 1.4375, "lr": 3e-05, "finish_rate": 0.107, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 35.8, "t_step_s": 73.1, "t_refresh_s": 0.3, "mem_gb": 10.47} {"step": 73, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.25415447360488275, "tokens": 120000, "cumulative_loss_tokens": 8760000, "grad_norm": 1.3359375, "lr": 3e-05, "finish_rate": 0.116, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.4, "t_step_s": 71.6, "t_refresh_s": 0.3, "mem_gb": 10.38} {"step": 74, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2598834279651443, "tokens": 120000, "cumulative_loss_tokens": 8880000, "grad_norm": 1.1875, "lr": 3e-05, "finish_rate": 0.109, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 35.8, "t_step_s": 72.9, "t_refresh_s": 0.3, "mem_gb": 10.46} {"step": 75, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.3036417054095616, "tokens": 120000, "cumulative_loss_tokens": 9000000, "grad_norm": 1.40625, "lr": 3e-05, "finish_rate": 0.133, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.9, "t_step_s": 71.1, "t_refresh_s": 0.3, "mem_gb": 10.41} {"step": 76, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.28468160825570427, "tokens": 120000, "cumulative_loss_tokens": 9120000, "grad_norm": 1.40625, "lr": 3e-05, "finish_rate": 0.124, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.7, "t_step_s": 71.7, "t_refresh_s": 0.3, "mem_gb": 10.45} {"step": 77, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.26907755986650783, "tokens": 120000, "cumulative_loss_tokens": 9240000, "grad_norm": 1.2109375, "lr": 3e-05, "finish_rate": 0.137, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.0, "t_step_s": 70.4, "t_refresh_s": 0.3, "mem_gb": 10.46} {"step": 78, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.23572471514294546, "tokens": 120000, "cumulative_loss_tokens": 9360000, "grad_norm": 1.3125, "lr": 3e-05, "finish_rate": 0.113, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 36.0, "t_step_s": 72.0, "t_refresh_s": 0.3, "mem_gb": 10.45} {"step": 79, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2453982294137279, "tokens": 120000, "cumulative_loss_tokens": 9480000, "grad_norm": 1.296875, "lr": 3e-05, "finish_rate": 0.156, "comp_len": 491.8, "t_data_s": 0.0, "t_rollout_s": 33.0, "t_step_s": 69.6, "t_refresh_s": 0.3, "mem_gb": 10.43} {"step": 80, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2668868999974181, "tokens": 120000, "cumulative_loss_tokens": 9600000, "grad_norm": 2.328125, "lr": 3e-05, "finish_rate": 0.149, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.2, "t_step_s": 70.1, "t_refresh_s": 0.3, "mem_gb": 10.4} [eval step 80] sample: 'To solve this problem, we need to determine the maximum number of maples that can be planted along the alley given the constraints:\n\n1. The total number of trees is 75.\n2. There are no two maples betw' {"step": 80, "gsm8k_n": 64, "gsm8k_quick_chat": 0.390625, "t_eval_s": 44.8} {"step": 81, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.27245304357651623, "tokens": 120000, "cumulative_loss_tokens": 9720000, "grad_norm": 1.46875, "lr": 3e-05, "finish_rate": 0.097, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 35.7, "t_step_s": 71.9, "t_refresh_s": 0.3, "mem_gb": 12.44} {"step": 82, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.26285181757311027, "tokens": 120000, "cumulative_loss_tokens": 9840000, "grad_norm": 1.1875, "lr": 3e-05, "finish_rate": 0.117, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 35.5, "t_step_s": 71.6, "t_refresh_s": 0.3, "mem_gb": 10.42} {"step": 83, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2504500327867145, "tokens": 120000, "cumulative_loss_tokens": 9960000, "grad_norm": 1.390625, "lr": 3e-05, "finish_rate": 0.157, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 34.4, "t_step_s": 71.2, "t_refresh_s": 0.3, "mem_gb": 10.46} {"step": 84, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.25419983249952394, "tokens": 120000, "cumulative_loss_tokens": 10080000, "grad_norm": 1.640625, "lr": 3e-05, "finish_rate": 0.165, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 33.9, "t_step_s": 70.4, "t_refresh_s": 0.3, "mem_gb": 10.43} {"step": 85, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.24969746678496402, "tokens": 120000, "cumulative_loss_tokens": 10200000, "grad_norm": 1.3203125, "lr": 3e-05, "finish_rate": 0.117, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 35.8, "t_step_s": 73.1, "t_refresh_s": 0.3, "mem_gb": 10.47} {"step": 86, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.26418633902855215, "tokens": 120000, "cumulative_loss_tokens": 10320000, "grad_norm": 1.3984375, "lr": 3e-05, "finish_rate": 0.172, "comp_len": 491.8, "t_data_s": 0.0, "t_rollout_s": 34.2, "t_step_s": 71.4, "t_refresh_s": 0.3, "mem_gb": 10.48} {"step": 87, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.25853364124981065, "tokens": 120000, "cumulative_loss_tokens": 10440000, "grad_norm": 1.3515625, "lr": 3e-05, "finish_rate": 0.137, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.3, "t_step_s": 70.2, "t_refresh_s": 0.3, "mem_gb": 10.4} {"step": 88, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.26424587182166676, "tokens": 120000, "cumulative_loss_tokens": 10560000, "grad_norm": 1.328125, "lr": 3e-05, "finish_rate": 0.132, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 34.2, "t_step_s": 70.9, "t_refresh_s": 0.3, "mem_gb": 10.48} {"step": 89, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.247395279224962, "tokens": 120000, "cumulative_loss_tokens": 10680000, "grad_norm": 1.2109375, "lr": 3e-05, "finish_rate": 0.12, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.5, "t_step_s": 70.4, "t_refresh_s": 0.3, "mem_gb": 10.37} {"step": 90, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.27036167939615746, "tokens": 120000, "cumulative_loss_tokens": 10800000, "grad_norm": 1.3359375, "lr": 3e-05, "finish_rate": 0.112, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 34.6, "t_step_s": 71.2, "t_refresh_s": 0.3, "mem_gb": 10.53} [eval step 90] sample: "To solve this problem, we need to maximize the number of maples \\( M \\) such that the total number of trees \\( T \\) is 75, and there are no two maples between which there are exactly 5 trees.\n\nLet's b" {"step": 90, "gsm8k_n": 64, "gsm8k_quick_chat": 0.359375, "t_eval_s": 44.7} {"step": 91, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2520888489575436, "tokens": 120000, "cumulative_loss_tokens": 10920000, "grad_norm": 1.203125, "lr": 3e-05, "finish_rate": 0.216, "comp_len": 489.8, "t_data_s": 0.0, "t_rollout_s": 33.6, "t_step_s": 70.4, "t_refresh_s": 0.3, "mem_gb": 12.44} {"step": 92, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.23066797492839397, "tokens": 120000, "cumulative_loss_tokens": 11040000, "grad_norm": 1.2421875, "lr": 3e-05, "finish_rate": 0.088, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 34.6, "t_step_s": 70.4, "t_refresh_s": 0.3, "mem_gb": 10.43} {"step": 93, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.23549754046387972, "tokens": 120000, "cumulative_loss_tokens": 11160000, "grad_norm": 1.3125, "lr": 3e-05, "finish_rate": 0.088, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 37.7, "t_step_s": 74.4, "t_refresh_s": 0.3, "mem_gb": 10.54} {"step": 94, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2587053025153776, "tokens": 120000, "cumulative_loss_tokens": 11280000, "grad_norm": 1.40625, "lr": 3e-05, "finish_rate": 0.126, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 35.1, "t_step_s": 71.5, "t_refresh_s": 0.3, "mem_gb": 10.45} {"step": 95, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.25772884955219927, "tokens": 120000, "cumulative_loss_tokens": 11400000, "grad_norm": 1.2265625, "lr": 3e-05, "finish_rate": 0.059, "comp_len": 506.3, "t_data_s": 0.0, "t_rollout_s": 37.7, "t_step_s": 74.7, "t_refresh_s": 0.3, "mem_gb": 10.52} {"step": 96, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2583792527654519, "tokens": 120000, "cumulative_loss_tokens": 11520000, "grad_norm": 1.296875, "lr": 3e-05, "finish_rate": 0.145, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 33.8, "t_step_s": 70.6, "t_refresh_s": 0.3, "mem_gb": 10.47} {"step": 97, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2516133760511875, "tokens": 120000, "cumulative_loss_tokens": 11640000, "grad_norm": 1.3984375, "lr": 3e-05, "finish_rate": 0.174, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 33.8, "t_step_s": 71.0, "t_refresh_s": 0.3, "mem_gb": 10.56} {"step": 98, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2511387197598815, "tokens": 120000, "cumulative_loss_tokens": 11760000, "grad_norm": 1.25, "lr": 3e-05, "finish_rate": 0.083, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 36.0, "t_step_s": 72.5, "t_refresh_s": 0.3, "mem_gb": 10.5} {"step": 99, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.25845928863280765, "tokens": 120000, "cumulative_loss_tokens": 11880000, "grad_norm": 1.328125, "lr": 3e-05, "finish_rate": 0.121, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 35.0, "t_step_s": 71.0, "t_refresh_s": 0.3, "mem_gb": 10.4} {"step": 100, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2377076770828416, "tokens": 120000, "cumulative_loss_tokens": 12000000, "grad_norm": 1.203125, "lr": 3e-05, "finish_rate": 0.177, "comp_len": 493.8, "t_data_s": 0.0, "t_rollout_s": 32.8, "t_step_s": 69.4, "t_refresh_s": 0.3, "mem_gb": 10.38} [eval step 100] sample: 'To solve this problem, we need to maximize the number of maples (M) that can be planted along the alley such that the total number of trees (T) is 75, and there are no two maples between which there a' {"step": 100, "gsm8k_n": 64, "gsm8k_quick_chat": 0.328125, "t_eval_s": 44.8} checkpoint snapshot queued -> outputs/healed/healing_breadth/glean_math_keep25_seed1224_long/step0100 {"step": 101, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2684837412368506, "tokens": 120000, "cumulative_loss_tokens": 12120000, "grad_norm": 1.296875, "lr": 3e-05, "finish_rate": 0.196, "comp_len": 489.8, "t_data_s": 0.0, "t_rollout_s": 33.0, "t_step_s": 69.4, "t_refresh_s": 0.3, "mem_gb": 12.44} {"step": 102, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.24094757018235202, "tokens": 120000, "cumulative_loss_tokens": 12240000, "grad_norm": 1.3046875, "lr": 3e-05, "finish_rate": 0.141, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.0, "t_step_s": 70.5, "t_refresh_s": 0.3, "mem_gb": 10.38} {"step": 103, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2634418107945472, "tokens": 120000, "cumulative_loss_tokens": 12360000, "grad_norm": 1.2265625, "lr": 3e-05, "finish_rate": 0.181, "comp_len": 493.8, "t_data_s": 0.0, "t_rollout_s": 33.4, "t_step_s": 70.0, "t_refresh_s": 0.3, "mem_gb": 10.4} {"step": 104, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22899044912308456, "tokens": 120000, "cumulative_loss_tokens": 12480000, "grad_norm": 1.171875, "lr": 3e-05, "finish_rate": 0.136, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 33.8, "t_step_s": 70.1, "t_refresh_s": 0.3, "mem_gb": 10.38} {"step": 105, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.23377793765577176, "tokens": 120000, "cumulative_loss_tokens": 12600000, "grad_norm": 1.203125, "lr": 3e-05, "finish_rate": 0.128, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 34.9, "t_step_s": 72.2, "t_refresh_s": 0.3, "mem_gb": 10.52} {"step": 106, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.23661676298119128, "tokens": 120000, "cumulative_loss_tokens": 12720000, "grad_norm": 1.1484375, "lr": 3e-05, "finish_rate": 0.121, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 35.6, "t_step_s": 72.4, "t_refresh_s": 0.3, "mem_gb": 10.46} {"step": 107, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22338223065932591, "tokens": 120000, "cumulative_loss_tokens": 12840000, "grad_norm": 1.1328125, "lr": 3e-05, "finish_rate": 0.145, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 34.3, "t_step_s": 71.2, "t_refresh_s": 0.3, "mem_gb": 10.49} {"step": 108, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.24060181738672157, "tokens": 120000, "cumulative_loss_tokens": 12960000, "grad_norm": 1.34375, "lr": 3e-05, "finish_rate": 0.168, "comp_len": 491.8, "t_data_s": 0.0, "t_rollout_s": 33.4, "t_step_s": 70.5, "t_refresh_s": 0.3, "mem_gb": 10.54} {"step": 109, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22999099592032532, "tokens": 120000, "cumulative_loss_tokens": 13080000, "grad_norm": 1.1796875, "lr": 3e-05, "finish_rate": 0.161, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 33.9, "t_step_s": 70.3, "t_refresh_s": 0.3, "mem_gb": 10.37} {"step": 110, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.24095743914954365, "tokens": 120000, "cumulative_loss_tokens": 13200000, "grad_norm": 1.578125, "lr": 3e-05, "finish_rate": 0.112, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 35.2, "t_step_s": 71.6, "t_refresh_s": 0.3, "mem_gb": 10.48} [eval step 110] sample: 'To solve this problem, we need to determine the maximum number of maples (\\(M\\)) that can be planted along the alley such that the total number of trees (\\(T\\)) is 75, and there are no two maples betw' {"step": 110, "gsm8k_n": 64, "gsm8k_quick_chat": 0.390625, "t_eval_s": 45.3} {"step": 111, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21752187181736032, "tokens": 120000, "cumulative_loss_tokens": 13320000, "grad_norm": 1.0859375, "lr": 3e-05, "finish_rate": 0.217, "comp_len": 481.9, "t_data_s": 0.0, "t_rollout_s": 30.9, "t_step_s": 68.1, "t_refresh_s": 0.3, "mem_gb": 12.44} {"step": 112, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2314429316678395, "tokens": 120000, "cumulative_loss_tokens": 13440000, "grad_norm": 1.3359375, "lr": 3e-05, "finish_rate": 0.12, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.8, "t_step_s": 71.5, "t_refresh_s": 0.3, "mem_gb": 10.52} {"step": 113, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2235619456699739, "tokens": 120000, "cumulative_loss_tokens": 13560000, "grad_norm": 1.1484375, "lr": 3e-05, "finish_rate": 0.137, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 33.9, "t_step_s": 71.1, "t_refresh_s": 0.3, "mem_gb": 10.55} {"step": 114, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.23131392294329903, "tokens": 120000, "cumulative_loss_tokens": 13680000, "grad_norm": 1.234375, "lr": 3e-05, "finish_rate": 0.124, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 34.3, "t_step_s": 71.2, "t_refresh_s": 0.3, "mem_gb": 10.38} {"step": 115, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22863815203296642, "tokens": 120000, "cumulative_loss_tokens": 13800000, "grad_norm": 1.2578125, "lr": 3e-05, "finish_rate": 0.104, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.9, "t_step_s": 71.7, "t_refresh_s": 0.3, "mem_gb": 10.4} {"step": 116, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22345312169020376, "tokens": 120000, "cumulative_loss_tokens": 13920000, "grad_norm": 1.3203125, "lr": 3e-05, "finish_rate": 0.125, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 35.2, "t_step_s": 72.0, "t_refresh_s": 0.3, "mem_gb": 10.48} {"step": 117, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22978670983004074, "tokens": 120000, "cumulative_loss_tokens": 14040000, "grad_norm": 1.1328125, "lr": 3e-05, "finish_rate": 0.108, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 35.8, "t_step_s": 73.0, "t_refresh_s": 0.3, "mem_gb": 10.52} {"step": 118, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2331024925402055, "tokens": 120000, "cumulative_loss_tokens": 14160000, "grad_norm": 1.28125, "lr": 3e-05, "finish_rate": 0.076, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 37.3, "t_step_s": 73.9, "t_refresh_s": 0.3, "mem_gb": 10.46} {"step": 119, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22393636475646247, "tokens": 120000, "cumulative_loss_tokens": 14280000, "grad_norm": 1.234375, "lr": 3e-05, "finish_rate": 0.13, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 35.4, "t_step_s": 71.5, "t_refresh_s": 0.3, "mem_gb": 10.44} {"step": 120, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2355032395929719, "tokens": 120000, "cumulative_loss_tokens": 14400000, "grad_norm": 1.203125, "lr": 3e-05, "finish_rate": 0.16, "comp_len": 491.8, "t_data_s": 0.0, "t_rollout_s": 33.1, "t_step_s": 68.8, "t_refresh_s": 0.3, "mem_gb": 10.36} [eval step 120] sample: 'To solve this problem, we need to determine the maximum number of maples that can be planted along the alley given the constraints:\n\n1. The total number of trees is 75.\n2. There are no two maples betw' {"step": 120, "gsm8k_n": 64, "gsm8k_quick_chat": 0.34375, "t_eval_s": 44.9} {"step": 121, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.24267969185064237, "tokens": 120000, "cumulative_loss_tokens": 14520000, "grad_norm": 1.171875, "lr": 3e-05, "finish_rate": 0.113, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 34.9, "t_step_s": 71.1, "t_refresh_s": 0.3, "mem_gb": 12.43} {"step": 122, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.23425059191025793, "tokens": 120000, "cumulative_loss_tokens": 14640000, "grad_norm": 1.1640625, "lr": 3e-05, "finish_rate": 0.093, "comp_len": 506.3, "t_data_s": 0.0, "t_rollout_s": 37.0, "t_step_s": 74.2, "t_refresh_s": 0.3, "mem_gb": 10.53} {"step": 123, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.26844764651320874, "tokens": 120000, "cumulative_loss_tokens": 14760000, "grad_norm": 1.359375, "lr": 3e-05, "finish_rate": 0.096, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 35.0, "t_step_s": 71.1, "t_refresh_s": 0.3, "mem_gb": 10.47} {"step": 124, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21723629867484173, "tokens": 120000, "cumulative_loss_tokens": 14880000, "grad_norm": 1.2578125, "lr": 3e-05, "finish_rate": 0.116, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 33.6, "t_step_s": 69.0, "t_refresh_s": 0.3, "mem_gb": 10.37} {"step": 125, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21943003199926267, "tokens": 120000, "cumulative_loss_tokens": 15000000, "grad_norm": 1.3515625, "lr": 3e-05, "finish_rate": 0.14, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 34.5, "t_step_s": 70.8, "t_refresh_s": 0.3, "mem_gb": 10.49} {"step": 126, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.214982470301042, "tokens": 120000, "cumulative_loss_tokens": 15120000, "grad_norm": 1.171875, "lr": 3e-05, "finish_rate": 0.13, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 34.9, "t_step_s": 70.3, "t_refresh_s": 0.3, "mem_gb": 10.37} {"step": 127, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22647308003318806, "tokens": 120000, "cumulative_loss_tokens": 15240000, "grad_norm": 1.34375, "lr": 3e-05, "finish_rate": 0.105, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 35.0, "t_step_s": 71.0, "t_refresh_s": 0.3, "mem_gb": 10.38} {"step": 128, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.23893264475651085, "tokens": 120000, "cumulative_loss_tokens": 15360000, "grad_norm": 1.4140625, "lr": 3e-05, "finish_rate": 0.104, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 35.3, "t_step_s": 71.4, "t_refresh_s": 0.3, "mem_gb": 10.5} {"step": 129, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.23907007528046767, "tokens": 120000, "cumulative_loss_tokens": 15480000, "grad_norm": 1.1796875, "lr": 3e-05, "finish_rate": 0.127, "comp_len": 491.8, "t_data_s": 0.0, "t_rollout_s": 31.7, "t_step_s": 69.0, "t_refresh_s": 0.3, "mem_gb": 10.47} {"step": 130, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.23370486757767697, "tokens": 120000, "cumulative_loss_tokens": 15600000, "grad_norm": 1.2421875, "lr": 3e-05, "finish_rate": 0.139, "comp_len": 489.8, "t_data_s": 0.0, "t_rollout_s": 32.2, "t_step_s": 68.9, "t_refresh_s": 0.3, "mem_gb": 10.43} [eval step 130] sample: 'To solve this problem, we need to determine the maximum number of maples that can be planted along the alley given the constraints:\n\n1. The total number of trees is 75.\n2. There are no two maples betw' {"step": 130, "gsm8k_n": 64, "gsm8k_quick_chat": 0.375, "t_eval_s": 45.3} {"step": 131, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22123020526878537, "tokens": 120000, "cumulative_loss_tokens": 15720000, "grad_norm": 1.2421875, "lr": 3e-05, "finish_rate": 0.088, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 36.5, "t_step_s": 73.1, "t_refresh_s": 0.3, "mem_gb": 12.44} {"step": 132, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22262640608406314, "tokens": 120000, "cumulative_loss_tokens": 15840000, "grad_norm": 1.203125, "lr": 3e-05, "finish_rate": 0.129, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 33.3, "t_step_s": 69.3, "t_refresh_s": 0.3, "mem_gb": 10.38} {"step": 133, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.23642245450820773, "tokens": 120000, "cumulative_loss_tokens": 15960000, "grad_norm": 1.3125, "lr": 3e-05, "finish_rate": 0.141, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 33.9, "t_step_s": 70.1, "t_refresh_s": 0.3, "mem_gb": 10.5} {"step": 134, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22181539568305014, "tokens": 120000, "cumulative_loss_tokens": 16080000, "grad_norm": 1.1484375, "lr": 3e-05, "finish_rate": 0.124, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 35.1, "t_step_s": 71.9, "t_refresh_s": 0.3, "mem_gb": 10.49} {"step": 135, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21886845049690457, "tokens": 120000, "cumulative_loss_tokens": 16200000, "grad_norm": 1.265625, "lr": 3e-05, "finish_rate": 0.088, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 35.4, "t_step_s": 71.3, "t_refresh_s": 0.3, "mem_gb": 10.43} {"step": 136, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.23528965907134117, "tokens": 120000, "cumulative_loss_tokens": 16320000, "grad_norm": 1.2421875, "lr": 3e-05, "finish_rate": 0.117, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 35.0, "t_step_s": 71.1, "t_refresh_s": 0.3, "mem_gb": 10.45} {"step": 137, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22028850772418082, "tokens": 120000, "cumulative_loss_tokens": 16440000, "grad_norm": 1.3828125, "lr": 3e-05, "finish_rate": 0.076, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 36.0, "t_step_s": 71.8, "t_refresh_s": 0.3, "mem_gb": 10.47} {"step": 138, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22388356751191119, "tokens": 120000, "cumulative_loss_tokens": 16560000, "grad_norm": 1.171875, "lr": 3e-05, "finish_rate": 0.123, "comp_len": 493.8, "t_data_s": 0.0, "t_rollout_s": 33.4, "t_step_s": 69.5, "t_refresh_s": 0.3, "mem_gb": 10.38} {"step": 139, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2186046968769282, "tokens": 120000, "cumulative_loss_tokens": 16680000, "grad_norm": 1.1171875, "lr": 3e-05, "finish_rate": 0.084, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 36.2, "t_step_s": 72.8, "t_refresh_s": 0.3, "mem_gb": 10.46} {"step": 140, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21518618610985576, "tokens": 120000, "cumulative_loss_tokens": 16800000, "grad_norm": 1.1171875, "lr": 3e-05, "finish_rate": 0.121, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 33.9, "t_step_s": 69.8, "t_refresh_s": 0.3, "mem_gb": 10.44} [eval step 140] sample: 'To solve this problem, we need to carefully analyze the constraints given:\n\n1. **Total number of trees:** 75 trees.\n2. **No two maples between which there are exactly 5 trees:** This means that no two' {"step": 140, "gsm8k_n": 64, "gsm8k_quick_chat": 0.359375, "t_eval_s": 44.8} {"step": 141, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21726824556073795, "tokens": 120000, "cumulative_loss_tokens": 16920000, "grad_norm": 1.1953125, "lr": 3e-05, "finish_rate": 0.072, "comp_len": 506.3, "t_data_s": 0.0, "t_rollout_s": 37.1, "t_step_s": 73.1, "t_refresh_s": 0.3, "mem_gb": 12.43} {"step": 142, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21776566527485847, "tokens": 120000, "cumulative_loss_tokens": 17040000, "grad_norm": 1.2109375, "lr": 3e-05, "finish_rate": 0.104, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 33.9, "t_step_s": 70.1, "t_refresh_s": 0.3, "mem_gb": 10.4} {"step": 143, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21226181217512738, "tokens": 120000, "cumulative_loss_tokens": 17160000, "grad_norm": 1.234375, "lr": 3e-05, "finish_rate": 0.128, "comp_len": 493.8, "t_data_s": 0.0, "t_rollout_s": 33.6, "t_step_s": 70.3, "t_refresh_s": 0.3, "mem_gb": 10.47} {"step": 144, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.20745828753014406, "tokens": 120000, "cumulative_loss_tokens": 17280000, "grad_norm": 1.09375, "lr": 3e-05, "finish_rate": 0.141, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 35.2, "t_step_s": 71.5, "t_refresh_s": 0.3, "mem_gb": 10.46} {"step": 145, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.20095169744286687, "tokens": 120000, "cumulative_loss_tokens": 17400000, "grad_norm": 1.1796875, "lr": 3e-05, "finish_rate": 0.067, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 37.0, "t_step_s": 72.8, "t_refresh_s": 0.3, "mem_gb": 10.5} {"step": 146, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21430944070120653, "tokens": 120000, "cumulative_loss_tokens": 17520000, "grad_norm": 1.15625, "lr": 3e-05, "finish_rate": 0.092, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 35.2, "t_step_s": 71.4, "t_refresh_s": 0.3, "mem_gb": 10.41} {"step": 147, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.20074411552796761, "tokens": 120000, "cumulative_loss_tokens": 17640000, "grad_norm": 1.1015625, "lr": 3e-05, "finish_rate": 0.092, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 35.3, "t_step_s": 71.3, "t_refresh_s": 0.3, "mem_gb": 10.51} {"step": 148, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2208368306187292, "tokens": 120000, "cumulative_loss_tokens": 17760000, "grad_norm": 1.1328125, "lr": 3e-05, "finish_rate": 0.148, "comp_len": 493.8, "t_data_s": 0.0, "t_rollout_s": 33.4, "t_step_s": 69.9, "t_refresh_s": 0.3, "mem_gb": 10.4} {"step": 149, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22132619506145518, "tokens": 120000, "cumulative_loss_tokens": 17880000, "grad_norm": 1.265625, "lr": 3e-05, "finish_rate": 0.108, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 36.8, "t_step_s": 74.1, "t_refresh_s": 0.3, "mem_gb": 10.47} {"step": 150, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2195085643796871, "tokens": 120000, "cumulative_loss_tokens": 18000000, "grad_norm": 1.1328125, "lr": 3e-05, "finish_rate": 0.112, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 36.1, "t_step_s": 73.3, "t_refresh_s": 0.3, "mem_gb": 10.53} [eval step 150] sample: "To solve this problem, we need to determine the maximum number of maples that can be placed along the alley such that there are no two maples between which there are exactly 5 trees.\n\nLet's break down" {"step": 150, "gsm8k_n": 64, "gsm8k_quick_chat": 0.3125, "t_eval_s": 44.9} checkpoint snapshot queued -> outputs/healed/healing_breadth/glean_math_keep25_seed1224_long/step0150 {"step": 151, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21052663061742982, "tokens": 120000, "cumulative_loss_tokens": 18120000, "grad_norm": 1.234375, "lr": 3e-05, "finish_rate": 0.1, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 35.0, "t_step_s": 70.8, "t_refresh_s": 0.3, "mem_gb": 12.43} {"step": 152, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.23573231500076752, "tokens": 120000, "cumulative_loss_tokens": 18240000, "grad_norm": 1.140625, "lr": 3e-05, "finish_rate": 0.076, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 37.6, "t_step_s": 74.7, "t_refresh_s": 0.3, "mem_gb": 10.54} {"step": 153, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21262744003695747, "tokens": 120000, "cumulative_loss_tokens": 18360000, "grad_norm": 1.125, "lr": 3e-05, "finish_rate": 0.152, "comp_len": 493.8, "t_data_s": 0.0, "t_rollout_s": 33.7, "t_step_s": 70.2, "t_refresh_s": 0.3, "mem_gb": 10.46} {"step": 154, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.20737704386965683, "tokens": 120000, "cumulative_loss_tokens": 18480000, "grad_norm": 1.125, "lr": 3e-05, "finish_rate": 0.155, "comp_len": 489.8, "t_data_s": 0.0, "t_rollout_s": 32.1, "t_step_s": 69.1, "t_refresh_s": 0.3, "mem_gb": 10.4} {"step": 155, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.20515705520591387, "tokens": 120000, "cumulative_loss_tokens": 18600000, "grad_norm": 1.125, "lr": 3e-05, "finish_rate": 0.125, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 34.8, "t_step_s": 72.9, "t_refresh_s": 0.3, "mem_gb": 10.47} {"step": 156, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.227416008350874, "tokens": 120000, "cumulative_loss_tokens": 18720000, "grad_norm": 1.1640625, "lr": 3e-05, "finish_rate": 0.143, "comp_len": 491.8, "t_data_s": 0.0, "t_rollout_s": 33.7, "t_step_s": 70.4, "t_refresh_s": 0.3, "mem_gb": 10.46} {"step": 157, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2179569326767077, "tokens": 120000, "cumulative_loss_tokens": 18840000, "grad_norm": 1.1953125, "lr": 3e-05, "finish_rate": 0.116, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.7, "t_step_s": 70.6, "t_refresh_s": 0.3, "mem_gb": 10.37} {"step": 158, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21790637281499803, "tokens": 120000, "cumulative_loss_tokens": 18960000, "grad_norm": 1.1953125, "lr": 3e-05, "finish_rate": 0.12, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 33.5, "t_step_s": 69.6, "t_refresh_s": 0.3, "mem_gb": 10.37} {"step": 159, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21037111974650374, "tokens": 120000, "cumulative_loss_tokens": 19080000, "grad_norm": 1.25, "lr": 3e-05, "finish_rate": 0.124, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.3, "t_step_s": 70.5, "t_refresh_s": 0.3, "mem_gb": 10.39} {"step": 160, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22135015249016385, "tokens": 120000, "cumulative_loss_tokens": 19200000, "grad_norm": 1.234375, "lr": 3e-05, "finish_rate": 0.1, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 35.2, "t_step_s": 72.0, "t_refresh_s": 0.3, "mem_gb": 10.44} [eval step 160] sample: "To solve this problem, we need to determine the maximum number of maples that can be placed along the alley such that there are no two maples between which there are exactly 5 trees.\n\nLet's break down" {"step": 160, "gsm8k_n": 64, "gsm8k_quick_chat": 0.375, "t_eval_s": 44.8} {"step": 161, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.19312054982806245, "tokens": 120000, "cumulative_loss_tokens": 19320000, "grad_norm": 1.078125, "lr": 3e-05, "finish_rate": 0.076, "comp_len": 506.3, "t_data_s": 0.0, "t_rollout_s": 37.0, "t_step_s": 73.0, "t_refresh_s": 0.3, "mem_gb": 12.43} {"step": 162, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21525976048012574, "tokens": 120000, "cumulative_loss_tokens": 19440000, "grad_norm": 1.2734375, "lr": 3e-05, "finish_rate": 0.177, "comp_len": 493.8, "t_data_s": 0.0, "t_rollout_s": 33.4, "t_step_s": 69.6, "t_refresh_s": 0.3, "mem_gb": 10.38} {"step": 163, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21099251879341901, "tokens": 120000, "cumulative_loss_tokens": 19560000, "grad_norm": 1.171875, "lr": 3e-05, "finish_rate": 0.108, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 35.3, "t_step_s": 71.9, "t_refresh_s": 0.3, "mem_gb": 10.41} {"step": 164, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2188319661433498, "tokens": 120000, "cumulative_loss_tokens": 19680000, "grad_norm": 1.1953125, "lr": 3e-05, "finish_rate": 0.136, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 33.7, "t_step_s": 70.1, "t_refresh_s": 0.3, "mem_gb": 10.47} {"step": 165, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21699987841347854, "tokens": 120000, "cumulative_loss_tokens": 19800000, "grad_norm": 1.1640625, "lr": 3e-05, "finish_rate": 0.137, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.7, "t_step_s": 72.0, "t_refresh_s": 0.3, "mem_gb": 10.54} {"step": 166, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.20422210597346227, "tokens": 120000, "cumulative_loss_tokens": 19920000, "grad_norm": 1.1875, "lr": 3e-05, "finish_rate": 0.088, "comp_len": 502.1, "t_data_s": 0.0, "t_rollout_s": 34.4, "t_step_s": 69.9, "t_refresh_s": 0.3, "mem_gb": 10.36} {"step": 167, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.21616086491048336, "tokens": 120000, "cumulative_loss_tokens": 20040000, "grad_norm": 1.109375, "lr": 3e-05, "finish_rate": 0.108, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 35.9, "t_step_s": 72.8, "t_refresh_s": 0.3, "mem_gb": 10.49} {"step": 168, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2041427251155178, "tokens": 120000, "cumulative_loss_tokens": 20160000, "grad_norm": 1.1796875, "lr": 3e-05, "finish_rate": 0.117, "comp_len": 500.0, "t_data_s": 0.0, "t_rollout_s": 35.2, "t_step_s": 71.1, "t_refresh_s": 0.3, "mem_gb": 10.39} {"step": 169, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.2007420184203113, "tokens": 120000, "cumulative_loss_tokens": 20280000, "grad_norm": 1.0546875, "lr": 3e-05, "finish_rate": 0.113, "comp_len": 504.2, "t_data_s": 0.0, "t_rollout_s": 36.2, "t_step_s": 72.4, "t_refresh_s": 0.3, "mem_gb": 10.51} {"step": 170, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.22800080033155778, "tokens": 120000, "cumulative_loss_tokens": 20400000, "grad_norm": 1.40625, "lr": 3e-05, "finish_rate": 0.153, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 33.6, "t_step_s": 69.9, "t_refresh_s": 0.3, "mem_gb": 10.42} [eval step 170] sample: 'To solve this problem, we need to carefully analyze the constraints and use reasoning to determine the maximum number of maples that can be planted.\n\n### Problem Breakdown:\n\n1. **Total Number of Trees' {"step": 170, "gsm8k_n": 64, "gsm8k_quick_chat": 0.34375, "t_eval_s": 44.8} {"step": 171, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.20762226390354335, "tokens": 120000, "cumulative_loss_tokens": 20520000, "grad_norm": 1.2421875, "lr": 3e-05, "finish_rate": 0.112, "comp_len": 497.9, "t_data_s": 0.0, "t_rollout_s": 34.9, "t_step_s": 72.1, "t_refresh_s": 0.3, "mem_gb": 12.43} {"step": 172, "epoch": 0, "training_mode": "on-policy", "reverse_kl": 0.19745559071643898, "tokens": 120000, "cumulative_loss_tokens": 20640000, "grad_norm": 1.1875, "lr": 3e-05, "finish_rate": 0.145, "comp_len": 495.9, "t_data_s": 0.0, "t_rollout_s": 33.4, "t_step_s": 69.5, "t_refresh_s": 0.3, "mem_gb": 10.38}