sol-high.h043.opsd-r2e-commit-small-replay.run_default.broadcasts.step_1

AgentPTB sweep checkpoint. Cell sol-high โ€” Codex / gpt-5.6-sol @ effort high.

field value
plot cell sol-high
driver Codex / gpt-5.6-sol
reasoning effort high
run boot (UTC) 2026-08-08T07:28:19Z
role intermediate
hours into run h43.29 of 100
checkpoint path in run outputs/opsd-r2e-commit-small-replay/run_default/broadcasts/step_1
shards 4
size 18.8 GB
base model Qwen/Qwen3.5-9B-Base
eos_token_id None โš ๏ธ MISSING 248046

Reading the eos field

248046 is <|im_end|>, the token the Qwen3.5 chat template ends every assistant turn with. Checkpoints missing it do not stop at end-of-turn and overrun the context window, so their eval numbers are a floor, not a measurement โ€” compare them only against other checkpoints with the same eos status, or re-package before evaluating.

Cell note: best cell in the sweep

Mapping back to the figures

The repo id is {cell}.h{HHH}.{family}.{step}, where hHHH is the hour of the 100-hour run at which this checkpoint was written โ€” the same x-axis the sweep figures use for eval panels (t_h). So a checkpoint drops onto the performance-over-time curve directly, and sorting repo ids within a cell sorts them chronologically.

hHHH is rounded down to whole hours for sortability; the exact value is the hours into run row above, and in agentic-ptb/INDEX.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
9B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for agentic-ptb/sol-high.h043.opsd-r2e-commit-small-replay.run_default.broadcasts.step_1

Finetuned
(582)
this model