DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing
Paper • 2601.09609 • Published • 4
Cold-start SFT model of a reproduction of DPWriter (ACL 2026, arXiv 2601.09609), trained on Pitt CRC with the vendored verl 0.4.1 code of the DPWriter repository.
<think> + <goal> <info> <struct> <lang> <pres> tags + response); prompt+response truncated to 8192 tokens.<think>, </think>, <goal>, </goal>, <info>, </info>, <struct>, </struct>,
<lang>, </lang>, <pres>, </pres>) are added as special tokens; their embeddings are initialised with the mean
of the existing embedding table plus small noise.fsdp_sft_trainer, 3 epochs, global batch 64, lr 1e-5 cosine with 10% warmup, bf16 compute,
2 x RTX PRO 6000 (96GB), 17 h. Final validation loss 1.554 (512 held-out training rows).chat_template.jinja); the model answers with <think>...</think> followed by the text.This checkpoint is the starting point of the GRPO + Diverse Planning Branching RL stage
(see DPWriter-Qwen3-4B-K32-ngram-lambda06-*).