DPWriter-Qwen3-4B-SFT

Cold-start SFT model of a reproduction of DPWriter (ACL 2026, arXiv 2601.09609), trained on Pitt CRC with the vendored verl 0.4.1 code of the DPWriter repository.

  • Base model: Qwen/Qwen3-4B-Base, full-parameter fine-tuning (no LoRA), weights stored in fp32.
  • Data: DPWriterData (38,328 rows), the 31,139 rows whose response carries the semi-structured plan (<think> + <goal> <info> <struct> <lang> <pres> tags + response); prompt+response truncated to 8192 tokens.
  • Tokenizer: the 12 plan tags (<think>, </think>, <goal>, </goal>, <info>, </info>, <struct>, </struct>, <lang>, </lang>, <pres>, </pres>) are added as special tokens; their embeddings are initialised with the mean of the existing embedding table plus small noise.
  • Training: verl fsdp_sft_trainer, 3 epochs, global batch 64, lr 1e-5 cosine with 10% warmup, bf16 compute, 2 x RTX PRO 6000 (96GB), 17 h. Final validation loss 1.554 (512 held-out training rows).
  • Chat template: Qwen3 template (chat_template.jinja); the model answers with <think>...</think> followed by the text.

This checkpoint is the starting point of the GRPO + Diverse Planning Branching RL stage (see DPWriter-Qwen3-4B-K32-ngram-lambda06-*).

Downloads last month
7
Safetensors
Model size
4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Johnny1024/DPWriter-Qwen3-4B-SFT

Finetuned
(487)
this model
Finetunes
2 models

Paper for Johnny1024/DPWriter-Qwen3-4B-SFT