Continual Learning with Soft Prompt Collection Private reproducibility artifacts for distilling a mathematical reasoning RL checkpoint into a 16-token soft prompt. • 3 items • Updated about 23 hours ago
Continual Learning with Soft Prompt Collection Private reproducibility artifacts for distilling a mathematical reasoning RL checkpoint into a 16-token soft prompt. • 3 items • Updated about 23 hours ago
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling Paper • 2604.27039 • Published Apr 29 • 25
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling Paper • 2604.27039 • Published Apr 29 • 25