Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories Paper • 2606.03979 • Published Jun 2 • 29
XiaoXu123123/academic-humanize-qwen25-7b-dpo-v2-lora Text Generation • Updated May 11 • 433 • 6
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices Paper • 2512.01374 • Published Dec 1, 2025 • 109