view article Article Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community ResterChed • 18 days ago • 192
RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation Paper • 2606.11709 • Published Jun 10 • 1 • 1
RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation Paper • 2606.11709 • Published Jun 10 • 1
Learning from Language Feedback via Variational Policy Distillation Paper • 2605.15113 • Published May 18 • 13
Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion Paper • 2606.14885 • Published Jun 12 • 12
Rethinking Continual Experience Internalization for Self-Evolving LLM Agents Paper • 2606.04703 • Published Jun 3 • 26