From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published 10 days ago • 88
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Paper • 2607.25659 • Published 8 days ago • 82
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 20 days ago • 103
Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs Paper • 2606.05846 • Published Jun 4 • 4
Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning Paper • 2605.11458 • Published May 12 • 7
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe Paper • 2604.13016 • Published Apr 14 • 114
Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs Paper • 2601.17058 • Published Jan 22 • 191
The Flexibility Trap: Why Arbitrary Order Limits Reasoning Potential in Diffusion Language Models Paper • 2601.15165 • Published Jan 21 • 75
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs Paper • 2601.11000 • Published Jan 16 • 27
PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models Paper • 2506.16054 • Published Jun 19, 2025 • 60
Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation Paper • 2506.03857 • Published Jun 4, 2025 • 2 • 2
Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation Paper • 2506.03857 • Published Jun 4, 2025 • 2