Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision Paper • 2604.12002 • Published Apr 13 • 12
Closing the Train-Test Gap in World Models for Gradient-Based Planning Paper • 2512.09929 • Published Dec 10, 2025
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence Paper • 2511.07384 • Published Nov 10, 2025 • 21
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning Paper • 2507.16746 • Published Jul 22, 2025 • 36
Reranking-based Generation for Unbiased Perspective Summarization Paper • 2506.15925 • Published Jun 19, 2025 • 5
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions Paper • 2502.04322 • Published Feb 6, 2025 • 3
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming Paper • 2501.18837 • Published Jan 31, 2025 • 10
Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution Paper • 2409.07072 • Published Sep 11, 2024
Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations Paper • 2307.08678 • Published Jul 17, 2023
Enhancing Few-shot Text-to-SQL Capabilities of Large Language Models: A Study on Prompt Design Strategies Paper • 2305.12586 • Published May 21, 2023
Contrastive Loss is All You Need to Recover Analogies as Parallel Lines Paper • 2306.08221 • Published Jun 14, 2023