Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models Paper • 2606.16700 • Published Jun 15 • 14
Skip a Layer or Loop It? Learning Program-of-Layers in LLMs Paper • 2606.06574 • Published Jun 4 • 25
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering Paper • 2401.13170 • Published Jan 24, 2024 • 4
PANDA (Pedantic ANswer-correctness Determination and Adjudication):Improving Automatic Evaluation for Question Answering and Text Generation Paper • 2402.11161 • Published Feb 17, 2024 • 2
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges Paper • 2501.02189 • Published Jan 4, 2025 • 1
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents Paper • 2511.07685 • Published Nov 10, 2025 • 10
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment Paper • 2512.04210 • Published Dec 3, 2025
SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones? Paper • 2605.30329 • Published May 28 • 8
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring Paper • 2604.19984 • Published Apr 21 • 1
Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning Paper • 2606.07631 • Published May 31 • 3
Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use Paper • 2605.14038 • Published May 13 • 15
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks Paper • 2604.20987 • Published Apr 22 • 22
Design-o-meter: Towards Evaluating and Refining Graphic Designs Paper • 2411.14959 • Published Nov 22, 2024 • 1