Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations Paper • 2609.22255 • Published 17 days ago • 13
Running 22 GLEE Competition — Live Leaderboard 🏆 22 Live leaderboard of the GLEE Competition @ NeurIPS 2026
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs Paper • 2604.14137 • Published Apr 16 • 10
LLM Explainability with Counterfactual Chains and Causal Graphs Paper • 2606.05972 • Published Jun 4 • 19
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks Paper • 2605.28556 • Published May 27 • 74
Efficient Agent Evaluation via Diversity-Guided User Simulation Paper • 2604.21480 • Published Apr 23 • 16
Alignment Makes Language Models Normative, Not Descriptive Paper • 2603.17218 • Published Mar 17 • 47
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs Paper • 2603.09906 • Published Mar 10 • 76
The Poisoned Apple Effect: Strategic Manipulation of Mediated Markets via Technology Expansion of AI Agents Paper • 2601.11496 • Published Jan 16 • 47