Runtime error Agents 3 Leaderboard 🥇 3 Explore the E2LMC leaderboard and filter top team submissions
When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models Paper • 2606.10740 • Published Jun 9 • 2
When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models Paper • 2606.10740 • Published Jun 9 • 2
Mellum: Production-Grade in-IDE Contextual Code Completion with Multi-File Project Understanding Paper • 2510.05788 • Published Oct 7, 2025 • 4
OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond Paper • 2605.19660 • Published May 19 • 40
SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training Paper • 2605.18719 • Published May 18 • 7
RDP LoRA: Geometry-Driven Identification for Parameter-Efficient Adaptation in Large Language Models Paper • 2604.19321 • Published Apr 21 • 8
RDP LoRA: Geometry-Driven Identification for Parameter-Efficient Adaptation in Large Language Models Paper • 2604.19321 • Published Apr 21 • 8
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data Paper • 2604.12633 • Published Apr 14 • 2
Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models Paper • 2604.01622 • Published Apr 2 • 7
Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework Paper • 2604.06170 • Published Apr 7 • 31
Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps Paper • 2510.13430 • Published Oct 15, 2025 • 2
3LM: Bridging Arabic, STEM, and Code through Benchmarking Paper • 2507.15850 • Published Jul 21, 2025 • 6
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Paper • 2506.07731 • Published Jun 9, 2025 • 2
Are Arabic Benchmarks Reliable? QIMMA's Quality-First Approach to LLM Evaluation Paper • 2604.03395 • Published Apr 3 • 2
HateMirage: An Explainable Multi-Dimensional Dataset for Decoding Faux Hate and Subtle Online Abuse Paper • 2603.02684 • Published Mar 3 • 1
HateMirage: An Explainable Multi-Dimensional Dataset for Decoding Faux Hate and Subtle Online Abuse Paper • 2603.02684 • Published Mar 3 • 1