MPCEval: A Benchmark for Multi-Party Conversation Generation Paper • 2603.04969 • Published Mar 5
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents Paper • 2606.28480 • Published Jun 26 • 49
Controlled Low-Rank Adaptation with Subspace Regularization for Continued Training on Large Language Models Paper • 2410.16801 • Published Mar 21, 2025
Batch-of-Thought: Cross-Instance Learning for Enhanced LLM Reasoning Paper • 2601.02950 • Published Mar 7 • 1
T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning Paper • 2603.03790 • Published Mar 4 • 79
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons Paper • 2503.05731 • Published Feb 19, 2025 • 3
A Survey of Vibe Coding with Large Language Models Paper • 2510.12399 • Published Oct 14, 2025 • 50
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks Paper • 2410.03769 • Published Oct 2, 2024