Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation Paper • 2510.04265 • Published May 12
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Paper • 2608.04001 • Published 25 days ago • 1
LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem Paper • 2403.00108 • Published Feb 29, 2024
More for Keys, Less for Values: Adaptive KV Cache Quantization Paper • 2502.15075 • Published Feb 20, 2025