MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following Paper • 2605.03858 • Published May 5 • 1
Towards Understanding Multimodal Fine-Tuning: Spatial Features Paper • 2602.08713 • Published Feb 6 • 1
EVCL: Elastic Variational Continual Learning with Weight Consolidation Paper • 2406.15972 • Published Jun 23, 2024 • 1
Reasoning Fine-Tuning Induces Persistent Latent Policy States Paper • 2607.18532 • Published 16 days ago • 1
Constitutional Midtraining: Content Presence Drives Alignment Gains Paper • 2607.26654 • Published 7 days ago • 6
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks Paper • 2605.01417 • Published May 2 • 2
Measuring what Matters: Construct Validity in Large Language Model Benchmarks Paper • 2511.04703 • Published Nov 3, 2025 • 8
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought Paper • 2403.05518 • Published Mar 8, 2024 • 3
Constitutional Midtraining: Content Presence Drives Alignment Gains Paper • 2607.26654 • Published 7 days ago • 6
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Paper • 2606.07433 • Published Jun 5 • 21
Hallucination in World Models is Predictable and Preventable Paper • 2606.27326 • Published Jun 25 • 9
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Paper • 2501.14818 • Published Jan 20, 2025 • 10
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Paper • 2502.05178 • Published Feb 7, 2025 • 10
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Paper • 2503.14734 • Published Mar 18, 2025 • 9
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models Paper • 2504.03624 • Published Apr 4, 2025 • 20