Agent Plasticity: Measuring Self-Improvement Through Experience Paper • 2610.08902 • Published 6 days ago • 5
EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Paper • 2606.18967 • Published Jun 17 • 25
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution Paper • 2604.07725 • Published Apr 10
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning Paper • 2601.10649 • Published Apr 7
CDLM: Consistency Diffusion Language Models For Faster Sampling Paper • 2511.19269 • Published Nov 24, 2025 • 1
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities Paper • 2507.06261 • Published Jul 7, 2025 • 69
Agent Plasticity: Measuring Self-Improvement Through Experience Paper • 2610.08902 • Published 6 days ago • 5
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research Paper • 2511.19399 • Published Nov 24, 2025 • 64
$V_1$: Unifying Generation and Self-Verification for Parallel Reasoners Paper • 2603.04304 • Published Mar 4 • 14
V_1: Unifying Generation and Self-Verification for Parallel Reasoners Paper • 2603.04304 • Published Mar 4 • 14
GraPE: A Generate-Plan-Edit Framework for Compositional T2I Synthesis Paper • 2412.06089 • Published Dec 8, 2024 • 4
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages Paper • 2404.16816 • Published Apr 25, 2024 • 3
Scaling Retrieval-Based Language Models with a Trillion-Token Datastore Paper • 2407.12854 • Published Jul 9, 2024 • 31
Language models scale reliably with over-training and on downstream tasks Paper • 2403.08540 • Published Mar 13, 2024 • 15