LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents Paper • 2608.17393 • Published 10 days ago • 24
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design Paper • 2608.10299 • Published 18 days ago • 135
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty Paper • 2412.20251 • Published May 25, 2025
CLLMate: A Multimodal Benchmark for Weather and Climate Events Forecasting Paper • 2409.19058 • Published Feb 16, 2025 • 1
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models Paper • 2605.27243 • Published May 26
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models Paper • 2605.14906 • Published May 14 • 79
LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing Paper • 2406.16253 • Published Jun 24, 2024
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design Paper • 2608.10299 • Published 18 days ago • 135
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published Jul 19 • 93
Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games Paper • 2606.19338 • Published Jun 17 • 51
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context Paper • 2605.13831 • Published May 13 • 90 • 3
Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation Paper • 2606.12594 • Published Jun 10 • 17
SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks Paper • 2605.31433 • Published May 29 • 28
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid? Paper • 2605.06527 • Published May 7 • 48
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation Paper • 2605.10912 • Published May 11 • 46
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models Paper • 2605.14906 • Published May 14 • 79
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models Paper • 2605.14906 • Published May 14 • 79
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid? Paper • 2605.06527 • Published May 7 • 48
MMProLong Collection A 7B LVLM with 128K context window and 512K generalization through long-context continued pre-training • 1 item • Updated May 15