ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research Paper • 2610.02202 • Published 4 days ago • 11
AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs Paper • 2609.31590 • Published 10 days ago • 13
Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning Paper • 2609.38147 • Published 6 days ago • 1
Finetuning with Sampling: SFT Learns Better Than You Think Paper • 2610.02140 • Published 4 days ago • 1
From cacophony to hierarchy: a principled framework for assessing AI consciousness Paper • 2609.35618 • Published 6 days ago • 1
How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text Paper • 2609.40295 • Published 5 days ago • 4
Invent a Dataset: Measuring dataset generation abilities with zero seed Paper • 2610.01674 • Published 4 days ago • 1
Improving Test-Time Scaling with Adaptive Looped Transformers Paper • 2609.35748 • Published 7 days ago • 58
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents Paper • 2609.27334 • Published 12 days ago • 56
Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model Paper • 2609.40358 • Published 5 days ago • 14
KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling Paper • 2609.27294 • Published 12 days ago • 2
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 13 days ago • 37
Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It? Paper • 2609.27891 • Published Aug 21 • 32
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published Aug 26 • 193
Negative Self-Distillation: Learning to Reason by Avoiding Flaws Paper • 2609.11699 • Published 25 days ago • 37