Hallucination in World Models is Predictable and Preventable Paper • 2606.27326 • Published Jun 25 • 9
MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery Paper • 2606.24595 • Published Jun 23 • 2
Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining Paper • 2606.16246 • Published Jun 19 • 4
EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video Paper • 2606.16202 • Published Jun 15 • 2
PIPE-Cypher: Automatic Enterprise Benchmark Generation for Text-to-Cypher Systems Paper • 2606.08481 • Published Jun 7
Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning Paper • 2606.03965 • Published Jun 2 • 1
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs Paper • 2606.00477 • Published May 30
BOOKMARKS: Efficient Active Storyline Memory for Role-playing Paper • 2605.14169 • Published May 13 • 8
FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration Paper • 2605.08520 • Published May 8 • 6
CellMaster: Collaborative Cell Type Annotation in Single-Cell Analysis Paper • 2602.13346 • Published Feb 12 • 2
Steer2Edit: From Activation Steering to Component-Level Editing Paper • 2602.09870 • Published Feb 10 • 1
scPilot: Large Language Model Reasoning Toward Automated Single-Cell Analysis and Discovery Paper • 2602.11609 • Published Feb 12 • 2
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors Paper • 2602.08934 • Published Feb 9
FIRE-Bench: Evaluating Agents on the Rediscovery of Scientific Insights Paper • 2602.02905 • Published Feb 2 • 5