Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs Paper • 2610.01428 • Published 9 days ago • 12
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 23 days ago • 115
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published Sep 7 • 376
ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation Paper • 2609.09076 • Published Sep 8 • 24
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published Sep 8 • 83
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published Sep 3 • 86
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review Paper • 2608.08975 • Published Aug 10 • 48
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 266
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI Paper • 2608.10915 • Published Aug 11 • 195
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation Paper • 2608.05042 • Published Aug 5 • 9
GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning Paper • 2608.02585 • Published Aug 3 • 25
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Paper • 2607.23782 • Published Jul 26 • 80