Call Neighbours Yourself: Graph Walks with Destination-Conditioned On-Policy Self-Distillation Paper • 2608.29588 • Published 7 days ago
DeepAgent: A General Reasoning Agent with Scalable Toolsets Paper • 2510.21618 • Published Oct 24, 2025 • 103
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping Paper • 2510.18927 • Published Oct 21, 2025 • 86