-
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Paper • 2608.26005 • Published • 163 -
WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation
Paper • 2608.24479 • Published • 134 -
Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning
Paper • 2608.23318 • Published • 23 -
D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation
Paper • 2608.24987 • Published • 23
~
guihuor
AI & ML interests
多智能体强化学习
Recent Activity
updated a collection about 15 hours ago
rl updated a collection about 15 hours ago
rl updated a collection about 15 hours ago
rlOrganizations
None yet