andersonbcdefg/sharegpt_reward_modeling_pairwise_no_as_an_ai Viewer • Updated Jun 6, 2023 • 11.8k • 41 • 2
Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models Paper • 2609.13680 • Published 7 days ago • 8
ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents Paper • 2609.17523 • Published 4 days ago • 25
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 12 days ago • 329
andersonbcdefg/red_teaming_reward_modeling_pairwise_no_as_an_ai Viewer • Updated Jun 1, 2023 • 35.3k • 27 • 3
jaehyeokdoo2/openpi-droid-pnpcarrot-singetask-qflow-offlinerl-criticwarmup2000-alpha100-bs8-test Updated Mar 4 • 1
MInTRL: Off-policy Intervention can boost On-policy RL Paper • 2609.12419 • Published 8 days ago • 11
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training Paper • 2609.15051 • Published 5 days ago • 12
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction Paper • 2609.13285 • Published 11 days ago • 69
MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control Paper • 2609.06251 • Published 14 days ago • 3
DataFlex-RL: An Evaluation Platform for RLVR Data Policies Paper • 2609.06107 • Published 14 days ago • 80