SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Paper • 2608.03092 • Published 7 days ago • 10
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Paper • 2608.03092 • Published 7 days ago • 10 • 1
DanhVuiVe/ChartQA_Benetech_PlotQa_DVQA_combined_matcha_complete Viewer • Updated Oct 29, 2024 • 535k • 112 • 2
alimama-creative/FLUX.1-dev-Controlnet-Inpainting-Beta Image-to-Image • 2B • Updated Oct 12, 2024 • 4.12k • 427
openai/whisper-large-v3-turbo Automatic Speech Recognition • 0.8B • Updated Oct 4, 2024 • 7.62M • • 3.23k