MMR-GRPO-lambda-0.9 / reward_data
318 MB
kangdawei's picture
Training in progress, step 500
4b7cf2d verified