philomath-1209/gpt2-reward_model_hh-rlhf Reinforcement Learning • 0.1B • Updated about 7 hours ago • 41