The SFT+RL+OPD model weights of paper "To mix or to merge: Toward multi-domain reinforcement learning for large language models"