mradermacher/Qwen3-8B-MiMo-Music-GRPO-GGUF Reinforcement Learning • 8B • Updated about 19 hours ago • 245