mradermacher/Hemlock-Apothecary-7B-GRPO-e3-GGUF Reinforcement Learning • 8B • Updated 16 days ago • 592