view article Article DeepSeek-R1 Dissection: Understanding PPO & GRPO Without Any Prior Reinforcement Learning Knowledge NormalUhr • Feb 7, 2025 • 302
view article Article Welcome Llama 3 - Meta's new open LLM +3 philschmid, osanseviero, pcuenq, ybelkada, lvwerra • Apr 18, 2024 • 295