Post
55
Something I really like when I study a subject is understanding its history, how it reached the point where it is today
I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words
This is the companion piece to Class 3 of our Training Agents series with @burtenshaw . The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale
https://huggingface.co/blog/sergiopaniego/agentic-rl-2026
I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words
This is the companion piece to Class 3 of our Training Agents series with @burtenshaw . The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale
https://huggingface.co/blog/sergiopaniego/agentic-rl-2026