Buckets:
| title: Policy-Gradient Methods for LLM Post-Training | |
| maturity: developing | |
| sources: | |
| - arxiv:1502.05477 | |
| - arxiv:1506.02438 | |
| - arxiv:1707.06347 | |
| - arxiv:2203.02155 | |
| open_questions: | |
| - "How much of the classical policy-gradient machinery (a learned value function, GAE, the old-vs-new-policy trust region) is actually load-bearing for LLM post-training, where episodes are short and rewards are terminal — versus inherited by convention?" | |
| - "Is the on-policy actor-critic stack (PPO+GAE) being displaced by critic-free group-relative methods for reasoning RL, or do the two coexist by regime? This needs a corpus-wide survey of recent recipes to answer (GRPO/R1 sources not yet processed)." | |
| # Policy-Gradient Methods for LLM Post-Training | |
| Policy-gradient (PG) methods are the family of reinforcement-learning algorithms | |
| that optimize a *parameterized stochastic policy* directly, by following a noisy | |
| estimate of the gradient of expected reward. They are the algorithmic backbone of | |
| RL-based LLM post-training: the dominant RLHF recipe optimizes the language model | |
| with **Proximal Policy Optimization (PPO)** [source:arxiv:1707.06347], a PG method, | |
| using **Generalized Advantage Estimation (GAE)** [source:arxiv:1506.02438] for the | |
| advantage signal, all popularized for language models by InstructGPT | |
| [source:arxiv:2203.02155]. This article covers the mechanism shared by the whole | |
| family — from the bare score-function estimator, through the variance-reduction and | |
| step-control machinery that made it work on neural networks, to how that machinery | |
| is adapted (and partly degenerates) when the "environment" is text generation. | |
| ## 1. The policy-gradient estimator | |
| All PG methods maximize the expected return $\eta=\mathbb{E}\!\left[\sum_t r_t\right]$ | |
| by ascending a stochastic estimate of $g:=\nabla_\theta\eta$. Every common form of | |
| the estimator shares one structure [source:arxiv:1506.02438]: | |
| $$ g=\mathbb{E}\!\left[\sum_{t=0}^{\infty}\Psi_t\,\nabla_\theta\log\pi_\theta(a_t\mid s_t)\right], $$ | |
| where $\Psi_t$ — the weight on each action's score $\nabla_\theta\log\pi_\theta(a_t\mid s_t)$ — | |
| can be the total return, the reward-to-go, a baselined reward-to-go, the state–action | |
| value $Q^\pi$, the advantage $A^\pi(s,a)=Q^\pi(s,a)-V^\pi(s)$, or the TD residual | |
| $r_t+\gamma V^\pi(s_{t+1})-V^\pi(s_t)$ [source:arxiv:1506.02438]. The bare | |
| total-return form (with no baseline) is the classic REINFORCE estimator. The whole | |
| art of practical PG is the **choice of $\Psi_t$**, because that choice governs the | |
| estimator's variance — and the variance of the naive estimator scales unfavorably | |
| with the time horizon, since an action's effect is confounded with the effects of | |
| past and future actions [source:arxiv:1506.02438]. | |
| Choosing $\Psi_t=A^\pi(s_t,a_t)$ — the **advantage** — yields almost the lowest | |
| possible variance, with a clean interpretation: a PG step should raise the | |
| probability of better-than-average actions and lower it for worse-than-average ones, | |
| and the advantage is exactly the "better or worse than this policy's default" signal | |
| [source:arxiv:1506.02438]. The catch is that $A^\pi$ is unknown and must be | |
| estimated, which is where the rest of the machinery comes from. A recurring theme, | |
| stated sharply in the GAE paper, is that **bias is more pernicious than variance**: | |
| high variance just demands more samples, whereas bias can make the algorithm fail to | |
| converge or converge to something that is not even a local optimum | |
| [source:arxiv:1506.02438]. | |
| ## 2. Variance reduction: baselines, advantage, and GAE | |
| Subtracting a state-dependent **baseline** $b(s_t)$ from the return leaves the | |
| gradient unbiased (the baseline term vanishes because | |
| $\mathbb{E}_{a_t}[\nabla_\theta\log\pi_\theta(a_t\mid s_t)]=0$) while reducing | |
| variance; using $b=V^\pi$ turns the reward-to-go into an advantage estimate | |
| [source:arxiv:1506.02438]. **GAE** generalizes this into a one-parameter family. With | |
| an approximate value function $V$ and its TD residual | |
| $\delta^V_t=r_t+\gamma V(s_{t+1})-V(s_t)$, | |
| $$ \hat A_t^{\mathrm{GAE}(\gamma,\lambda)}=\sum_{l=0}^{\infty}(\gamma\lambda)^l\,\delta^V_{t+l}, $$ | |
| an exponentially-weighted average of $k$-step advantage estimators that collapses to | |
| a $(\gamma\lambda)$-discounted sum of Bellman residuals [source:arxiv:1506.02438]. The | |
| parameter $\lambda$ interpolates between a high-bias/low-variance one-step estimate | |
| ($\lambda=0$, just $\delta^V_t$) and an unbiased/high-variance Monte-Carlo estimate | |
| ($\lambda=1$, empirical returns minus the baseline) [source:arxiv:1506.02438]. Crucially, | |
| $\gamma$ and $\lambda$ are **not interchangeable**: $\gamma$ sets the scale/horizon of | |
| the value function and introduces bias by discounting long-range credit, while | |
| $\lambda$ trades bias for variance *given* the value function and "introduces far less | |
| bias than $\gamma$ for a reasonably accurate value function" — which is why the best | |
| $\lambda$ (empirically $\in[0.9,0.99]$) is typically lower than the best $\gamma$ | |
| [source:arxiv:1506.02438]. | |
| GAE needs a value function, and fitting $V_\phi$ robustly is its own problem; the GAE | |
| paper fits it by regression to discounted returns under a **trust region** (a bound on | |
| the change in $V_\phi$, equivalent to an average-KL constraint on a Gaussian view of | |
| the value function), solved with the same conjugate-gradient machinery TRPO uses for | |
| the policy [source:arxiv:1506.02438]. This pairing — GAE advantages plus a | |
| trust-region policy update — is the actor-critic stack that the RLHF pipeline | |
| inherited. | |
| ## 3. Controlling the step: trust regions (TRPO) and clipping (PPO) | |
| The second practical problem is step size: a single overlarge PG update can collapse | |
| the policy, from which on-policy learning may never recover. **TRPO** addresses this | |
| with theory. Starting from the identity that expresses a new policy's return via the | |
| old policy's advantages, it optimizes a local surrogate $L_\pi(\tilde\pi)$ and proves | |
| a monotonic-improvement bound | |
| $\eta(\tilde\pi)\ge L_\pi(\tilde\pi)-C\,D_{\mathrm{KL}}^{\max}(\pi,\tilde\pi)$ with | |
| $C=4\epsilon\gamma/(1-\gamma)^2$ [source:arxiv:1502.05477]. Because the | |
| theory-prescribed penalty forces tiny steps, the practical algorithm instead | |
| maximizes the surrogate subject to a hard constraint on the **average** KL between | |
| new and old policies, $\bar D_{\mathrm{KL}}\le\delta$, solved with conjugate gradient | |
| on Fisher-vector products plus a backtracking line search | |
| [source:arxiv:1502.05477]. TRPO also unifies the family: natural policy gradient, | |
| vanilla PG, and policy iteration are all special/limiting cases of its constrained | |
| update [source:arxiv:1502.05477]. | |
| **PPO** keeps TRPO's goal — bounded, stable steps — but discards the second-order | |
| machinery for a *clipped surrogate* optimized by ordinary SGD | |
| [source:arxiv:1707.06347]. With the probability ratio | |
| $r_t(\theta)=\pi_\theta(a_t\mid s_t)/\pi_{\theta_{\text{old}}}(a_t\mid s_t)$, | |
| $$ L^{\mathrm{CLIP}}(\theta)=\mathbb{E}_t\!\left[\min\!\big(r_t\hat A_t,\;\operatorname{clip}(r_t,1-\epsilon,1+\epsilon)\hat A_t\big)\right], $$ | |
| whose $\min$ makes it a pessimistic lower bound on the unclipped surrogate: once the | |
| ratio moves past $1\pm\epsilon$ in the improving direction the gradient flattens, | |
| removing the incentive for destructive steps [source:arxiv:1707.06347]. This first-order | |
| form is what lets PPO safely run **several epochs of minibatch SGD per batch** of | |
| rollouts — the clip is precisely what keeps those reused updates safe as $r_t$ drifts | |
| from 1 [source:arxiv:1707.06347]. PPO also studied an adaptive KL-penalty variant but | |
| reported it performs *worse* than clipping [source:arxiv:1707.06347]. The net trade — | |
| near-TRPO stability with vastly simpler implementation — is why PPO, not TRPO, became | |
| the workhorse optimizer of RLHF [source:arxiv:1707.06347]. | |
| ## 4. The LLM adaptation: PG methods inside RLHF | |
| When the policy is a language model, the "MDP" is degenerate in a specific way: a | |
| prompt is the initial state, each generated **token is an action**, and (in the | |
| standard RLHF setup) a single scalar reward from a reward model arrives only at the | |
| end of the sequence — i.e. a **contextual bandit at the sequence level** | |
| [source:arxiv:2203.02155]. InstructGPT instantiates the PG stack as: supervised | |
| fine-tuning (SFT) → reward model (RM) → PPO, optimizing | |
| $$ \text{objective}(\phi)=\mathbb{E}_{(x,y)\sim\pi_\phi^{RL}}\!\left[r_\theta(x,y)-\beta\log\frac{\pi_\phi^{RL}(y\mid x)}{\pi^{SFT}(y\mid x)}\right]+\gamma\,\mathbb{E}_{x\sim D_{\text{pretrain}}}\!\left[\log\pi_\phi^{RL}(x)\right], $$ | |
| with a value head initialized from the RM, KL coefficient $\beta=0.02$, PPO clip | |
| $0.2$, batch size 512, a single inner epoch, and — tellingly — **no discount when | |
| estimating GAE** [source:arxiv:2203.02155]. | |
| That last detail is the key conceptual link back to Sections 2–3: because an LLM | |
| generation is a short, single-terminal-reward episode, the long-horizon | |
| credit-assignment problem GAE was built for is largely **degenerate** — with | |
| $\gamma=1$ and a terminal reward, $\lambda$ matters far less than it does in | |
| locomotion [source:arxiv:2203.02155][source:arxiv:1506.02438]. Several other | |
| adaptations distinguish LLM-PPO from the canonical control algorithm: | |
| - **Two different KLs.** TRPO/PPO use a new-vs-old-*policy* KL as a *step-size control* | |
| [source:arxiv:1502.05477][source:arxiv:1707.06347]; RLHF *additionally* adds a | |
| per-token KL penalty to a **frozen reference (SFT) policy** as a *regularizer* | |
| against reward-model over-optimization [source:arxiv:2203.02155]. These play | |
| conceptually distinct roles and should not be conflated — the RLHF penalty is closer | |
| in spirit to PPO's (dispreferred) adaptive-KL-penalty variant than to its clip. | |
| - **Few epochs, large batches.** Where the PPO paper reuses each batch for $K=3$–$10$ | |
| epochs [source:arxiv:1707.06347], InstructGPT runs a single inner epoch on very large | |
| batches [source:arxiv:2203.02155]. | |
| - **Auxiliary pretraining loss (PPO-ptx).** To counter the "alignment tax" — PPO | |
| regressing on public NLP benchmarks — InstructGPT mixes pretraining gradients into | |
| the update with coefficient $\gamma=27.8$, which recovers regressions better than | |
| simply raising the reference-KL coefficient [source:arxiv:2203.02155]. | |
| - **A small fixed critic for a large policy.** A 6B RM and 6B value function are used | |
| even for the 175B policy [source:arxiv:2203.02155]. | |
| The headline payoff of the recipe is behavioral: labelers prefer 175B InstructGPT | |
| over 175B GPT-3 about 85% of the time, and even the 1.3B InstructGPT model is | |
| preferred over 175B GPT-3 despite ~100× fewer parameters | |
| [source:arxiv:2203.02155]. | |
| ## 5. Relationships to neighboring method families | |
| PG-with-a-critic is one corner of a larger space; two neighbors matter most for | |
| orientation (each has — or will have — its own article): | |
| - **Critic-free / group-relative methods** (`algorithms/grpo-and-group-relative`): | |
| drop the learned value function entirely and estimate advantages from the reward | |
| statistics of a *group* of samples for the same prompt. This removes GAE and the | |
| value-function trust region from the stack — attractive precisely because, per | |
| Section 4, the critic's long-horizon role is weak in the terminal-reward LLM | |
| setting. *(The GRPO and DeepSeek-R1 sources are on the reading frontier but not yet | |
| processed; this pointer is intentionally light pending their capture.)* | |
| - **RL-free preference optimization** (`algorithms/dpo-and-offline-po`): skips the PG | |
| loop altogether, turning the RLHF objective into a supervised loss on preference | |
| pairs. It is the main "no-RL" baseline against which PG-based RLHF is measured. | |
| ## 6. Current status and trajectory | |
| *(Hedged, and grounded in the merged corpus; trend claims here cite their evidence | |
| base rather than a single paper, and "not-reported ≠ not-used" applies throughout.)* | |
| Within the corpus processed so far, the **PPO + GAE actor-critic stack is the | |
| reference RLHF optimizer**: it is what InstructGPT used and popularized | |
| [source:arxiv:2203.02155][source:arxiv:1707.06347], and GAE remains the default | |
| advantage estimator wherever a learned critic is in play | |
| [source:arxiv:1506.02438]. TRPO is essentially never used directly for LLMs — its | |
| role is ancestral, the trust-region idea that PPO simplified | |
| [source:arxiv:1502.05477][source:arxiv:1707.06347]. | |
| The visible trajectory is a **partial move away from the learned critic** for | |
| reasoning-oriented RL: critic-free, group-relative methods drop the value function | |
| (and thus GAE), motivated by the same observation that the critic's long-horizon | |
| machinery is largely idle when rewards are terminal. This is a *trend statement* and | |
| must be treated as such — it should be firmed up by a corpus-wide survey of recent | |
| recipes (which report a value function vs. which do not), not asserted from any single | |
| paper, and the relevant GRPO/DeepSeek-R1 sources are queued but not yet processed in | |
| this wiki. What is safe to say now: the *score-function gradient itself* (Section 1) | |
| is common to PPO and to the group-relative methods alike, so "policy-gradient methods" | |
| as a family are not fading even where one specific member (PPO-with-GAE) may be ceding | |
| ground in the reasoning regime. | |
| ## 7. Open questions | |
| - How much of the classical PG machinery (learned $V$, GAE, old-vs-new trust region) | |
| is actually load-bearing for LLM post-training versus inherited by convention, given | |
| the degenerate terminal-reward episode structure? [source:arxiv:2203.02155] | |
| - What is the right way to set/adapt $\gamma,\lambda$ (or to dispense with them) | |
| automatically — flagged as future work already in the GAE paper | |
| [source:arxiv:1506.02438]? | |
| - Does the on-policy PPO+GAE stack get displaced by critic-free methods across the | |
| board, or do they partition by regime (broad preference RLHF vs. verifiable-reward | |
| reasoning RL)? Unresolved pending more of the corpus. | |
| ## Runnable check: the baseline is unbiased and cuts variance | |
| The policy-gradient estimator is $\nabla_\theta \mathbb{E}[R] = \mathbb{E}_\pi[(R-b)\,\nabla_\theta\log\pi]$. | |
| A state-independent baseline $b$ leaves the *expected* gradient unchanged (because the score | |
| function has zero mean, $\mathbb{E}_\pi[\nabla_\theta\log\pi]=0$) while reducing its variance. | |
| This enumerates a 2-action softmax bandit exactly (no sampling) and asserts both properties: | |
| ```python | |
| import math | |
| def softmax(z): | |
| m = max(z); e = [math.exp(x - m) for x in z]; s = sum(e) | |
| return [x / s for x in e] | |
| probs = softmax([0.3, -0.2]); rewards = [1.0, 3.0] | |
| # softmax score fn wrt logit j: dlog pi(a)/dz_j = 1{a==j} - pi(j); use component j=0 | |
| j = 0 | |
| score = [(1.0 if a == j else 0.0) - probs[j] for a in (0, 1)] | |
| b = sum(p * r for p, r in zip(probs, rewards)) # baseline = E[R] | |
| g_nob = sum(probs[a] * rewards[a] * score[a] for a in (0, 1)) | |
| g_bl = sum(probs[a] * (rewards[a] - b) * score[a] for a in (0, 1)) | |
| assert abs(sum(probs[a] * score[a] for a in (0, 1))) < 1e-12 # E[score] = 0 | |
| assert abs(g_nob - g_bl) < 1e-12 # baseline: same expected gradient | |
| var_nob = sum(probs[a] * (rewards[a] * score[a]) ** 2 for a in (0, 1)) - g_nob ** 2 | |
| var_bl = sum(probs[a] * ((rewards[a] - b) * score[a]) ** 2 for a in (0, 1)) - g_bl ** 2 | |
| assert var_bl < var_nob # ...but lower variance | |
| ``` | |
| ## References | |
| - **TRPO** — Schulman et al. 2015 [source:arxiv:1502.05477]: trust-region policy | |
| update with a monotonic-improvement guarantee; the ancestor PPO simplifies. | |
| - **GAE** — Schulman et al. 2015/16 [source:arxiv:1506.02438]: the | |
| exponentially-weighted advantage estimator and the variance/bias analysis behind | |
| $\Psi_t$. | |
| - **PPO** — Schulman et al. 2017 [source:arxiv:1707.06347]: the clipped first-order | |
| surrogate that became the RLHF workhorse optimizer. | |
| - **InstructGPT** — Ouyang et al. 2022 [source:arxiv:2203.02155]: the canonical | |
| SFT→RM→PPO RLHF recipe and the source of the LLM-specific adaptations. | |
| - Forward links (articles): `algorithms/rlhf-ppo-pipeline`, | |
| `algorithms/grpo-and-group-relative`, `algorithms/dpo-and-offline-po`, | |
| `foundations/kl-regularization`. | |
Xet Storage Details
- Size:
- 16.1 kB
- Xet hash:
- 18c1c3f1f5dd9ad059cbbb69a768527814164e5badd23d790c29e32885dd1b2d
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.