Buckets:
| title: KL Regularization (pointer) | |
| maturity: stub | |
| sources: | |
| - arxiv:1909.08593 | |
| - arxiv:2203.02155 | |
| # KL Regularization | |
| The **reference-model KL penalty** — penalizing divergence from a frozen reference | |
| policy (usually the SFT model) — is the most universal regularizer in RL-based LLM | |
| post-training: it keeps the policy in the region where the reward is trustworthy, | |
| preserves generation diversity, and is the front-line control against reward | |
| over-optimization. It was introduced for language models by Ziegler et al. as | |
| $R = r - \beta\,\mathbb{D}_{\mathrm{KL}}(\pi\,\|\,\pi_{\text{ref}})$ | |
| [source:arxiv:1909.08593] and carried into InstructGPT with $\beta=0.02$ | |
| [source:arxiv:2203.02155]. | |
| > **This topic is treated comprehensively at | |
| > `objectives-and-regularization/reference-model-and-kl`.** See there for the | |
| > KL-control derivation and the closed-form Boltzmann optimum, the three jobs of the | |
| > penalty (anti-over-optimization, diversity/entropy, task definition), fixed-vs-adaptive | |
| > $\beta$ across recipes, KL-in-reward vs KL-in-loss placement, the **two distinct KLs** | |
| > (reference regularizer vs PPO/TRPO's step-size KL), the KL-vs-alignment-tax tradeoff, | |
| > and reference-free variants. | |
| This page is a deliberate pointer: the `foundations/kl-regularization` and | |
| `objectives-and-regularization/reference-model-and-kl` taxonomy nodes were near-synonymous, | |
| so the canonical treatment lives at the latter to keep one source of truth (a `meta:` | |
| taxonomy note will alias this node). | |
| See also: `foundations/policy-gradient-methods`, | |
| `reward-modeling/reward-model-overoptimization`. | |
Xet Storage Details
- Size:
- 1.6 kB
- Xet hash:
- c706cb63c9d2b082355b59f12629a061637470a501a49c0cdcda6620ff4e5d86
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.