← Atelier
Lesson 02 · Reinforcement learning for LLMs · ~40 min + code

From a reward to a gradient: REINFORCE, PPO, GRPO

A model answers What is 17 × 24? A checker says right or wrong. Nobody hands the model the correct answer to imitate; it has to find it by trying. How do you turn a score into a gradient?

This lesson builds that step by step. Each step fixes a concrete failure of the one before, and the result is the GRPO objective, term by term. Lesson 1's softmax gradient and KL are the only prerequisites.

1 · Reward → gradient

No label, only a score

Keep it small. The model's policy \(\pi\) is a softmax over six candidate answers to the prompt, exactly like lesson 1's four words. A grader gives partial credit for being close to 408:

With supervised fine-tuning you would raise \(\log \pi(\text{408})\) because a human wrote 408. Here nobody did. You sample an answer \(y\) from your own policy, and the grader returns \(r(y)\).

Which update increases the expected reward \(J = \sum_y \pi(y)\, r(y)\)?

The grader is never differentiated. Only the policy is:

\[ \nabla J = \sum_y r(y)\, \nabla \pi(y) = \sum_y \pi(y)\, r(y)\, \nabla \log \pi(y) = \mathbb{E}_{y\sim\pi}\big[\, r(y)\, \nabla \log \pi(y) \,\big] \]

The middle step is the log-derivative trick, \(\nabla \pi = \pi\, \nabla\log\pi\). It turns a sum over all answers into an average over samples, so one sample gives an unbiased estimate: \(r(y)\nabla\log\pi(y)\). This is REINFORCE (Williams, 1992).

On the logits you already know \(\nabla_z \log\pi(y) = e_y - \pi\) from lesson 1 (it's minus the cross-entropy gradient). So one REINFORCE step is

\[ z \leftarrow z + \eta\; r(y)\,\big(e_y - \pi\big) \]

It's SFT on your own sample, weighted by its reward. Try a few steps by hand:

One sample, one update (η = 1)

Bars show π, the probability of each answer. Green is the best answer.

2 · Baselines

Why subtract anything?

Every answer here scores between 0.4 and 1.0, so every reward is positive.

The policy samples 380, the worst answer (r = 0.4). What does the REINFORCE step do to π(380)?

It raises it. \(r = 0.4 > 0\), so the step is \(+0.4\,(e_y - \pi)\), which pushes the sampled answer up. On average the updates still favor better answers, because they get bigger pushes. But each single step pushes up whatever it happened to draw, and early luck compounds: the more likely an answer becomes, the more often it's sampled and pushed again.

Subtract a baseline \(b\), for instance the average reward in the batch, and use the advantage \(A = r - b\) instead of \(r\). An answer that beats the batch average is pushed up, and one that falls below it is pushed down. Here are 30 training runs of each version, 4 samples per step:

Expected reward during training: 30 runs each, same learning rate

Each thin line is one run; the thick line is their mean. "Ended on 408" counts runs with π(408) > 0.9 after 150 steps.

Subtracting \(b\) changes every single update. Why is that allowed?

For any \(b\) that doesn't depend on which answer was sampled:

\[ \mathbb{E}_{y\sim\pi}\big[\, b\,\nabla\log\pi(y) \,\big] = b \sum_y \nabla \pi(y) = b\, \nabla \!\sum_y \pi(y) = b\,\nabla 1 = 0 \]

The expected gradient is unchanged, and its variance can drop a lot. The best constant baseline is close to the expected reward. PPO learns it with a value network \(V(s)\). GRPO, next, gets it for free from a group of samples.

3 · Group-relative advantages

GRPO: let the group be the baseline

PPO for LLMs trains a second network, a value model as large as the policy, just to predict expected reward. GRPO (Shao et al., DeepSeekMath, 2024) removes it. For each prompt it samples a group of \(G\) completions, scores them, and compares each one to its own group:

\[ \hat A_i = \frac{r_i - \operatorname{mean}(r_1,\dots,r_G)}{\operatorname{std}(r_1,\dots,r_G)} \]

From here on the checker is binary: 1 if the answer is 408, 0 otherwise, the usual setup for math with a verifier. Before any formula work, build the instinct.

Game: predict the push

For each completion in the group, predict whether the update makes it more likely (↑), less likely (↓), or leaves it alone (0). Then check.

4 · Ratios and clipping

Reusing samples without trusting them too far

Sampling completions from an LLM is the expensive part, so PPO and GRPO take several gradient steps on each batch. After the first step, the policy being trained, \(\pi_\theta\), is no longer the policy that produced the samples, \(\pi_{\text{old}}\). The importance ratio corrects for that:

\[ \rho = \frac{\pi_\theta(y)}{\pi_{\text{old}}(y)}, \qquad \mathbb{E}_{y\sim\pi_{\text{old}}}\big[\rho\, A\big] = \mathbb{E}_{y\sim\pi_\theta}\big[A\big] \]

But a batch drawn from \(\pi_{\text{old}}\) only says something reliable about policies near \(\pi_{\text{old}}\).

A good completion (A > 0) is already 30% more likely than when it was sampled: ρ = 1.3. With ε = 0.2, should the next step keep pushing it up?

Stop, but don't reverse. PPO's clipped objective takes the minimum of the plain term and a clipped one:

\[ L(\rho) = \min\!\big(\rho A,\; \operatorname{clip}(\rho,\, 1-\varepsilon,\, 1+\varepsilon)\, A\big) \]

Beyond \(1+\varepsilon\) the clipped term is constant, the min picks it, and its gradient is zero. The completion stays where it is, and nothing pulls it back.

The clipped objective for one completion (ε = 0.2)

Drag ρ. The slope of the curve at the dot is the gradient the completion receives. Flat means clipped: no gradient.

Flip to the bad completion and look at where the flat part is. Clipping only removes the incentive to move further in the direction the advantage wants. If a bad completion became more likely (ρ > 1), nothing is clipped, and the full gradient pushes it back down. Now play it with ratios:

Game: predict the push, second step on the same batch
5 · The full objective

Tokens, length, and the KL leash

A completion is a sequence of tokens, and the ratio is taken per token. Here is GRPO as written in DeepSeekMath:

\[ \begin{aligned} J(\theta) = \mathbb{E}\Bigg[ \frac{1}{G}\sum_{i=1}^{G} \frac{1}{|o_i|} \sum_{t=1}^{|o_i|} \Big\{ &\min\!\big(\rho_{i,t}\hat A_{i},\ \operatorname{clip}(\rho_{i,t}, 1\!-\!\varepsilon, 1\!+\!\varepsilon)\,\hat A_{i}\big) \\ &- \beta\, \mathrm{KL}\big[\pi_\theta \,\|\, \pi_{\text{ref}}\big] \Big\} \Bigg] \end{aligned} \]

You've built everything inside the braces. Three details are left, and each one is a common interview follow-up.

A 5-token completion is wrong because of token 3. Which tokens does this update push down?

All five get the same \(\hat A_i\), because the reward describes the whole sequence. GRPO has no per-token credit assignment. PPO with a learned value function can give each token its own advantage, but that needs the critic GRPO removed. Over many samples the correct tokens also show up in good completions, so their pushes roughly cancel. That cancellation is slow and noisy.

Two wrong completions, both with \(\hat A = -1\): one is 10 tokens long, the other 100. With the \(1/|o_i|\) average, how hard is each token of the long one pushed down, compared with the short one?

10× less. Each token's weight is \(\hat A_i / |o_i|\). Long wrong answers are punished less per token, and long right answers are rewarded less per token. The net pressure favors making wrong answers longer. Liu et al. (2025, "Dr. GRPO") identify this length bias and remove the \(1/|o_i|\). DAPO (Yu et al., 2025) averages over all tokens in the batch instead. Try it:

Per-token push in one group of four

Click a reward to flip it. The last two columns are the weight each token of that completion gets in the loss, ×1000. GRPO averages within each sequence: \(\hat A_i / (G\,|o_i|)\). DAPO averages over all tokens in the group: \(\hat A_i / \sum_j |o_j|\).

The KL term is lesson 1's reverse KL, \(\mathrm{KL}(\pi_\theta\|\pi_{\text{ref}})\), where \(\pi_{\text{ref}}\) is the model before RL. It's estimated per token on the sampled completions with \(\frac{\pi_{\text{ref}}}{\pi_\theta} - \ln\frac{\pi_{\text{ref}}}{\pi_\theta} - 1\) and subtracted with weight \(\beta\). That keeps the policy from wandering where the reference model puts little mass, a partial brake on reward hacking. Unlike InstructGPT-style PPO, where the KL penalty is added to the reward, GRPO keeps it as a separate loss term. DAPO removes it entirely. Name your variant.

6 · New cases, no hints

Apply it somewhere else

You multiply every reward by 10. What changes in GRPO's update? In plain REINFORCE?

Standardizing by the group's mean and std makes \(\hat A\) invariant to shifting or scaling rewards. REINFORCE's step is proportional to \(r\). This is also why GRPO needs no reward normalization and why its effective step size changes from prompt to prompt: a group whose rewards barely differ still gets advantages of order 1.

Your model already solves a prompt in all 8 of 8 samples. What does GRPO learn from that prompt?

Every \(r_i\) equals the mean, so every \(\hat A_i = 0\). All-wrong groups are the same. Those prompts cost a full generation and give no gradient. As training progresses, more prompts become all-correct and the effective batch shrinks. DAPO's "dynamic sampling" keeps sampling until the batch is full of groups with mixed outcomes. A curriculum of prompts at the edge of the model's ability does the same job.

A learned reward model gives a bonus to answers that sound confident. After a while, reward keeps climbing. What is the first thing to check?

Reward that keeps rising while held-out accuracy stalls is the signature of reward hacking. The policy found a way to please the grader that doesn't solve the task. Also watch the KL to the reference, the output length, and actual samples. Fixes: repair or ensemble the reward model, raise \(\beta\), cap training steps, or use a verifiable reward where possible.

A group of G = 2 with binary rewards. What advantages can occur (population std)?

With one right and one wrong, the mean is 0.5 and the std 0.5, so \(\hat A = (+1, -1)\). With equal rewards, both are 0. Every informative group is a pure pairwise comparison: raise the winner, lower the loser. That's close in spirit to preference learning, except the pairs come from the current policy (online), not a fixed dataset as in DPO.

7 · Implement, unaided

Write the pieces of the GRPO loss

Plain editor, no completion, hidden tests. State shapes and edge cases out loud first. ⌘/Ctrl+Enter runs.

Exercise 1: group-relative advantages

Exercise 2: the clipped surrogate and its gradient

Exercise 3: the per-token KL estimator

8 · Staff-level follow-ups

Answer out loud, then check

About a minute each. Short answer first, then the reason.

What does GRPO remove from PPO, and what does that cost?
  • It removes the value network: no critic to train, and no second policy-sized model in memory.
  • Costs: G samples per prompt, a baseline that's noisy when G is small, no per-token credit assignment, and zero signal from all-correct or all-wrong groups.
  • It fits verifiable, sequence-level rewards (math, code) where a good critic is hard to learn anyway.
Why is a baseline allowed, and what is the best one?
  • \(\mathbb{E}[b\,\nabla\log\pi] = b\,\nabla\sum\pi = 0\) for any \(b\) that doesn't depend on the sampled action, so the gradient stays unbiased.
  • Variance is roughly minimized by \(b \approx \mathbb{E}[r]\); the state value \(V(s)\) is the standard choice, which gives advantage \(A = Q - V\).
  • GRPO's group mean is a Monte Carlo estimate of \(V\) for that prompt. Including \(r_i\) in its own mean only rescales the expected gradient by \((G-1)/G\). Dividing by the group std, which also depends on \(r_i\), is what makes GRPO not exactly an unbiased policy gradient. RLOO uses a leave-one-out mean and no std.
Walk through PPO's clipped objective for A > 0 and A < 0.
  • A > 0: gradient flows while ρ < 1+ε, then the objective is flat, so there's no incentive to increase the probability further.
  • A < 0: gradient flows while ρ > 1−ε, and it's flat below that. A bad action that became more likely is never clipped.
  • It's a cheap, first-order stand-in for TRPO's KL trust region. It limits incentive, not the step itself, so monitor KL(π_old‖π) and clip fraction, and stop epochs early if they spike.
Where does GRPO's length bias come from, and how do you fix it?
  • The per-sequence mean \(1/|o_i|\) gives each token weight \(\hat A_i/|o_i|\), so long wrong answers are penalized less per token. Response length drifts upward, especially for incorrect outputs.
  • Fixes: token-level averaging across the batch (DAPO), or dropping the length normalization (Dr. GRPO). Also check overlong truncation and any length terms in the reward.
  • Diagnose with length curves split by correct vs incorrect.
DPO vs online RL with a reward (PPO/GRPO): what's actually different?
  • DPO optimizes the same KL-regularized objective using its closed-form optimum \(\pi^* \propto \pi_{\text{ref}}\,e^{r/\beta}\). This turns preference pairs into a classification loss on log-ratios, with no sampling and no reward model.
  • It's offline: it learns only from pairs in the dataset. Online RL samples from the current policy, so it can explore and correct its own new mistakes. It also costs generation, and it can hack the reward.
  • Choose by data: fixed preference pairs favor DPO; a checkable reward and compute favor online RL.
How would you show that an RL run actually improved the model?
  • Held-out tasks not used for training, checked for contamination. Report pass@1 and pass@k with confidence intervals over seeds, not just reward.
  • Baselines: the SFT model, best-of-n sampling from it (does RL beat reranking?), and an ablation per change.
  • Watch for regressions: KL, length, format hacking, and general capabilities on other benchmarks. Read samples.

Sources: R. J. Williams, Simple statistical gradient-following algorithms (1992); Schulman et al., Proximal Policy Optimization Algorithms (2017); Shao et al., DeepSeekMath (2024); Liu et al., Understanding R1-Zero-Like Training (2025); Yu et al., DAPO (2025); Rafailov et al., Direct Preference Optimization (2023).