Econometrics bridge: Likelihood scores, pairwise choice models, control variates, constrained optimization
Estimated time: 135 min
Lab: Open browser lab
Code: Python - R
Why this should feel familiar
Post-training mixes several statistical problems that are easy to collapse into one phrase such as “RLHF.” The course should keep them separate:
- supervised estimation from demonstrations;
- preference or reward estimation;
- value or baseline estimation for variance reduction;
- policy optimization;
- direct preference objectives that avoid a separate online RL loop.
Stage 1: supervised fine-tuning
Given prompt-response pairs \((x,y)\), SFT usually continues token-level likelihood training:
\[ \max_\theta \sum_{(x,y)}\log \pi_\theta(y\mid x). \]The output remains a language model or policy. SFT changes its conditional distribution so desired instruction-response patterns receive higher probability.
Stage 2: preference and reward estimation
Suppose annotators prefer \(y^+\) to \(y^-\) for prompt \(x\). A common pairwise reward model uses
\[ P(y^+\succ y^-\mid x) =\sigma\left(r_\phi(x,y^+)-r_\phi(x,y^-)\right). \]The reward model estimates which complete outputs are preferred. It is not the same as the policy, and its scalar output is not automatically a calibrated measure of truth or social value.
Value model versus reward model
A reward model scores an outcome or transition according to the learned training signal.
A value model estimates expected future return from a state or partial trajectory:
\[ V_\psi(s_t)\approx \mathbb E[G_t\mid s_t]. \]In policy-gradient training, the value estimate acts as a baseline. The advantage
\[ \hat A_t=\hat G_t-V_\psi(s_t) \]centers the policy-gradient signal and can reduce variance.
This is different from a reference model, which is a policy used to measure or penalize drift from a baseline distribution.
PPO
For policy \(\pi_\theta\) and old policy \(\pi_{old}\), define
\[ r_t(\theta)=\frac{\pi_\theta(a_t\mid s_t)}{\pi_{old}(a_t\mid s_t)}. \]The clipped PPO surrogate is
\[ L^{CLIP}=\mathbb E_t\left[ \min\left(r_t\hat A_t, \operatorname{clip}(r_t,1-\epsilon,1+\epsilon)\hat A_t\right) \right]. \]Clipping discourages locally destructive probability-ratio changes. It is not a guarantee that the full policy lies inside a hard global trust region.
RLHF-style objectives often also include a KL penalty to a reference policy:
\[ \mathbb E[r_\phi(x,y)]-\beta D_{KL}(\pi_\theta\|\pi_{ref}). \]The value baseline and reference policy therefore serve different purposes.
DPO
DPO trains directly from preferred and rejected responses. In simplified form it increases the preference margin of the current policy relative to a reference policy. A typical pairwise term depends on
[ \beta\left[ \log\frac{\pi_\theta(y^+\mid x)}{\pi_{ref}(y^+\mid x)}
\log\frac{\pi_\theta(y^-\mid x)}{\pi_{ref}(y^-\mid x)} \right]. ]
DPO therefore uses preference pairs without separately fitting a reward model and then running an online PPO loop. It is more precise to call it direct preference optimization than “SFT on preferred/rejected pairs,” because the objective depends on a preference contrast and a reference policy, not ordinary likelihood of only the preferred text.
GRPO
GRPO is a PPO-related approach that can avoid a separately trained value model by comparing rewards within a group of sampled responses. For rewards \(R_1,\ldots,R_G\), a simplified standardized group advantage is
\[ A_i=\frac{R_i-\bar R}{s_R+\epsilon}. \]The policy is then updated using relative group performance. The key change is the baseline construction: group-relative rewards replace a learned state-value critic in the advantage calculation. The policy still changes through gradient optimization.
When does policy change?
A policy does not update only when a reward is higher than the previous reward. Gradient-based methods use the signed training signal across sampled actions. A negative or below-baseline advantage can decrease the probability of an action; a positive advantage can increase it. The optimizer aggregates these contributions across a batch.
Econometrician’s checkpoint
Keep this object map visible:
| Object | Statistical role |
|---|---|
| policy / language model | distribution over outputs or actions |
| reward model | learned scalar preference signal |
| value model | expected future return / variance-reduction baseline |
| reference model | anchor distribution for policy drift |
| advantage | relative signal used in policy-gradient updates |
| preference pair | observed comparison used by reward or direct-preference methods |
Reward optimization can amplify misspecification. Preference data can be narrow or inconsistent. Reference-policy constraints can preserve both useful behavior and defects. These are statistical design questions, not implementation footnotes.
Interactive browser lab
Use three linked calculations:
- PPO: vary old probability, new probability, advantage, and clipping epsilon;
- DPO: vary the preferred-versus-rejected log-probability margin and beta;
- GRPO: change one reward in a fixed group and inspect the group-centered standardized advantages.
The lab prints every intermediate quantity so the different baselines and ratios remain visible.
Python and R lab
Reproduce the PPO clipped contribution, a simplified DPO preference probability, and GRPO standardized group advantages. Include both positive and negative advantages.
Practice:
- Explain why a reward model and value model are not interchangeable.
- Explain why comparing the policy with a reference model does not replace the need for a value baseline in PPO.
- State precisely what DPO removes from the conventional reward-model-plus-PPO pipeline.
- Explain what GRPO changes relative to PPO’s learned critic/value baseline.
- Explain why policy parameters can change after a below-baseline reward.
Learner output
Explain the distinct roles of the reward model, value baseline, and reference policy, then compare PPO, DPO, and GRPO at the objective level.