Econometrics bridge: Likelihood scores, pairwise choice models, control variates, constrained optimization
Estimated time: 135 min
Lab: Open browser lab
Code: Python - R

Why this should feel familiar

Post-training mixes several statistical problems that are easy to collapse into one phrase such as “RLHF.” The course should keep them separate:

  1. supervised estimation from demonstrations;
  2. preference or reward estimation;
  3. value or baseline estimation for variance reduction;
  4. policy optimization;
  5. direct preference objectives that avoid a separate online RL loop.

Stage 1: supervised fine-tuning

Given prompt-response pairs \((x,y)\), SFT usually continues token-level likelihood training:

\[ \max_\theta \sum_{(x,y)}\log \pi_\theta(y\mid x). \]

The output remains a language model or policy. SFT changes its conditional distribution so desired instruction-response patterns receive higher probability.

Stage 2: preference and reward estimation

Suppose annotators prefer \(y^+\) to \(y^-\) for prompt \(x\). A common pairwise reward model uses

\[ P(y^+\succ y^-\mid x) =\sigma\left(r_\phi(x,y^+)-r_\phi(x,y^-)\right). \]

The reward model estimates which complete outputs are preferred. It is not the same as the policy, and its scalar output is not automatically a calibrated measure of truth or social value.

Value model versus reward model

A reward model scores an outcome or transition according to the learned training signal.

A value model estimates expected future return from a state or partial trajectory:

\[ V_\psi(s_t)\approx \mathbb E[G_t\mid s_t]. \]

In policy-gradient training, the value estimate acts as a baseline. The advantage

\[ \hat A_t=\hat G_t-V_\psi(s_t) \]

centers the policy-gradient signal and can reduce variance.

This is different from a reference model, which is a policy used to measure or penalize drift from a baseline distribution.

PPO

For policy \(\pi_\theta\) and old policy \(\pi_{old}\), define

\[ r_t(\theta)=\frac{\pi_\theta(a_t\mid s_t)}{\pi_{old}(a_t\mid s_t)}. \]

The clipped PPO surrogate is

\[ L^{CLIP}=\mathbb E_t\left[ \min\left(r_t\hat A_t, \operatorname{clip}(r_t,1-\epsilon,1+\epsilon)\hat A_t\right) \right]. \]

Clipping discourages locally destructive probability-ratio changes. It is not a guarantee that the full policy lies inside a hard global trust region.

RLHF-style objectives often also include a KL penalty to a reference policy:

\[ \mathbb E[r_\phi(x,y)]-\beta D_{KL}(\pi_\theta\|\pi_{ref}). \]

The value baseline and reference policy therefore serve different purposes.

DPO

DPO trains directly from preferred and rejected responses. In simplified form it increases the preference margin of the current policy relative to a reference policy. A typical pairwise term depends on

[ \beta\left[ \log\frac{\pi_\theta(y^+\mid x)}{\pi_{ref}(y^+\mid x)}

\log\frac{\pi_\theta(y^-\mid x)}{\pi_{ref}(y^-\mid x)} \right]. ]

DPO therefore uses preference pairs without separately fitting a reward model and then running an online PPO loop. It is more precise to call it direct preference optimization than “SFT on preferred/rejected pairs,” because the objective depends on a preference contrast and a reference policy, not ordinary likelihood of only the preferred text.

GRPO

GRPO is a PPO-related approach that can avoid a separately trained value model by comparing rewards within a group of sampled responses. For rewards \(R_1,\ldots,R_G\), a simplified standardized group advantage is

\[ A_i=\frac{R_i-\bar R}{s_R+\epsilon}. \]

The policy is then updated using relative group performance. The key change is the baseline construction: group-relative rewards replace a learned state-value critic in the advantage calculation. The policy still changes through gradient optimization.

When does policy change?

A policy does not update only when a reward is higher than the previous reward. Gradient-based methods use the signed training signal across sampled actions. A negative or below-baseline advantage can decrease the probability of an action; a positive advantage can increase it. The optimizer aggregates these contributions across a batch.

Econometrician’s checkpoint

Keep this object map visible:

Object Statistical role
policy / language model distribution over outputs or actions
reward model learned scalar preference signal
value model expected future return / variance-reduction baseline
reference model anchor distribution for policy drift
advantage relative signal used in policy-gradient updates
preference pair observed comparison used by reward or direct-preference methods

Reward optimization can amplify misspecification. Preference data can be narrow or inconsistent. Reference-policy constraints can preserve both useful behavior and defects. These are statistical design questions, not implementation footnotes.

Interactive browser lab

Use three linked calculations:

  • PPO: vary old probability, new probability, advantage, and clipping epsilon;
  • DPO: vary the preferred-versus-rejected log-probability margin and beta;
  • GRPO: change one reward in a fixed group and inspect the group-centered standardized advantages.

The lab prints every intermediate quantity so the different baselines and ratios remain visible.

Python and R lab

Reproduce the PPO clipped contribution, a simplified DPO preference probability, and GRPO standardized group advantages. Include both positive and negative advantages.

Practice:

  1. Explain why a reward model and value model are not interchangeable.
  2. Explain why comparing the policy with a reference model does not replace the need for a value baseline in PPO.
  3. State precisely what DPO removes from the conventional reward-model-plus-PPO pipeline.
  4. Explain what GRPO changes relative to PPO’s learned critic/value baseline.
  5. Explain why policy parameters can change after a below-baseline reward.

Learner output

Explain the distinct roles of the reward model, value baseline, and reference policy, then compare PPO, DPO, and GRPO at the objective level.