Econometrics bridge: Maximum likelihood, low-rank factorization, mixture models, model complexity
Estimated time: 120 min
Lab: Open browser lab
Code: Python - R

Why this should feel familiar

The transformer explains an architecture. A foundation model additionally depends on a training regime: a broad self-supervised objective, large and diverse data, substantial model capacity, and transfer to many downstream tasks.

The core language-model objective remains conditional maximum likelihood:

\[ \mathcal L(\theta)=-\sum_t\log p_\theta(x_t\mid x_{The important change is scale and reuse. One model is trained on a broad distribution and then adapted through context, parameter updates, or system components.

Four different adaptation mechanisms

1. Pretraining changes the weights

Pretraining estimates a large parameter vector \(\theta\) from a broad self-supervised corpus. The learned representations become a reusable prior for later tasks.

2. In-context learning changes the input, not the weights

Given demonstrations \(D\) inside the prompt,

\[ p_\theta(y\mid x,D) \]

can differ substantially from \(p_\theta(y\mid x)\) even though \(\theta\) is unchanged. This is inference-time conditioning, not gradient-based fine-tuning.

3. Fine-tuning changes some or all weights

Supervised fine-tuning continues optimization on a narrower dataset. Full fine-tuning updates the base parameters. Parameter-efficient methods restrict the trainable update.

4. LoRA constrains the update to low rank

For a pretrained matrix \(W\), LoRA represents an update as

\[ W'=W+\Delta W, \qquad \Delta W=BA, \]

where \(A\in\mathbb R^{r\times d_{in}}\) and \(B\in\mathbb R^{d_{out}\times r}\) with small rank \(r\).

Instead of learning \(d_{out}d_{in}\) new entries freely, the trainable update uses

\[ r(d_{in}+d_{out}) \]

parameters. This creates a direct bridge to reduced-rank approximation and factor models.

Scaling is more than parameter count

Modern training decisions involve model size, data volume, training compute, context length, and data quality. Empirical scaling laws study how loss changes as these resources increase. The durable lesson is not one fixed exponent; it is that capacity and data must be considered jointly under a compute budget.

Inference also has separate resource dimensions: parameter memory, active FLOPs, key-value cache, memory bandwidth, and communication.

Distillation and mixture of experts

Knowledge distillation trains a smaller student to match information from a larger teacher, often using soft target distributions rather than only hard labels.

Mixture-of-experts scaling takes a different route. A router sends each token to a subset of expert networks. For expert outputs \(f_k(x)\) and routing weights \(p_k(x)\),

\[ y(x)=\sum_k p_k(x)f_k(x), \]

with sparse systems evaluating only the top few experts. This increases total parameter capacity without activating every parameter for every token.

The mixture-model analogy is useful, but a sparse MoE router also controls computation. Load balance and communication therefore become part of the modeling problem.

Econometrician’s checkpoint

Do not collapse these operations:

Mechanism What changes?
prompting / in-context learning input context
full fine-tuning many or all model weights
LoRA a low-rank parameter update
distillation a new student model is trained
MoE routing which parameter blocks are active for a token

A better prompt is not a weight update. LoRA is not merely “training fewer layers.” Sparse MoE is not the same as pruning. All may reduce cost or specialize behavior, but through different mechanisms.

Interactive browser lab

Change model dimension, LoRA rank, number of experts, and active experts per token. The lab compares full-matrix parameter count with LoRA trainable parameters and compares total MoE expert capacity with the share activated per token.

A separate toggle labels whether an example change is an inference-time context change or a parameter update.

Python and R lab

Construct a small low-rank update \(BA\), apply it to a base matrix, and report the full-update versus LoRA trainable parameter counts. Then compute total and active expert parameter counts for a toy sparse MoE layer.

Practice:

  1. Explain why in-context learning can change predictions without changing model weights.
  2. Compute LoRA trainable parameters for a square \(d\times d\) matrix and rank \(r\).
  3. Distinguish total MoE parameters from active parameters per token.
  4. Explain why more pretraining compute does not remove the need for downstream evaluation.

Learner output

Distinguish pretraining, in-context learning, full fine-tuning, LoRA, and sparse MoE by stating which parameters or activations change in each case.