Econometrics bridge: Maximum likelihood, low-rank factorization, mixture models, model complexity
Estimated time: 120 min
Lab: Open browser lab
Code: Python - R
Why this should feel familiar
The transformer explains an architecture. A foundation model additionally depends on a training regime: a broad self-supervised objective, large and diverse data, substantial model capacity, and transfer to many downstream tasks.
The core language-model objective remains conditional maximum likelihood:
\[ \mathcal L(\theta)=-\sum_t\log p_\theta(x_t\mid x_{Four different adaptation mechanisms
1. Pretraining changes the weights
Pretraining estimates a large parameter vector \(\theta\) from a broad self-supervised corpus. The learned representations become a reusable prior for later tasks.
2. In-context learning changes the input, not the weights
Given demonstrations \(D\) inside the prompt,
\[ p_\theta(y\mid x,D) \]can differ substantially from \(p_\theta(y\mid x)\) even though \(\theta\) is unchanged. This is inference-time conditioning, not gradient-based fine-tuning.
3. Fine-tuning changes some or all weights
Supervised fine-tuning continues optimization on a narrower dataset. Full fine-tuning updates the base parameters. Parameter-efficient methods restrict the trainable update.
4. LoRA constrains the update to low rank
For a pretrained matrix \(W\), LoRA represents an update as
\[ W'=W+\Delta W, \qquad \Delta W=BA, \]where \(A\in\mathbb R^{r\times d_{in}}\) and \(B\in\mathbb R^{d_{out}\times r}\) with small rank \(r\).
Instead of learning \(d_{out}d_{in}\) new entries freely, the trainable update uses
\[ r(d_{in}+d_{out}) \]parameters. This creates a direct bridge to reduced-rank approximation and factor models.
Scaling is more than parameter count
Modern training decisions involve model size, data volume, training compute, context length, and data quality. Empirical scaling laws study how loss changes as these resources increase. The durable lesson is not one fixed exponent; it is that capacity and data must be considered jointly under a compute budget.
Inference also has separate resource dimensions: parameter memory, active FLOPs, key-value cache, memory bandwidth, and communication.
Distillation and mixture of experts
Knowledge distillation trains a smaller student to match information from a larger teacher, often using soft target distributions rather than only hard labels.
Mixture-of-experts scaling takes a different route. A router sends each token to a subset of expert networks. For expert outputs \(f_k(x)\) and routing weights \(p_k(x)\),
\[ y(x)=\sum_k p_k(x)f_k(x), \]with sparse systems evaluating only the top few experts. This increases total parameter capacity without activating every parameter for every token.
The mixture-model analogy is useful, but a sparse MoE router also controls computation. Load balance and communication therefore become part of the modeling problem.
Econometrician’s checkpoint
Do not collapse these operations:
| Mechanism | What changes? |
|---|---|
| prompting / in-context learning | input context |
| full fine-tuning | many or all model weights |
| LoRA | a low-rank parameter update |
| distillation | a new student model is trained |
| MoE routing | which parameter blocks are active for a token |
A better prompt is not a weight update. LoRA is not merely “training fewer layers.” Sparse MoE is not the same as pruning. All may reduce cost or specialize behavior, but through different mechanisms.
Interactive browser lab
Change model dimension, LoRA rank, number of experts, and active experts per token. The lab compares full-matrix parameter count with LoRA trainable parameters and compares total MoE expert capacity with the share activated per token.
A separate toggle labels whether an example change is an inference-time context change or a parameter update.
Python and R lab
Construct a small low-rank update \(BA\), apply it to a base matrix, and report the full-update versus LoRA trainable parameter counts. Then compute total and active expert parameter counts for a toy sparse MoE layer.
Practice:
- Explain why in-context learning can change predictions without changing model weights.
- Compute LoRA trainable parameters for a square \(d\times d\) matrix and rank \(r\).
- Distinguish total MoE parameters from active parameters per token.
- Explain why more pretraining compute does not remove the need for downstream evaluation.
Learner output
Distinguish pretraining, in-context learning, full fine-tuning, LoRA, and sparse MoE by stating which parameters or activations change in each case.