Econometrics bridge: Weighted averages, kernels, distributed lags, and state-space alternatives
Estimated time: 120 min
Lab: Open browser lab
Code: Python - R

Why this should feel familiar

A weighted average is familiar. The conceptual jump in attention is that the weights are not fixed after estimation. They are recomputed from the current representations for every input.

A distributed-lag regression might use fixed coefficients

\[ y_t=\sum_j \beta_j x_{t-j}. \]

Attention instead uses weights \(\alpha_{tj}\) that depend on the current query and candidate context:

\[ z_t=\sum_j \alpha_{tj}v_j. \]

Mathematical core

For input representations collected in matrix \(X\), learned projections produce

\[ Q=XW_Q, \qquad K=XW_K, \qquad V=XW_V. \]

Scaled dot-product attention is

\[ \operatorname{Attention}(Q,K,V) =\operatorname{softmax}\left(\frac{QK^\top}{\sqrt{d_k}}+M\right)V, \]

where \(M\) is a mask. In a causal language model, future positions receive effectively minus-infinite score so token \(t\) cannot attend to tokens that have not yet occurred.

For one query \(q_t\),

\[ \alpha_{tj}=\frac{\exp(q_t^\top k_j/\sqrt{d_k})} {\sum_{\ell}\exp(q_t^\top k_\ell/\sqrt{d_k})}. \]

The output is a context-dependent weighted combination of value vectors.

Multi-head attention

Multiple heads use different projection matrices. Each head can emphasize different relationships before their outputs are combined. This increases representational flexibility, but an individual attention head should not automatically be interpreted as a stable semantic or causal mechanism.

A transformer block

A modern transformer block contains more than attention:

  1. normalized hidden representations;
  2. multi-head attention;
  3. a residual connection;
  4. a position-wise feed-forward network;
  5. another residual path;
  6. positional information supplied explicitly or through a positional mechanism.

The feed-forward network transforms each token representation separately, while attention exchanges information across token positions.

Why attention changed sequence modeling

RNNs compress previous information into a recurrent state. Self-attention gives a token direct access to many earlier token representations. For sequence length \(n\), full self-attention forms \(n^2\) pairwise scores, which improves connectivity but creates a quadratic computation and memory cost in sequence length.

Econometrician’s checkpoint

Attention weights are data-dependent computational weights, not regression coefficients estimated once for the population. A high weight says that a value vector contributes strongly to this particular forward pass under this trained model. It does not by itself establish feature importance, structural interpretation, or causality.

The transformer also changes the role of sequence state. Instead of one recurrent vector carrying history forward, the model can repeatedly recompute interactions across a context window.

The language-model objective still looks familiar:

\[ \max_\theta \sum_t \log p_\theta(x_t\mid x_{What changed from the n-gram module is the representation of \(x_{

Interactive browser lab

Use four fixed token representations. Change the attention temperature and choose whether causal masking is active. The lab prints the score matrix, the normalized attention rows, and the resulting weighted values.

The textual matrix makes two facts visible: weights change with the current query, and causal masking removes future positions from the normalization set.

Python and R lab

Compute a small attention matrix from fixed query, key, and value vectors. Compare dense and causal attention and verify that each allowed row sums to one. Then report how the first token’s allowed context differs from the last token’s.

Practice:

  1. Derive the softmax attention weights for one query and two keys.
  2. Explain why attention weights are not coefficient estimates with ordinary econometric interpretation.
  3. Explain why causal masking is an information-boundary rule rather than a regularization trick.
  4. Compare the information path in an RNN with the information path in full self-attention.

Learner output

Explain how attention weights differ from fixed regression coefficients and why causal masking is required for autoregressive language modeling.