Econometrics bridge: Distributed lags, moving filters, AR models, state-space recursion
Estimated time: 110 min
Lab: Open browser lab
Code: Python - R
Why this should feel familiar
CNNs and RNNs are different architectural answers to the same broad question: how should a neural network exploit structure rather than treating every input coordinate as unrelated to every other coordinate?
A convolution resembles a learned local filter with shared coefficients. A recurrent network resembles a nonlinear state recursion whose hidden state carries information forward through a sequence.
Mathematical core
Convolution: local structure plus weight sharing
For a one-dimensional input \(x_t\) and kernel \(w_0,\ldots,w_{K-1}\),
\[ h_t=\sigma\left(b+\sum_{j=0}^{K-1}w_j x_{t-j}\right). \]The same kernel weights are used at every position. This gives two assumptions:
- nearby observations are especially relevant;
- the same local pattern can matter in different positions.
The analogy to a distributed-lag filter is useful, but a CNN usually stacks many learned channels and nonlinearities. Its weights are computational features, not automatically interpretable lag coefficients.
Recurrence: learned hidden state
A simple recurrent network is
\[ h_t=\tanh(W_x x_t+W_h h_{t-1}+b), \]with an output such as
\[ y_t=W_y h_t+c. \]The hidden state summarizes previous information. This resembles a state-space recursion, but the state is optimized for the task rather than assigned a structural economic interpretation.
Why recurrent gradients can vanish or explode
Backpropagation through time repeatedly multiplies Jacobians. In a scalar simplification,
\[ \frac{\partial h_t}{\partial h_{t-k}} \approx \prod_{j=t-k+1}^t w_h \sigma'(a_j). \]Repeated factors below one shrink gradients; factors above one can amplify them. Gated recurrent architectures such as LSTMs and GRUs were designed partly to make long-range information and gradients easier to preserve.
Architecture as an assumption
The important lesson is not a contest between CNNs and RNNs. Both impose structure on the function class.
- CNN: repeated local transformation and shared parameters;
- RNN: sequential state update and parameter sharing through time;
- fully connected network: fewer built-in assumptions about locality or sequence.
These restrictions can improve statistical efficiency when they match the data structure and hurt when they do not.
Econometrician’s checkpoint
Econometricians often write the data structure into a model explicitly: lags, moving averages, latent states, fixed effects, or panel indexes. Deep learning also encodes assumptions, but often through architecture rather than through a small set of interpretable coefficients.
A high recurrent weight is not analogous to a single persistent AR coefficient in every respect. It interacts with nonlinear activations, other hidden dimensions, gates, normalization, and the training objective.
This module is intentionally historical as well as practical. CNNs remain central in vision and local-signal processing; RNNs remain useful and conceptually important. But transformers will soon replace recurrence with direct context-dependent interactions for many language tasks.
Interactive browser lab
Use the same short numeric sequence in two mechanisms.
- Change a three-element convolution kernel and inspect the local-filter response.
- Change the recurrent state weight and inspect the hidden-state path from the same sequence.
The lab reports the arrays as text so the mechanism remains accessible without a visual plot.
Python and R lab
Compute a valid one-dimensional convolution and a scalar tanh recurrence. Compare recurrent weights below, near, and above one and report how the hidden-state path and a simple gradient-magnitude proxy change.
Practice:
- Explain why weight sharing reduces parameter count in a convolution.
- State the modeling assumption implied by a local receptive field.
- Compare an AR state variable with an RNN hidden state.
- Explain why a recurrent architecture can have long memory in principle but still be difficult to train.
Learner output
Compare the information-flow assumptions of a convolutional filter and a recurrent hidden state.