Econometrics bridge: Nearest-neighbor prediction, information retrieval, measurement error, model combination
Estimated time: 105 min
Lab: Open browser lab
Code: Python - R

Why this should feel familiar

A pretrained language model stores broad statistical regularities in its parameters. Retrieval-augmented generation adds a second information source: documents selected at inference time from an external collection.

This changes the system from

\[ x\rightarrow p_\theta(y\mid x) \]

to something closer to

\[ x\rightarrow \text{retrieve }z\rightarrow p_\theta(y\mid x,z). \]

The original RAG formulation made the separation explicit by combining a retriever over documents with a generator conditioned on retrieved evidence.

Mathematical core

Let \(z\) denote a document or chunk. A retriever defines a distribution

\[ p_\eta(z\mid x), \]

and a generator defines

\[ p_\theta(y\mid x,z). \]

A latent-document formulation can write

\[ p(y\mid x)=\sum_z p_\eta(z\mid x)p_\theta(y\mid x,z). \]

Practical systems often approximate the sum with top-k retrieved chunks and place them in the model context.

Lexical and dense retrieval

Lexical retrieval scores overlap in observed terms. Dense retrieval maps queries and documents into learned vectors and uses a similarity such as cosine similarity or a dot product.

Neither dominates universally. Exact names, numbers, and rare identifiers can favor lexical signals; semantic paraphrases can favor dense representations. Production systems often combine them.

Retrieval creates a new measurement problem

The generator can only use evidence that reaches its context. End-to-end failure can therefore arise from several stages:

  1. the relevant document is absent from the corpus;
  2. chunking separates the needed evidence;
  3. the retriever ranks the relevant chunk too low;
  4. too much context dilutes useful evidence;
  5. the generator ignores or misuses retrieved text;
  6. the retrieved source is stale, conflicting, or untrustworthy.

This is why RAG evaluation must not be reduced to answer quality alone.

RAG is not merely prompt engineering

A prompt-only system changes the text presented to a fixed model. RAG adds a retrieval model, an external corpus, ranking, chunking, and a context-construction policy. Some implementations use prompts to tell the generator how to use evidence, but retrieval itself is a separate statistical and systems component.

The distinction between parametric memory and external/non-parametric memory is useful:

  • parametric knowledge is encoded in model weights;
  • external knowledge can be updated without retraining the base model;
  • retrieval decides what portion of that external store enters the context for a particular query.

Econometrician’s checkpoint

Treat retrieval quality as an upstream measurement process. If relevant evidence is systematically missing or mis-ranked, the downstream generator is conditioned on a biased information set.

Useful diagnostics include recall@k for known relevant documents, rank of the relevant chunk, context precision, answer accuracy, citation/evidence support, and robustness to distracting retrieved text.

A larger top-k is not automatically better. It can increase recall while also increasing irrelevant context, token cost, and opportunities for contradictory evidence.

Interactive browser lab

Choose one of several fixed queries and change top-k. The lab scores a tiny local document collection with deterministic term-frequency cosine similarity, prints the ranking, and constructs the exact retrieved context that a generator would receive.

Because the corpus is built into the lab, no network or external API is used.

Python and R lab

Implement the same local retriever. Report top-k rankings for each query and show one case where increasing k adds irrelevant context without changing whether the relevant document was already retrieved.

Practice:

  1. Distinguish a retrieval failure from a generation failure.
  2. Explain why updating an external corpus differs from fine-tuning model weights.
  3. Give one case where lexical retrieval may outperform dense semantic similarity.
  4. Design separate retrieval and answer-quality metrics for a business knowledge assistant.

Learner output

Diagnose a RAG failure by separating retrieval failure, context failure, and generation failure.