Econometrics bridge: Sequential decision rules, state updating, search, verification, and measurement
Estimated time: 135 min
Lab: Open browser lab
Code: Python - R
Why this should feel familiar
A language model maps context to a distribution over next tokens. An agentic system adds state, actions, external tools, observations, memory, and a control rule that decides what happens next.
This returns to the broad AI idea of an agent, but now with a modern language model inside the loop.
A useful system abstraction is
\[ \text{observe}\rightarrow\text{update state}\rightarrow\text{choose operation}\rightarrow\text{act}\rightarrow\text{observe again}. \]The operation may be another model call, retrieval, a calculator, code execution, database lookup, or a request for human approval.
Reasoning as inference-time computation
A direct answer uses one forward generation path. Reasoning methods allocate additional inference-time computation through intermediate steps, multiple candidate paths, verification, search, or repeated sampling.
Chain-of-thought prompting made intermediate natural-language steps a visible interface for this computation. Later systems may hide, compress, verify, or replace free-form reasoning with structured search or external tools. The durable concept is test-time computation, not one prompting template.
More reasoning tokens do not guarantee a better answer. Additional computation helps only if the search, decomposition, or verification procedure adds useful information rather than compounding error.
Search and planning return in modern systems
Classical AI distinguishes learning from search. A learned model can propose candidate actions or states, while a search procedure evaluates or explores alternatives.
A generic search objective can be written as
\[ a_{1:T}^*=\arg\max_{a_{1:T}} U(a_{1:T}), \]subject to the environment transition rules. Modern reasoning systems can use an LLM to generate candidates and a verifier, reward model, program executor, or environment response to evaluate them.
ReAct: reasoning plus acting
A ReAct-style loop interleaves model-generated decisions with external actions and observations:
\[ \text{state}_t \rightarrow \text{action}_t \rightarrow \text{observation}_{t+1} \rightarrow \text{state}_{t+1}. \]The important architectural point is that the model does not need to contain every fact or calculation internally. It can decide to retrieve a document, call a calculator, execute code, or query a local system, then condition subsequent behavior on the returned observation.
Memory
Agent memory can refer to several different mechanisms:
- current context window;
- retrieved external records;
- a running structured state or scratch record;
- summaries of prior interaction;
- persistent stores outside the model weights.
Do not call all of these “memory” without saying where the information lives, how it is written, how it is retrieved, and when it can become stale.
Multi-agent systems
Multiple model instances can divide roles such as planner, executor, critic, or verifier. This can improve modularity, but more agents do not automatically create more intelligence. Coordination introduces communication cost, correlated errors, role ambiguity, and new failure modes.
The right comparison is against a simpler single-agent or non-agent baseline under the same task and tool access.
Evaluation
Agent evaluation must include more than final-answer accuracy. Useful measurements include:
- task completion;
- tool-call accuracy;
- number and cost of steps;
- recovery after tool failure;
- retrieval support;
- state consistency over long horizons;
- unsafe or unauthorized actions;
- human-intervention rate;
- reproducibility under fixed seeds or deterministic tools where possible.
For high-impact operations, the action boundary matters. A system that can recommend, draft, or simulate is different from one that can execute an irreversible action.
Econometrician’s checkpoint
Agent traces create selection and dependence. The next observation depends on earlier actions, tool results, and routing choices. Treating every step as an independent row of data is usually wrong.
The course therefore ends where its two spines meet:
- learned representations and language-model policies provide flexible inference;
- Markov/sequential ideas describe evolving state and action;
- retrieval supplies external information;
- tools change the environment or compute exact operations;
- evaluation must cover the whole closed loop.
Interactive browser lab
Choose a fixed business task and compare a direct-answer mode with a deterministic tool-using loop. The local agent can call only two built-in tools: a small lookup table and a calculator. The lab prints a compact execution trace containing planned operations, tool calls, returned observations, and the final result.
The trace is a designed symbolic workflow for teaching system structure. It is not a hidden model chain-of-thought.
Python and R lab
Run the same deterministic state machine for three tasks. Count model-decision steps, tool calls, and verification checks. Compare the tool-using answer with a deliberately naive direct baseline.
Practice:
- Distinguish a language model from an agentic system.
- Give one example where search or verification adds inference-time computation without changing model weights.
- Specify where memory lives in a RAG-plus-agent system.
- Design one baseline that would test whether a multi-agent design adds value beyond one agent with the same tools.
- Mark the point in an agent workflow where human approval should be required for an irreversible action.
Learner output
Draw an agent loop that separates model inference, retrieval, tool execution, memory, verification, and environment state.