Cover image

Watching a Signal the Model Can Move

TL;DR for operators A latent-space monitor is valuable only if the internal signal it reads is sufficiently difficult for the monitored model to manipulate. Measuring Activation Control in Large Language Models tests that assumption directly.1 Across 25 open-weight instruction-tuned models, the authors find meaningful but coarse control over concept-related internal representations. Models can raise a concept signal, suppress it toward its ordinary baseline, order several requested intensity levels, and move signal toward broad regions of a sentence. They are much less successful at targeting a particular layer or restricting modulation to particular token groups. ...

September 28, 2026 · 9 min · Zelina
Cover image

Catch the Drift Before the Answer: Reasoning Trajectories as a Runtime Control Surface

TL;DR for operators When a model is midway through a difficult task, the expensive choice is often whether to leave it alone, verify it, or spend more inference on correction. Doing that for every request wastes compute and can damage answers that were already on track. Lihao Sun and colleagues show that the model’s internal state changes systematically as reasoning progresses, and that late-stage changes contain a useful warning signal for eventual failure.1 Their trajectory-based correctness features reach a seed-averaged best-layer ROC-AUC of about 0.85, versus 0.765 for the strongest reported LogitLens baseline and 0.649 for reasoning step count alone. ...

September 16, 2026 · 8 min · Zelina
Cover image

One Score, More Signals: Making Reward Models Easier to Rank and Audit

TL;DR for operators A system generating several candidate answers eventually needs a ranking decision: which response should be shown, which should be discarded, and which should receive additional review. A reward model commonly reduces that decision to one scalar score derived from the prompt and response text. Oprea and Bâra test whether that score improves when the model is also given four explicit signals—response length, toxicity, refusal behavior, and prompt-response semantic similarity—and allowed to interpret those signals jointly with the text representation.1 On Anthropic HH-RLHF, the answer is consistently yes across ten evaluated model configurations. The strongest DeBERTa-v3 reward model moves from 0.74 to 0.84 ROC-AUC and from 0.72 to 0.83 pairwise accuracy. ...

September 15, 2026 · 6 min · Zelina
Cover image

Role Call: Who Your Agents Are Actually Listening To

TL;DR for operators Teams are easy to label. Understanding who actually listens to whom is harder. Hong’s paper on learned coordination conventions proposes a diagnostic for inspecting how cooperative reinforcement-learning agents route information between predefined roles.1 The central move is architectural: place role labels in both the querying agent’s representation and each ally’s representation, then use cross-attention to expose a role-to-role routing matrix. ...

July 10, 2026 · 20 min · Zelina
Cover image

The Sticker on the Dashboard Is Not Steering

TL;DR for operators A policy, prompt, adapter, steering vector, or internal patch can make a model look more orderly. That does not mean it controls the model. The paper’s central distinction is brutal and useful: order is visible structure; control is validated movement through the right receiver under the right conditions, with side effects bounded.1 ...

June 27, 2026 · 20 min · Zelina
Cover image

Mind the Readout: Why AI Gets Smarter When We Stop Worshipping the Output

The current AI industry has a strangely theatrical relationship with intelligence. We judge models by the visible performance: the answer they print, the image they reconstruct, the attention map they expose, the number of reasoning steps they perform, the architectural flourish in the diagram. If the output looks sophisticated, we call the system capable. If the output looks wrong, we assume the capability is missing. This is convenient, measurable, and often completely misleading. Naturally, it is popular. ...

June 13, 2026 · 15 min · Zelina
Cover image

Think Meter, Not Think Bigger: The New Control Layer for AI Reasoning

Most companies do not actually want an AI system that “thinks longer.” They want one that knows when extra thinking is worth the bill. That distinction is becoming more important. Reasoning models are moving from demo-stage math puzzles into document review, financial research, compliance analysis, customer support escalation, and agentic workflows. In these settings, reasoning has three costs: latency, compute, and misplaced confidence. A model that spends 30 seconds producing an elegant wrong answer has not reasoned. It has performed expensive theatre. Very fluent theatre, admittedly. ...

June 2, 2026 · 14 min · Zelina
Cover image

High Entropy, Low Drama: The Internal Fingerprint of LLM Reasoning

Debugging a reasoning model usually starts at the wrong end. A model gives a wrong mathematical answer, so we inspect the final output. Then we inspect the chain-of-thought. Then we compare benchmark scores, sample more answers, compute pass rates, and hope the model’s visible reasoning trace tells us what happened inside. This is convenient. It is also a little like diagnosing a factory by reading only the shipping label. ...

May 31, 2026 · 15 min · Zelina
Cover image

Reasonable Doubt: Why LLM Reasoning Needs Process Control

Why this matters now The business case for LLMs has quietly moved from chatbot answers to agentic work: legal review, compliance checking, market research, document synthesis, internal analytics, coding support, and decision preparation. That shift changes the risk profile. A wrong chatbot answer is annoying. A wrong agent that looks coherent, cites documents, calls tools, updates files, and confidently stops too early is a workflow liability wearing a productivity costume. ...

May 31, 2026 · 12 min · Zelina
Cover image

Pre-Decision Intelligence: When AI Decides Before It Thinks

Audit logs are comforting things. They tell managers that a system took an action, they tell engineers which step fired, and they tell compliance teams that someone, somewhere, has a line of text to point at when the incident review begins. Now imagine an AI agent inside a business workflow. It has a customer request, a list of available tools, and a visible reasoning trace. The trace says it carefully considered whether to call an API, ask for missing information, or answer directly. It sounds deliberate. It sounds inspectable. It sounds like governance. ...

April 2, 2026 · 16 min · Zelina