Cover image

Watching a Signal the Model Can Move

TL;DR for operators A latent-space monitor is valuable only if the internal signal it reads is sufficiently difficult for the monitored model to manipulate. Measuring Activation Control in Large Language Models tests that assumption directly.1 Across 25 open-weight instruction-tuned models, the authors find meaningful but coarse control over concept-related internal representations. Models can raise a concept signal, suppress it toward its ordinary baseline, order several requested intensity levels, and move signal toward broad regions of a sentence. They are much less successful at targeting a particular layer or restricting modulation to particular token groups. ...

September 28, 2026 · 9 min · Zelina