TL;DR for operators

A latent-space monitor is valuable only if the internal signal it reads is sufficiently difficult for the monitored model to manipulate. Measuring Activation Control in Large Language Models tests that assumption directly.1

Across 25 open-weight instruction-tuned models, the authors find meaningful but coarse control over concept-related internal representations. Models can raise a concept signal, suppress it toward its ordinary baseline, order several requested intensity levels, and move signal toward broad regions of a sentence. They are much less successful at targeting a particular layer or restricting modulation to particular token groups.

That distinction matters because suppression alone is enough to interfere with several monitors. Without fine-tuning against those monitors, suppression instructions substantially reduced detection by linear probes, the Jacobian lens, activation oracles, and natural-language autoencoders.

For teams using internal probes in release evaluation or runtime safety systems, ordinary probe accuracy is therefore not the complete acceptance test. Monitor robustness should also be evaluated under model-controlled activation changes. The paper does not establish that hardened production monitoring systems fail: its evasion tests use off-the-shelf monitors and simple safety-neutral concepts, and its harder-task experiment is limited to one model.

An internal signal is only safer if the model cannot move it

A safety team may inspect internal model states because visible outputs are incomplete evidence. A model could produce an acceptable answer while concealing information relevant to deception, sandbagging, or some other behavior the evaluator wants to detect.

That gives latent-space monitoring an apparent advantage over behavioral monitoring: the evaluator observes an internal representation rather than relying entirely on what the model chooses to say.

The assumption becomes weaker if the model can itself change the representation being observed.

Kowalski and colleagues test this by instructing models to alter concept-related activity while performing an otherwise fixed task. The internal channel they record is the residual stream, a representation carried through successive transformer layers. For each concept, the researchers construct a direction in that representation space and measure how strongly generated-token activations project onto it.

The relevant capability is therefore narrower than arbitrary self-editing. The question is whether a model can deliberately move a measured concept signal enough to affect what an observer sees.

The answer, in these experiments, is yes.

The benchmark separates six kinds of internal control

The Activation Controllability Benchmark does not treat control as a single binary property. It measures six dimensions:

Dimension What the test asks
Engage Can the model increase a named concept signal on instruction?
Suppress Can it move that signal downward relative to its ordinary level?
Dial Rank Can requested intensity levels be placed in the correct order?
Temporal Control Can activation be concentrated in a requested broad sentence region?
Coverage Does control appear across different token categories?
Layer Targeting Can modulation be concentrated at a specifically requested layer?

Five dimensions are standardized with a $d’$-style effect size, while intensity ordering uses Spearman correlation. The authors then map the measures onto a common scale and combine them with a geometric mean.

That aggregation is intentionally conjunctive: weakness on one dimension pulls down the overall score rather than being completely offset by strength elsewhere. The composite score $S$ ranges from 0.059 to 0.538 across the 25 models.

The spread is more informative than a simple capability/no-capability label. Controllability also does not track parameter count in the tested roster; Llama 3.1 8B Instruct receives the highest composite score. This is a comparative result, not evidence that smaller models are generally more controllable or that scale causes controllability to fall.

For model evaluators, direct measurement is therefore preferable to treating size or model generation as a proxy.

Current control is broad, not surgically precise

The strongest evidence concerns magnitude.

All tested models show the ability to raise concept-related activation under an engagement instruction, while suppression generally moves the signal toward the no-instruction baseline. Virtually all models can also place four requested intensity levels in approximately monotonic order.

Models have some control over when the signal appears as well. They can preferentially move activation toward broad instructed regions of a sentence, with control often strongest near the end. But the timing remains imprecise.

Spatial control inside the network is much weaker. No tested model reliably concentrates activation at a specifically instructed layer. Attempts to restrict modulation to particular token groups are near zero or negative for essentially the entire roster.

This rules out an overly strong reading of the result. The study does not show models precisely rewriting arbitrary internal states.

What it shows is a less dramatic but operationally consequential capability: models can move the magnitude and broad placement of a monitored representation without possessing fine-grained control over exactly where that modulation occurs.

The activation also tends to be strongest in the deeper half of the network and is visible across token categories. Engagement is particularly strong on low-content tokens such as punctuation and determiners. The authors suggest that these tokens may provide representational capacity for carrying the modulation, but they present this as a plausible explanation rather than a demonstrated mechanism.

Harder tasks do not erase the effect in the model tested

A laboratory transcription task leaves open an important question: does internal control disappear once the model has something difficult to solve?

The paper treats this as a robustness test rather than a second benchmark. Gemma 3 27B Instruct solves 800 polynomial-factorization problems spanning five preregistered difficulty levels while receiving activation-control instructions.

Baseline task accuracy falls sharply across those recipes, from 96.25% to 10.63%. Yet activation modulation remains positive across the difficulty range. The estimated slope is $-0.0255$, with a 95% confidence interval of $[-0.0428,+0.0109]$.

That interval crosses zero. The experiment therefore provides no evidence of a systematic monotonic increase or decrease in activation control as task difficulty rises. The concept-thinking instruction also produces no detectable reduction in answer accuracy.

This finding is useful mainly for what it prevents us from assuming: competing task demands do not automatically eliminate activation modulation. But the test uses only Gemma 3 27B Instruct, so it cannot establish task-load invariance across model families.

A separate checkpoint analysis offers descriptive evidence that the capability can increase through ordinary training. In the Olmo models examined, measured controllability rises from lower levels early in pretraining to higher levels at the base and aligned checkpoints, despite no explicit training objective for activation control. Because these checkpoint results are descriptive and several lack bootstrap intervals, they support tracking the capability over development rather than a causal claim about which training stage produces it.

Probe accuracy and probe robustness are separate release criteria

The monitor experiments show why the benchmark belongs in safety evaluation rather than only interpretability research.

The monitors are given residual activations from the generated text but not the hidden instruction tokens. Differences in detection therefore reflect changes in the generated-span internal representation rather than a monitor simply noticing the suppression instruction itself.

At a 1% false-positive rate, linear-probe detection under suppression falls to roughly 5–21%, while engagement remains near ceiling for most tested models. The Jacobian lens reports the suppressed concept in fewer than 10% of trials for 11 of 13 fitted models. For Gemma 3 27B, the natural-language autoencoder verbalizer names the suppressed concept in 0 of 500 suppression trials. Activation-oracle results point in the same overall direction.

For a safety team, this creates two different evaluation questions:

  1. Does the monitor detect the target representation under ordinary conditions?
  2. Does it remain informative when the model deliberately changes that representation?

A monitor can perform well on the first test and poorly on the second.

That distinction can be incorporated into release gating. Model developers can measure controllability for each new checkpoint or model generation, then rerun monitor evaluations under engagement, suppression, timing, and other control conditions. Large changes in temporal or layer-specific control would be particularly relevant because they could alter which monitoring locations remain dependable.

The benchmark’s diagnostic dimensions can also guide monitor design. Current models show weak layer-specific and token-group targeting, so monitoring architectures that aggregate evidence across locations the model cannot selectively manipulate may be harder to evade than a single fixed probe. That is an engineering inference from the reported capability pattern, not a result directly demonstrated by the paper.

What the study does not establish

The monitor-evasion result should not be read as evidence that latent-space monitoring is unusable.

The evaluated monitors are off-the-shelf rather than adversarially hardened or trained on evasive activations. A production system could respond by training monitors on suppression attempts, combining multiple representation spaces, or exploiting dimensions where current models have weaker control. This study does not test whether those defenses succeed.

The concepts are also deliberately simple and safety-neutral. A representation for “bread” may be substantially easier to identify and manipulate than a distributed state associated with deception, evaluation awareness, or another safety-relevant property.

The composite score has its own interpretation boundary. Its six dimensions are not statistically independent; Engage and Coverage are especially correlated. $S$ is therefore a compact summary of several aspects of control, not six independent confirmations of one latent capability.

These boundaries narrow the conclusion without removing the practical concern. The paper demonstrates that some internal observation channels are already manipulable through ordinary natural-language instruction, and that the resulting changes can materially weaken several standard monitors.

Test the observation channel, not just the detector

Latent-space monitoring adds information that behavioral evaluation cannot see. That advantage remains valuable, but it does not make internal activations intrinsically tamper-resistant.

The relevant engineering question is no longer only whether a probe can decode a useful concept. It is also whether the monitored model can move that signal, under what instructions, at which depths, across which tokens, and far enough to change a safety decision.

The Activation Controllability Benchmark provides a way to measure those properties directly. Its current results describe a capability that is real but coarse: strong enough to interfere with several monitors, yet far from precise arbitrary control of internal computation.

For organizations relying on activation probes, that is sufficient reason to add controllability to the test suite. A monitor should be evaluated not only on what it can read, but on whether the system being monitored can change what there is to read.

Cognaptus: Automate the Present, Incubate the Future.


  1. Marek Mateusz Kowalski and Joshua Fonseca Rivera and Uzay Macar and David Demitri Africa (2026). Measuring Activation Control in Large Language Models. arXiv:2608.21664. https://arxiv.org/abs/2608.21664 ↩︎