Cover image

Readable Before Steerable: Why Probe Accuracy Is Not a Control Test

TL;DR for operators An internal probe can be highly accurate without providing a reliable way to change model behavior. In Xining Xun’s study of language-model development,1 internal probe readability was already near ceiling at the earliest 1,000-step checkpoint across all six Pythia sizes. Yet interventions that pushed activations along those same probe directions were statistically indistinguishable from matched random-direction edits in 43 of 48 model-by-checkpoint cells. ...

October 2, 2026 · 7 min · Zelina