The Model Felt the Tampering. It Couldn’t Name the Cause
TL;DR for operators Suppose middleware silently rewrites part of an AI agent’s own generated answer before the model continues. A reasonable expectation is that a capable model would notice the interference, or at least diagnose why its continuation has become strange. The Sleight of Word benchmark tests exactly that expectation.1 Across 19 open-weight instruction-tuned models, covert substitutions consistently change the models’ predictive distributions: post-swap surprisal and entropy rise for every model tested. But correct identification of the intervention is almost absent. No model exceeds 1.3% explicit switch awareness, and the pooled rate is below 0.1%. ...