The Guard Is Already in the Draft: Reusing Speculative Decoding for LLM Monitoring
TL;DR for operators Production monitoring often creates an unattractive choice. A small probe is cheap enough to run on every request but may compress away evidence that occurred earlier in a sequence. A stronger position-aware classifier preserves more information but adds computation to an inference path that is already expensive. Speculative Probing: LLM Monitoring at Speculative-Decoding Cost1 asks whether part of that cost has already been paid. Some LLM deployments use a smaller auxiliary prediction component to accelerate generation. Speculative Probing reuses that component—its speculative-decoding head—as a frozen feature extractor for monitoring. The base model and draft head remain unchanged; only a few learned task-specific query vectors and a small classifier are trained. ...