Cover image

Before You Retrain the Guardrail, Ask What It Already Knows

TL;DR for operators When a deployed safety classifier begins making the wrong decisions, retraining the classifier is not necessarily the first intervention to test. Sandoval and Topcu’s Regime-Conditional Verification (RCV)1 adds a small correctness layer around a frozen classifier. It asks a narrower question than the classifier itself: given that the classifier just said “safe” or “unsafe,” how likely is that verdict to agree with the deployer’s policy? ...

August 29, 2026 · 9 min · Zelina