Cover image

Before You Retrain the Guardrail, Ask What It Already Knows

TL;DR for operators When a deployed safety classifier begins making the wrong decisions, retraining the classifier is not necessarily the first intervention to test. Sandoval and Topcu’s Regime-Conditional Verification (RCV)1 adds a small correctness layer around a frozen classifier. It asks a narrower question than the classifier itself: given that the classifier just said “safe” or “unsafe,” how likely is that verdict to agree with the deployer’s policy? ...

August 29, 2026 · 9 min · Zelina
Cover image

Take Over, Then Let Go: AutoIntervene Makes Robot Recovery a Training Signal

TL;DR for operators When an imitation-learned robot drifts outside the situations represented in its demonstrations, detecting that something looks unusual solves only half the operational problem. A deployment system must also decide when human takeover is warranted, when the recovery has progressed far enough to return control, and whether that intervention can improve the next policy rather than disappear as one-off operational labor. ...

August 22, 2026 · 8 min · Zelina