Make Verification Expensive: RedHerring Turns an Agent’s Search Budget Into a Security Surface
TL;DR for operators An autonomous security agent can notice far more suspicious code paths than it can fully prove vulnerable during a fixed audit window. In the paper’s 300-trajectory analysis, verification consumed 73.8% of wall-clock time, compared with 7.3% for hypothesis formation; verifying one hypothesis took about 11.1 times as much time as generating one. ...