Cover image

Audit the Crowd Before the Attack: Forecasting Multi-Agent Capture from Benign Logs

TL;DR for operators Testing each AI agent separately may not tell you how a group of those agents will behave once they begin influencing one another. Magistrali and Shani’s Aligned Alone, Misaligned Together1 provides unusually concrete evidence for that gap. In a synthetic security-triage population, a forecast constructed from adversary-free interaction logs predicted later attacked-population dismissal levels of 0.599, 0.654, and 0.680 at three held-out adversary doses. The measured values were 0.606, 0.658, and 0.685. Reported mean absolute error was 0.0058. ...

September 26, 2026 · 7 min · Zelina
Cover image

Cheap Seats, Sharp Eyes: Reward-Hack Detection Without the Frontier Judge

TL;DR for operators A frontier LLM judge is an expensive way to inspect every agent trajectory for reward hacking. This paper asks whether a much smaller detector can do most of that monitoring job at much lower cost. The answer is: yes, under the same information condition, and with important caveats. A 13.8M-parameter transformer encoder plus a logistic regression probe detects reward hacking in cleaned Terminal-Wrench trajectories with 0.9467 AUC and 0.8296 TPR@5%FPR. In the authors’ matched comparison, a reproduced gpt-5.4 judge reaches 0.9510 AUC and 0.7130 TPR@5%FPR on the cleaned sanitized-vs-baseline split.1 ...

June 15, 2026 · 17 min · Zelina