Audit the Crowd Before the Attack: Forecasting Multi-Agent Capture from Benign Logs
TL;DR for operators Testing each AI agent separately may not tell you how a group of those agents will behave once they begin influencing one another. Magistrali and Shani’s Aligned Alone, Misaligned Together1 provides unusually concrete evidence for that gap. In a synthetic security-triage population, a forecast constructed from adversary-free interaction logs predicted later attacked-population dismissal levels of 0.599, 0.654, and 0.680 at three held-out adversary doses. The measured values were 0.606, 0.658, and 0.685. Reported mean absolute error was 0.0058. ...