Cover image

Safety Without the Data Lake: Federating the Guard, Not the Traces

TL;DR for operators Suppose several business units or partner organizations run different agent workflows. Each has its own prompts, tools, communication patterns, and failure cases. They want a common safety layer, but centralizing those interaction traces would expose precisely the operational data they are trying to protect. The harder problem is that a guard trained elsewhere may not transfer well enough to solve this. In the reported experiments, an architecture-matched topology guard scores 0.512 AUROC when transferred off the shelf to Agent-SafetyBench, but 0.695 after in-domain retraining. Local adaptation helps, yet isolated local training is also weaker than collaborative training and becomes fragile when client labels are highly skewed. ...

September 27, 2026 · 7 min · Zelina