Cover image

Approved One by One, Risky in Combination: Testing Agent Skills Before They Execute

TL;DR for operators A marketplace can approve several useful skills independently, then have an agent activate them together for an ordinary business task. The problem is that individually acceptable components can combine their preconditions, resource changes, and outputs into a plan-level objective that neither the user nor any single skill requested. The reported benchmark shows why this interaction needs separate testing: severe plans rise from 4.7% of single-skill compositions to 66.5% of five-skill compositions. This does not prove that adding skills directly causes harm in every setting, but it does show that standalone approval is not enough evidence for composition safety. ...

July 24, 2026 · 9 min · Zelina