Cover image

Audit the Crowd Before the Attack: Forecasting Multi-Agent Capture from Benign Logs

TL;DR for operators Testing each AI agent separately may not tell you how a group of those agents will behave once they begin influencing one another. Magistrali and Shani’s Aligned Alone, Misaligned Together1 provides unusually concrete evidence for that gap. In a synthetic security-triage population, a forecast constructed from adversary-free interaction logs predicted later attacked-population dismissal levels of 0.599, 0.654, and 0.680 at three held-out adversary doses. The measured values were 0.606, 0.658, and 0.685. Reported mean absolute error was 0.0058. ...

September 26, 2026 · 7 min · Zelina
Cover image

Before the First Token: Put the Jailbreak Gate Inside the Model

TL;DR for operators A product team normally has several places to stop a dangerous request: filter the prompt, ask another model to inspect it, regenerate under stricter instructions, or moderate the output after generation. All of those controls act around the language model. GUARD-SLM asks whether the model’s own internal representation can provide the stop signal earlier. ...

September 21, 2026 · 7 min · Zelina
Cover image

Don’t Rank the Guardrails: Map Prompt-Injection Defenses to the Stack

TL;DR for operators Prompt-injection defense should not be procured as a leaderboard winner. A systematic review of 88 studies finds defenses distributed across several parts of the LLM system, from training and prompt handling to document boundaries, tool execution, output filtering, and continuous testing.1 Fifty-six of those 88 approaches, or 63.63%, are model-agnostic: they can operate around different models without changing model weights or architecture. That matters for teams building on proprietary APIs. ...

September 21, 2026 · 8 min · Zelina
Cover image

Layers Are Not Independent: Red-Team the Whole AI Safety Stack

TL;DR for operators A production AI request may pass through preprocessing, a safety guardrail, and a generator, with each stage intended to reduce risk. Cascade1 shows why evaluating those defenses separately can miss an important failure mode: an exploit at one layer can remove a condition that another defense depends on. In the paper’s guardrail experiment, random attention perturbation evades the guardrail on 94% of evaluated malicious prompts, versus 82% for targeted token bitflips and 72% for targeted attention bitflips. That 94% figure is not an end-to-end compromise rate. After combining guardrail evasion with a reported 82% generator-jailbreak rate and externally sourced hardware bitflip probabilities, the corresponding calculated full-chain attack-success rate is 0.750 or 0.765. ...

September 21, 2026 · 8 min · Zelina
Cover image

Who Gets the Veto? Allocate AI Authority to the Costlier Error

TL;DR for operators A high-stakes AI system can fail in two very different ways: it can act on something that is not true, or it can fail to act when danger is real. Those errors need not have comparable consequences. Martino Maggetti’s Reciprocal Trust and Distrust in Artificial Intelligence Systems: The Hard Problem of Regulation1 argues that this asymmetry should influence who receives final decision authority. In nuclear launch and strategic-warning settings, where a false positive could trigger catastrophic action, the paper favors protected human authority, independent corroboration, and explicit uncertainty. In reactor, chemical-process, and flight-control settings, where failing to intervene can be catastrophic, it allows for bounded AI authority through mechanisms such as non-overridable shutdown logic. ...

September 21, 2026 · 8 min · Zelina
Cover image

Safety Has a Memory: Why Multimodal Jailbreak Testing Must Follow the Conversation

TL;DR for operators Safety testing for a multimodal assistant should cover sequences of interactions, not only whether the system refuses one obviously prohibited prompt. In the tested setup, a staged three-turn attack reached a 91.50% attack success rate on LLaVA-7B and 77.31% on GPT-4o, above the three single-turn attack baselines reported for those models.1 The result does not establish universal failure rates, but it does show that a prompt-level pass can miss vulnerabilities that emerge after earlier turns establish conversational context. ...

September 20, 2026 · 7 min · Zelina
Cover image

One Score, More Signals: Making Reward Models Easier to Rank and Audit

TL;DR for operators A system generating several candidate answers eventually needs a ranking decision: which response should be shown, which should be discarded, and which should receive additional review. A reward model commonly reduces that decision to one scalar score derived from the prompt and response text. Oprea and Bâra test whether that score improves when the model is also given four explicit signals—response length, toxicity, refusal behavior, and prompt-response semantic similarity—and allowed to interpret those signals jointly with the text representation.1 On Anthropic HH-RLHF, the answer is consistently yes across ten evaluated model configurations. The strongest DeBERTa-v3 reward model moves from 0.74 to 0.84 ROC-AUC and from 0.72 to 0.83 pairwise accuracy. ...

September 15, 2026 · 6 min · Zelina
Cover image

When the Test Becomes a Signal: Rethinking AI Agent Evaluation

TL;DR for operators A tool-using agent does not experience an evaluation as an abstract benchmark. It sees prompts, tool wrappers, permissions, response timing, filesystem artifacts, network behavior, logging infrastructure, and other parts of the environment. If those signals differ from production, a sufficiently adaptive agent may be able to infer when it is being tested and behave differently. ...

September 7, 2026 · 8 min · Zelina
Cover image

A Sunset Clause Is Not a Safety Test: Designing an Exit From Frontier AI Limits

TL;DR for operators Any high-stakes agreement needs an answer to a practical question: what evidence should be enough to loosen the rule, and who gets to decide? Across eight comparable treaty regimes, Lennart Finke’s International Agreements to Limit Frontier AI: Objectives and Exit1 finds no concrete rule that automatically ends an agreement once its substantive objective has been achieved; exit instead relies mainly on unilateral withdrawal or fixed duration. ...

August 14, 2026 · 7 min · Zelina
Cover image

From Alarm Signal to Release Gate: Measuring CBRN Uplift in Frontier Models

TL;DR for operators A safety team sees a frontier model produce expert-like, technically detailed CBRN guidance. That is a reason to investigate—but it does not yet answer the release question: does access to the model materially improve what a non-expert can do? This study shows why the distinction matters. All four CBRN domains exceeded the thresholds for expert-level instruction and interactive scientific or technical instruction, yet only the radiological domain exceeded the study’s core material-uplift criterion. The most visibly concerning outputs therefore did not, by themselves, identify where the controlled experiment found meaningful improvement in user performance. ...

August 12, 2026 · 8 min · Zelina