Cover image

A Clean Jailbreak Cluster Can Still Miss Unsafe Compliance

TL;DR for operators A safety monitor can become highly accurate at recognizing jailbreak-shaped prompts without becoming equally accurate at predicting unsafe model behavior. Delcon, Algaba, and Ginis demonstrate this gap across six Qwen and Llama instruction-tuned models.1 Their internal embeddings separate control and jailbreak prompts with balanced accuracy around 0.99–1.00. Yet in Qwen-2.5-7B, where refusal and compliance observations are comparatively balanced, the corresponding refusal-versus-compliance separation reaches only 0.677. ...

August 28, 2026 · 7 min · Zelina
Cover image

Audit the Step, Not the Aftershock

TL;DR for operators When an autonomous analysis agent makes one questionable operation, every later operation may inherit the damage. An audit that merely asks which steps look unusual can therefore produce a long queue of downstream symptoms rather than isolate where new violations occurred. Ahmed Hassoon and Mark Dredze formalize a different target: score each operation according to whether it behaved as expected given the state it actually received.1 Under their assumptions, a correctly executed downstream step remains statistically null even when its input was already corrupted. This makes one-step scoring useful for narrowing a review queue. ...

August 28, 2026 · 8 min · Zelina
Cover image

Decision Rights, Not More Layers: What an Auditable Fraud Pipeline Actually Earns

TL;DR for operators A fraud classifier has already scored a transaction. The next operational choice is whether extra context—relationship patterns, anomaly signals, explanations, or an LLM investigator—should merely inform the case or be allowed to change the decision. In Rahil Sharma’s evaluation, the answer is component-specific.1 The bounded LLM investigator was correct on 39 of 60 deliberately balanced difficult cases, versus 43 of 60 for simply applying a 0.5 threshold to the classifier: 65.0% versus 71.7%. The agent changed eight classifier decisions. Two changes fixed mistakes; six replaced correct decisions with incorrect ones. ...

August 24, 2026 · 7 min · Zelina
Cover image

The Reviewer Was Right. The Workflow Still Failed.

TL;DR for operators A quality-control component can correctly identify a defect and still add little value if the next stage ignores the correction. That distinction matters for AI workflows built around critics, reviewers, validators, or approval agents: reviewer accuracy measures whether the warning is right, not whether the warning changes what the system ultimately does. ...

August 18, 2026 · 7 min · Zelina
Cover image

When Saying Less Scores More: The Win-by-Silence Failure in AI Plan Evaluation

TL;DR for operators If an AI-generated plan is scored before its real-world outcome is known, a higher score need not mean a stronger or more complete plan. In the benchmark studied here, every one of 26 routes had at least one intermediate transition whose deletion increased the fixed-parameter score. Across all 57 admissible deletions, 27 raised the score. ...

August 17, 2026 · 8 min · Zelina
Cover image

A Sunset Clause Is Not a Safety Test: Designing an Exit From Frontier AI Limits

TL;DR for operators Any high-stakes agreement needs an answer to a practical question: what evidence should be enough to loosen the rule, and who gets to decide? Across eight comparable treaty regimes, Lennart Finke’s International Agreements to Limit Frontier AI: Objectives and Exit1 finds no concrete rule that automatically ends an agreement once its substantive objective has been achieved; exit instead relies mainly on unilateral withdrawal or fixed duration. ...

August 14, 2026 · 7 min · Zelina
Cover image

Control in Degrees: Why Reliable AI Needs Calibrated Intervention

TL;DR for operators Reliability is often treated as a binary control problem: approve or reject an agent action, preserve or replace a learned component. The evidence here points to a second question that can matter just as much: how strongly should the system intervene, where, and under what conditions? The clearest technical example comes from continual reinforcement learning. In a 400-million-step SlipperyAnt stress test, CPR recorded zero policy collapses across all 15 seeds under the paper’s main collapse criterion, while Adam and binary-reset baselines experienced collapses. Rather than fully replacing every selected component, CPR changes it by an amount tied to measured utility—preserving more useful learned state while refreshing low-utility state more aggressively. ...

August 11, 2026 · 8 min · Zelina
Cover image

Reasoning Tokens Are Compute, Not an Audit Trail

TL;DR for operators Giving an AI system more reasoning tokens can improve difficult answers because each generated token triggers another round of model computation and preserves intermediate information for the next step. For problems requiring a sequence of dependent operations, this can create additional computational depth rather than merely reveal reasoning that was already complete inside the model. ...

August 11, 2026 · 7 min · Zelina
Cover image

Common Is Not Defining: Testing Whether Language Models Understand Category Relations

TL;DR for operators A model-review team may need to decide whether a feature is essential to a category or merely common in the data. That distinction matters because a strong association can otherwise become an unsupported ontology rule, automated policy, risk classification, or product requirement. For six embedding-based transformer models, scores initially appeared to separate defining properties from properties that were only statistically common. Once researchers controlled for human-rated prevalence—how often each property occurs—most of that separation disappeared. The same raw score that seemed to reveal conceptual structure was largely explained by frequency. GPT-4 retained a substantially stronger distinction under the same control. ...

August 8, 2026 · 6 min · Zelina
Cover image

Right Answer, Wrong Evidence: A Deployment Gate for Grid-Diagnosis LLMs

TL;DR for operators A grid operator may see topology, live measurements, and an incident narrative all point to the same diagnosis. The decision is not only whether the answer is correct, but whether the model relied on evidence that the diagnostic task permits it to use. In the study, shortcut incident text produced a mean signed utility effect of +0.062 even though its preregistered engineering importance was zero. The model therefore became more accurate by using evidence that should not have determined the answer. Accuracy and a plausible explanation cannot reveal that divergence on their own. ...

August 7, 2026 · 7 min · Zelina