Cover image

A Clean Jailbreak Cluster Can Still Miss Unsafe Compliance

TL;DR for operators A safety monitor can become highly accurate at recognizing jailbreak-shaped prompts without becoming equally accurate at predicting unsafe model behavior. Delcon, Algaba, and Ginis demonstrate this gap across six Qwen and Llama instruction-tuned models.1 Their internal embeddings separate control and jailbreak prompts with balanced accuracy around 0.99–1.00. Yet in Qwen-2.5-7B, where refusal and compliance observations are comparatively balanced, the corresponding refusal-versus-compliance separation reaches only 0.677. ...

August 28, 2026 · 7 min · Zelina
Cover image

Prompted to Death: When Words Become a Denial-of-Service

A customer asks an AI assistant a question. The assistant begins answering, continues answering, wanders into repetition, and eventually reaches the maximum output limit. Nobody stole a password. No prohibited content appeared. The model may even have remained grammatically competent throughout the ordeal. It simply consumed far more computation than the request deserved. ...

January 4, 2026 · 19 min · Zelina