Cover image

A Refusal Is Only One Turn: PsychJail Tests Safety Under Adaptive Persuasion

TL;DR for operators A model that refuses a harmful request once has demonstrated one response under one conversational strategy. That is weaker evidence than it may appear. In the experiments behind PsychJail,1 successful attacks were strongly front-loaded: the median successful turn was 1 for all four tested victim models, while the mean successful turn ranged only from 1.20 to 1.37. Many failures therefore did not require a long campaign of conversational erosion. They appeared as soon as the attacker found a more effective way to frame the interaction. ...

September 21, 2026 · 7 min · Zelina
Cover image

A Refusal Is Not a Safety Test: Probe Harm After the Prompt Changes Form

TL;DR for operators A chatbot that refuses a plainly written harmful request has passed one test of its safety behavior, not the whole test. In Emoji-Based Jailbreaking of Large Language Models, Gopinadh and Hussain submit the same 50 emoji-augmented adversarial prompts to four locally deployed open-source models.1 They report successful jailbreak rates of 10% for Gemma 2 9B, 10% for Mistral 7B, 6% for Llama 3 8B, and 0% for Qwen 2 7B. ...

September 20, 2026 · 7 min · Zelina
Cover image

Safety Has a Memory: Why Multimodal Jailbreak Testing Must Follow the Conversation

TL;DR for operators Safety testing for a multimodal assistant should cover sequences of interactions, not only whether the system refuses one obviously prohibited prompt. In the tested setup, a staged three-turn attack reached a 91.50% attack success rate on LLaVA-7B and 77.31% on GPT-4o, above the three single-turn attack baselines reported for those models.1 The result does not establish universal failure rates, but it does show that a prompt-level pass can miss vulnerabilities that emerge after earlier turns establish conversational context. ...

September 20, 2026 · 7 min · Zelina
Cover image

A Clean Jailbreak Cluster Can Still Miss Unsafe Compliance

TL;DR for operators A safety monitor can become highly accurate at recognizing jailbreak-shaped prompts without becoming equally accurate at predicting unsafe model behavior. Delcon, Algaba, and Ginis demonstrate this gap across six Qwen and Llama instruction-tuned models.1 Their internal embeddings separate control and jailbreak prompts with balanced accuracy around 0.99–1.00. Yet in Qwen-2.5-7B, where refusal and compliance observations are comparatively balanced, the corresponding refusal-versus-compliance separation reaches only 0.677. ...

August 28, 2026 · 7 min · Zelina
Cover image

The Refusal Rate That Refuses to Reassure

TL;DR for operators The reassuring headline is that both evaluated frontier models rejected most automated jailbreak attempts. The operationally useful headline is that they still produced 1,620 and 702 panel-confirmed harmful completions, respectively, across every top-level harm category in the benchmark.1 The strongest adaptive attack succeeded on 11.51% of attempts against one model and 6.10% against the other. Static encodings and familiar jailbreak templates, by contrast, were almost entirely neutralised. ...

July 16, 2026 · 18 min · Zelina
Cover image

Mind the Slot: Jailbreak Prompts Have Weak Points, Not Just Bad Words

Security teams like to search for suspicious strings. That habit is understandable. Strings are visible. They can be logged, filtered, matched, scored, and proudly displayed in dashboards. A bad suffix at the end of a prompt looks like a bad suffix at the end of a prompt. Convenient. Almost too convenient. The problem is that prompts are not flat text boxes. They are transformed into token sequences, wrapped in chat templates, and passed through attention layers that do not treat every position equally. Some positions receive more influence over the model’s next-token behavior than others. Put adversarial tokens there, and the same amount of “badness” can travel farther. ...

June 6, 2026 · 19 min · Zelina
Cover image

Jailbreak Risk Needs a Stopwatch, Not Just a Scorecard

Jailbreak Risk Needs a Stopwatch, Not Just a Scorecard For many organizations, LLM safety is still treated like a checkpoint: run a benchmark, report an attack success rate, add a few guardrails, and move on. The resulting dashboard looks reassuringly official. It may even have decimals. Unfortunately, adversarial users do not attack dashboards. They attack systems. ...

May 30, 2026 · 17 min · Zelina
Cover image

Jailbreak ASR Is Wearing a Costume

The number looked safe. Then someone ran it twice. A familiar business problem: one vendor says its model resists jailbreaks. Another red-team report says a new attack reaches a spectacular Attack Success Rate. A compliance team sees a percentage, puts it into a risk register, and moves on. Unfortunately, that percentage may be doing more acting than measuring. ...

May 29, 2026 · 14 min · Zelina
Cover image

Red Queen Receipts: AI Security Testing Needs Logs, Not Vibes

Security testing is not a screenshot. A model gives a dangerous answer. Someone posts the transcript. A vendor says the model has been updated. A consultant turns the incident into a slide titled “AI risk is real.” Everyone nods gravely. Very mature. Very enterprise. The harder question is less theatrical: can the same vulnerability be tested again, under controlled conditions, with visible logs, a consistent evaluator, repeatable statistics, and enough human inspection to make the result defensible? ...

May 22, 2026 · 14 min · Zelina
Cover image

Context Is the New Attack Surface

A benchmark score is easy to quote. It is harder to know what broke. In Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models, Pavlos Ntais reports an 81.0% attack success rate against GPT-OSS-20B on a held-out 200-item test set.1 That number is attention-grabbing. It is also not the main lesson. ...

May 16, 2026 · 13 min · Zelina