Cover image

A Refusal Is Not a Safety Test: Probe Harm After the Prompt Changes Form

TL;DR for operators A chatbot that refuses a plainly written harmful request has passed one test of its safety behavior, not the whole test. In Emoji-Based Jailbreaking of Large Language Models, Gopinadh and Hussain submit the same 50 emoji-augmented adversarial prompts to four locally deployed open-source models.1 They report successful jailbreak rates of 10% for Gemma 2 9B, 10% for Mistral 7B, 6% for Llama 3 8B, and 0% for Qwen 2 7B. ...

September 20, 2026 · 7 min · Zelina