Cover image

Following Instructions Is Not the Same as Knowing More

TL;DR for operators A multimodal model can become much better at obeying instructions without becoming much better at the underlying tasks those instructions govern. In the experiments examined here, one 8B vision-language model gains 10.58 percentage points on a targeted instruction-following benchmark and another gains 22.91 points. Yet their average results across broader STEM, VQA, OCR, and document-understanding tests move by only +0.33 and -0.21 points. ...

October 1, 2026 · 7 min · Zelina
Cover image

Give the Generator Less to Leak: KFS-RAG Moves Privacy to the Retrieval Boundary

TL;DR for operators An enterprise assistant may retrieve a long internal document because one sentence is relevant to a user question. If the entire passage is then handed to the generator, the model receives much more information than it needs—and a successful prompt injection has more material available to disclose. KFS-RAG addresses that exposure point after retrieval but before generation. Instead of forwarding raw passages, it identifies query-relevant evidence and converts it into a compact set of facts. In the paper’s open-domain QA evaluation, this reduced untargeted chunk recovery from 70.8% with Vanilla RAG to 21.03%, while answer BLEU-1 and ROUGE-L remained close to the raw-context baseline. ...

September 22, 2026 · 8 min · Zelina
Cover image

Don’t Rank the Guardrails: Map Prompt-Injection Defenses to the Stack

TL;DR for operators Prompt-injection defense should not be procured as a leaderboard winner. A systematic review of 88 studies finds defenses distributed across several parts of the LLM system, from training and prompt handling to document boundaries, tool execution, output filtering, and continuous testing.1 Fifty-six of those 88 approaches, or 63.63%, are model-agnostic: they can operate around different models without changing model weights or architecture. That matters for teams building on proprietary APIs. ...

September 21, 2026 · 8 min · Zelina
Cover image

The Planner Trusted the Wrong State: A New Security Boundary for Embodied Agents

TL;DR for operators An embodied agent can receive the correct user instruction and still plan toward the wrong objective if the internal description of its environment has been manipulated. Liu et al. test this failure mode by altering planner-visible state semantics rather than changing the instruction, model, planner, executor, or environment itself.1 ...

September 6, 2026 · 7 min · Zelina
Cover image

The Model Remembers What the Firewall Forgets

TL;DR for operators The current enterprise LLM mistake is to treat safety as a chat-interface problem. Put a filter in front, add a stern system prompt, run a cheerful demo, and hope the model has suddenly acquired a moral philosophy and a compliance department. Charming. Also insufficient. Two recent papers make a stronger and more operationally useful point. The first evaluates open-source LLMs under prompt-injection and jailbreak attacks, then compares lightweight inference-time defenses such as input filtering, system-prompt hardening, vector detection, voting, and self-examination.1 The second studies whether fine-tuned language models memorize sensitive personally identifiable information that appears only in inputs, not in the desired outputs, and benchmarks mitigation methods such as differential privacy, UnDial, regularization, and DPO.2 ...

July 6, 2026 · 18 min · Zelina
Cover image

The Tool Response Is Not Your Boss

TL;DR for operators The paper’s useful message is not “LLM agents are unsafe,” which is too vague to help anyone do anything before lunch. The useful message is narrower and more operational: agents become vulnerable when untrusted content from SaaS integrations is read into the agent context and then treated as authority for a later action. ...

July 1, 2026 · 19 min · Zelina
Cover image

Feedback Is the New Attack Surface

TL;DR for operators AI agents are not only vulnerable because someone can hide a bad instruction in an email, document, web page, Slack message, or tool output. They are vulnerable because attackers can now automate the search for bad instructions that work. That changes the security problem. A one-off prompt injection is annoying. An automated attack loop is strategic. It generates candidate injections, observes the agent’s response, scores partial progress, keeps the promising branches, and tries again. Very entrepreneurial, in the worst possible way. ...

June 23, 2026 · 21 min · Zelina
Cover image

The Jailbreak Wasn’t Written. It Was Bred.

TL;DR for operators The paper introduces GAS-Leak-LLM, a black-box method that uses a genetic algorithm to evolve adversarial suffixes: small text sequences appended to harmful prompts to increase the chance that a model produces unsafe content.1 The important part is not that another jailbreak exists. We have enough of those. The important part is that jailbreak discovery is framed as a repeatable optimization loop using only model queries. ...

June 23, 2026 · 15 min · Zelina
Cover image

Stop Signs Are Not Steering Wheels: TRIAD and the Case for Repairable Agent Guardrails

TL;DR for operators Most agent guardrails behave like stop signs. They inspect a proposed action, decide whether it looks safe, and then allow or block execution. This is neat, legible, and often operationally clumsy. Real agent failures are not always cleanly harmful from the first word. A useful business request can be contaminated by a prompt injection, a malicious tool response, or an unsafe intermediate plan. Blocking the whole task may reduce risk, but it also throws away the legitimate work. Excellent safety theatre, less excellent operations. ...

June 19, 2026 · 20 min · Zelina
Cover image

Mind the Slot: Jailbreak Prompts Have Weak Points, Not Just Bad Words

Security teams like to search for suspicious strings. That habit is understandable. Strings are visible. They can be logged, filtered, matched, scored, and proudly displayed in dashboards. A bad suffix at the end of a prompt looks like a bad suffix at the end of a prompt. Convenient. Almost too convenient. The problem is that prompts are not flat text boxes. They are transformed into token sequences, wrapped in chat templates, and passed through attention layers that do not treat every position equally. Some positions receive more influence over the model’s next-token behavior than others. Put adversarial tokens there, and the same amount of “badness” can travel farther. ...

June 6, 2026 · 19 min · Zelina