Cover image

Restore the Image, Preserve the Poisoning Signal

TL;DR for operators Publishing images that remain useful to people while making them poor training material for unauthorized models creates an awkward engineering constraint: stronger interference often comes with more visible image degradation. The paper Leveraging Imperfect Restoration for Data Availability Attack1 shows that those two objectives need not move together. Its proposed method deliberately restores some of the visual damage created by an existing convolution-based poisoning technique while leaving enough class-specific structure to continue disrupting learning. ...

September 29, 2026 · 7 min · Zelina
Cover image

Safety Without the Data Lake: Federating the Guard, Not the Traces

TL;DR for operators Suppose several business units or partner organizations run different agent workflows. Each has its own prompts, tools, communication patterns, and failure cases. They want a common safety layer, but centralizing those interaction traces would expose precisely the operational data they are trying to protect. The harder problem is that a guard trained elsewhere may not transfer well enough to solve this. In the reported experiments, an architecture-matched topology guard scores 0.512 AUROC when transferred off the shelf to Agent-SafetyBench, but 0.695 after in-domain retraining. Local adaptation helps, yet isolated local training is also weaker than collaborative training and becomes fragile when client labels are highly skewed. ...

September 27, 2026 · 7 min · Zelina
Cover image

Privacy Starts Before the First Gradient

TL;DR for operators Federated fine-tuning keeps raw examples on the client, but that does not mean the client begins from a neutral model state. A malicious coordinating server can send an adapter deliberately structured so that private examples produce recoverable traces during training. Privacy risk can therefore enter through what the client downloads, not only through what it later uploads. ...

August 28, 2026 · 6 min · Zelina
Cover image

Black Box Is Too Blunt: What AI Interfaces Reveal to Attackers

TL;DR for operators What exactly should a deployed AI system reveal to users, vendors, insiders, or connected applications? Hiding model architecture and weights does not make every interface equally opaque. Mahbub and colleagues separate access into six operational categories—None, Metadata, Decision-Only, Score/Rank, Embedding, and White-Box—because each exposes a different signal to an adversary.1 A binary decision permits probing, while numerical confidence or similarity values provide directional feedback; internal feature representations expose still richer information. The paper’s synthesis suggests that richer signals generally reduce attacker uncertainty and query burden while enabling additional risks such as model extraction and biometric template inversion. ...

August 15, 2026 · 7 min · Zelina
Cover image

Feedback Is the New Attack Surface

TL;DR for operators AI agents are not only vulnerable because someone can hide a bad instruction in an email, document, web page, Slack message, or tool output. They are vulnerable because attackers can now automate the search for bad instructions that work. That changes the security problem. A one-off prompt injection is annoying. An automated attack loop is strategic. It generates candidate injections, observes the agent’s response, scores partial progress, keeps the promising branches, and tries again. Very entrepreneurial, in the worst possible way. ...

June 23, 2026 · 21 min · Zelina
Cover image

Jailbreak ASR Is Wearing a Costume

The number looked safe. Then someone ran it twice. A familiar business problem: one vendor says its model resists jailbreaks. Another red-team report says a new attack reaches a spectacular Attack Success Rate. A compliance team sees a percentage, puts it into a risk register, and moves on. Unfortunately, that percentage may be doing more acting than measuring. ...

May 29, 2026 · 14 min · Zelina
Cover image

Red Queen Receipts: AI Security Testing Needs Logs, Not Vibes

Security testing is not a screenshot. A model gives a dangerous answer. Someone posts the transcript. A vendor says the model has been updated. A consultant turns the incident into a slide titled “AI risk is real.” Everyone nods gravely. Very mature. Very enterprise. The harder question is less theatrical: can the same vulnerability be tested again, under controlled conditions, with visible logs, a consistent evaluator, repeatable statistics, and enough human inspection to make the result defensible? ...

May 22, 2026 · 14 min · Zelina
Cover image

Receipts, Please: RAG’s New Evidence Stack

Opening — Why this matters now The original business pitch for retrieval-augmented generation was wonderfully simple: connect the model to your documents, ask questions, get grounded answers. No need to retrain the model. No need to wait for the next foundation-model release. Just give the chatbot some files and let productivity bloom. ...

May 7, 2026 · 17 min · Zelina
Cover image

Phantasia and the Illusion of Safety: When AI Lies Without Looking Wrong

Safety checks usually look for the model doing something strange. That sounds reasonable. A compromised model should produce a strange phrase, repeat a suspicious payload, ignore the image, or behave in a way that feels obviously detached from the input. This is the comforting version of AI security: attackers leave fingerprints, defenders look for fingerprints, and everyone goes home after filling out a procurement checklist. ...

April 12, 2026 · 17 min · Zelina
Cover image

Protocol Over Prompts: Why ANX Rewrites the Rules of AI Agent Interaction

Forms are boring until an AI agent has to fill one. Then the boring form becomes a surprisingly expensive machine. The agent reads the page, interprets the fields, finds the dropdowns, waits for the browser, loads dynamic options, decides what to click, serializes actions, and tries not to leak whatever the user typed into the wrong place. This is not intelligence in the glamorous sense. It is office work wearing a robotic costume. ...

April 7, 2026 · 18 min · Zelina