When Not to Send Data to a Public LLM

How to decide when a business workflow should avoid public LLM endpoints, based on data sensitivity, contractual exposure, and safer design alternatives.

March 16, 2026 · 7 min · Michelle

How to Design Human Review for AI Systems

How to build a risk-tiered human review model so oversight is meaningful, efficient, and matched to business impact rather than added as a vague slogan.

March 16, 2026 · 6 min · Michelle

AI Access Control, Logging, and Retention Policies

How to design access controls, prompt/output logging, and retention rules for AI systems so governance remains practical, auditable, and proportional to risk.

March 16, 2026 · 7 min · Michelle

AI Vendor Risk Assessment and Procurement Checklist

How to evaluate AI vendors before rollout, using a practical checklist for data handling, governance, contract risk, security posture, and operational fit.

March 16, 2026 · 7 min · Michelle

AI Evaluation, Monitoring, and Incident Response for Production Systems

How to evaluate, monitor, and respond to failures in production AI systems so quality, safety, and governance remain active after launch.

March 16, 2026 · 6 min · Michelle
Cover image

When the Test Becomes a Signal: Rethinking AI Agent Evaluation

TL;DR for operators A tool-using agent does not experience an evaluation as an abstract benchmark. It sees prompts, tool wrappers, permissions, response timing, filesystem artifacts, network behavior, logging infrastructure, and other parts of the environment. If those signals differ from production, a sufficiently adaptive agent may be able to infer when it is being tested and behave differently. ...

September 7, 2026 · 8 min · Zelina
Cover image

Important, but Not Direct: When Time-Series Attribution Misstates Model Dependencies

TL;DR for operators A forecasting dashboard can correctly report that an earlier observation influenced a prediction and still give the wrong impression about how that influence enters the model. Amadeo Tunyi’s paper, The Failures of Marginal Influence-Based Attribution Methods for Global Time Series Explanations, argues that familiar scalar attribution methods cannot in general recover the model’s direct temporal dependency structure.1 Marginal methods can assign importance to an earlier variable whose influence is entirely mediated through a later, autocorrelated observation. Gradient methods can report sensitivity that exists only outside the support of the data the model actually sees. ...

August 19, 2026 · 8 min · Zelina
Cover image

Agree Once, Remember Later: The Commit Boundary in Personal Agents

TL;DR for operators A personal assistant can hear a confident user claim, store it as a preference or rule, and rely on it during a later task after the original conversation is gone. The safety problem is therefore not only the agreeable reply. It is the write that lets the claim survive. ...

July 24, 2026 · 8 min · Zelina
Cover image

Frame Before You Aim: Why AI Needs the Right Reference Point

Business AI has acquired a slightly dangerous reflex: when a system underperforms, reach for a stronger model, a faster pipeline, or a more elaborate scoring function. Very enterprise. Very expensive. Occasionally useful. The more interesting failure mode is quieter. A system may have enough intelligence, enough data, and enough compute, yet still be solving the wrong version of the problem because it inherited the wrong reference frame. It reads a wearable signal as if it were clinical instrumentation. It schedules network traffic as if packets only matter after they announce themselves. It ranks alternatives as if the best and worst items in the current dataset were the same thing as business aspiration and business refusal. ...

June 14, 2026 · 15 min · Zelina
Cover image

Jailbreak ASR Is Wearing a Costume

The number looked safe. Then someone ran it twice. A familiar business problem: one vendor says its model resists jailbreaks. Another red-team report says a new attack reaches a spectacular Attack Success Rate. A compliance team sees a percentage, puts it into a risk register, and moves on. Unfortunately, that percentage may be doing more acting than measuring. ...

May 29, 2026 · 14 min · Zelina