Cover image

XAI, But Make It Scalable: Why Experts Should Stop Writing Rules

Churn is a wonderfully inconvenient business problem. Customers do not leave in one elegant, universal way. Some leave because price finally annoyed them. Some leave because support failed at exactly the wrong moment. Some leave because a monthly contract made exit frictionless. Some leave because they were already mentally gone and the invoice merely made it official. ...

December 23, 2025 · 15 min · Zelina
Cover image

When Agents Agree Too Much: Emergent Bias in Multi‑Agent AI Systems

When Agents Agree Too Much: Emergent Bias in Multi-Agent AI Systems Credit review is not supposed to work like a group chat. A bank cannot defend a biased lending workflow by saying, “each analyst looked fair on their own.” The decision process matters. Who sees whose opinion matters. Whether dissent survives matters. Whether the final answer comes from independent judgment or from a politely self-reinforcing committee definitely matters. ...

December 21, 2025 · 14 min · Zelina
Cover image

The Latent Truth: Why Prototype Explanations Need a Reality Check

The Latent Truth: Why Prototype Explanations Need a Reality Check Audit starts with a simple request: show me why. For prototype-based neural networks, that request has always had a pleasantly visual answer. The model points to a learned prototype from training data and says, in effect, “this part of the image looks like that part of an example I already know.” This is the interpretability sales pitch in its most charming form. No opaque wall of logits. No post-hoc heatmap pretending to be a confession. Just a case-based explanation: this resembles that. ...

November 22, 2025 · 15 min · Zelina
Cover image

Steering the Schemer: How Test-Time Alignment Tames Machiavellian Agents

A procurement agent does not need a villain moustache to become unpleasant. Give it a target, a reward function, and enough freedom, and it may discover that squeezing suppliers, hiding trade-offs, or exploiting procedural loopholes is not “unethical” in its world. It is just efficient. That is the point of the MACHIAVELLI benchmark, and also the reason the paper Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping is worth reading carefully.1 The paper is not selling a new moral soul for AI agents. Thankfully. We have enough vendors selling souls already. It proposes something more operationally useful: a runtime steering layer that adjusts an already-trained reinforcement learning agent’s action choices using attribute classifiers. ...

November 17, 2025 · 15 min · Zelina
Cover image

Graph Crimes of the Temporal Kind: How LoReTTA Quietly Breaks Time

A fraud model does not only learn from transactions. It learns from sequence. Who interacted with whom. When. How often. After what previous event. Before which next event. In temporal graph systems, the order is not metadata. It is the thing being modelled. That is why LoReTTA is an uncomfortable paper.1 It does not argue that Temporal Graph Neural Networks can be broken only by a powerful adversary with model access, expensive surrogate training, and a theatrical pile of fake edges. It argues something more operationally annoying: a continuous-time graph can be poisoned by removing influential interactions and replacing them with plausible ones. The resulting history still looks enough like history. The model quietly learns the wrong temporal structure. Very civilised, as crimes go. ...

November 16, 2025 · 16 min · Zelina
Cover image

Better Wrong Than Certain: How AI Learns to Know When It Doesn’t Know

A credit model approves the familiar applicant. A diagnostic model reads the common scan. A pricing model values the house in a neighbourhood it has seen a thousand times before. Everyone relaxes. The model is “confident”. Then a strange case arrives. The applicant has an unusual income pattern. The scan comes from an underrepresented patient group. The house sits outside the areas covered by historic transactions. The model still produces an answer, because that is what models are trained to do. Press button, receive number. Very efficient. Occasionally ridiculous. ...

November 10, 2025 · 14 min · Zelina
Cover image

FAITH in Numbers: Stress-Testing LLMs Against Financial Hallucinations

TL;DR for operators FAITH is useful because it changes the hallucination question from “Does the model sound right?” to “Can the model reconstruct a known financial number from the exact tables and surrounding text that justify it?”1 That sounds modest. It is not. In finance, modest is usually where the damage hides. ...

August 8, 2025 · 18 min · Zelina
Cover image

Curvature in the Jump: Geometrizing Financial Lévy Models

TL;DR for operators Jaehyung Choi’s paper does not offer a new trading strategy, volatility forecast, or backtest that makes the Sharpe ratio stand up and sing.1 Its contribution is more structural: it builds an information-geometric framework for Lévy processes, the family of stochastic processes often used when financial returns refuse to behave like polite Gaussian increments. ...

August 3, 2025 · 17 min · Zelina
Cover image

Signed, Sealed, Delivered: A Rough Path to Better Volatility Models

TL;DR for operators Options calibration has a familiar operational problem: the model that is fast enough to run every day is usually the model that assumes the market is behaving politely. The market, naturally, has other hobbies. This paper compares two ways of calibrating implied volatility surfaces. The first is the classical route: use model-specific analytical approximations for Heston and rough Bergomi. The second is the rough-path route: represent volatility as a linear functional of the truncated signature of a primary stochastic process.1 ...

August 3, 2025 · 15 min · Zelina
Cover image

OneShield Against the Storm: A Smarter Firewall for LLM Risks

TL;DR for operators Enterprise LLM safety is often discussed as if the main question is whether the model has been trained to “behave”. That is the comforting version of the story. It is also too small. IBM’s OneShield paper argues for a different operating model: treat safety as a separate, model-agnostic guardrail layer that sits around the LLM, runs multiple specialised detectors in parallel, and then applies explicit policy decisions through a separate policy manager.1 In plain business terms, OneShield is less like teaching the model good manners and more like installing a configurable safety-control plane around every AI interaction. Glamorous? Not especially. Operationally useful? Very much so. ...

July 30, 2025 · 18 min · Zelina