Cover image

Safety Without the Data Lake: Federating the Guard, Not the Traces

TL;DR for operators Suppose several business units or partner organizations run different agent workflows. Each has its own prompts, tools, communication patterns, and failure cases. They want a common safety layer, but centralizing those interaction traces would expose precisely the operational data they are trying to protect. The harder problem is that a guard trained elsewhere may not transfer well enough to solve this. In the reported experiments, an architecture-matched topology guard scores 0.512 AUROC when transferred off the shelf to Agent-SafetyBench, but 0.695 after in-domain retraining. Local adaptation helps, yet isolated local training is also weaker than collaborative training and becomes fragile when client labels are highly skewed. ...

September 27, 2026 · 7 min · Zelina
Cover image

Decision Rights, Not More Layers: What an Auditable Fraud Pipeline Actually Earns

TL;DR for operators A fraud classifier has already scored a transaction. The next operational choice is whether extra context—relationship patterns, anomaly signals, explanations, or an LLM investigator—should merely inform the case or be allowed to change the decision. In Rahil Sharma’s evaluation, the answer is component-specific.1 The bounded LLM investigator was correct on 39 of 60 deliberately balanced difficult cases, versus 43 of 60 for simply applying a 0.5 threshold to the classifier: 65.0% versus 71.7%. The agent changed eight classifier decisions. Two changes fixed mistakes; six replaced correct decisions with incorrect ones. ...

August 24, 2026 · 7 min · Zelina
Cover image

Curved Space, Straighter Retrieval: Why Graph RAG Needs Geometry

Curved Space, Straighter Retrieval: Why Graph RAG Needs Geometry Retrieval looks simple until the wrong thing keeps showing up. A company builds a graph model over products, papers, suppliers, users, or transactions. The model performs reasonably well inside familiar territory. Then the data shifts. New products appear. A new research domain enters the citation graph. A social platform changes user behavior. The model’s internal knowledge, frozen inside parameters, starts behaving like yesterday’s org chart: technically structured, operationally stale. ...

June 6, 2026 · 15 min · Zelina
Cover image

Regrets, Graphs, and the Price of Privacy: Federated Causal Discovery Grows Up

A hospital changes its treatment protocol. Another keeps the old one. A third removes an approval step that had quietly influenced several downstream decisions. Their datasets now disagree. The usual federated-learning instinct is to treat that disagreement as a problem: smooth it, average it, or design an aggregation rule robust enough to survive it. In causal discovery, however, some disagreements contain precisely the information the global model lacks. Removing a local dependency can expose a previously hidden causal pattern. A policy difference that looks like statistical inconvenience may function as an accidental experiment. ...

December 30, 2025 · 17 min · Zelina
Cover image

Benchmarking Without Borders: How GraphBench Rewrites the Rules of Graph Learning

Benchmarks Are Where Models Stop Being Inspirational Benchmarks are not glamorous. They are where models go after the demo video, after the conference slide, and after the sentence “this generalizes beautifully” has done its little dance in front of investors. Graph learning badly needs that room. For years, graph machine learning has been evaluated on comfortable territory: molecular graphs, citation networks, small academic datasets, and carefully packaged tasks that are useful but narrow. That helped the field grow. It also created a quiet distortion. A model could look impressive while never having to deal with a social network that changes over time, a circuit whose tiny structural error destroys correctness, a SAT instance where solver choice matters, or a weather graph where the planet is inconveniently spherical. ...

December 7, 2025 · 16 min · Zelina
Cover image

Tools of Habit: Why LLM Agents Benefit from a Little Inertia

Tools are where many agent demos quietly become invoices. A multi-step LLM agent may look intelligent because it reasons, acts, observes, and repeats. Under the hood, though, it often pays the model to decide every small next move: search here, load that node, look around, check valid actions, fill this argument, try again. Some of those decisions need judgement. Others are basically muscle memory wearing a lab coat. ...

November 20, 2025 · 14 min · Zelina
Cover image

Quantum Bridges: Crossing the Label Gap with ILQSSL and IPQSSL

TL;DR for operators Labels are expensive. That is the clean business problem behind this paper. In healthcare, credit review, fraud triage, and scientific classification, organisations often have many observations and too few trusted labels. Semi-supervised learning tries to stretch those scarce labels across the structure of the data rather than pretending every missing label is merely a procurement problem with a nicer dashboard. ...

August 9, 2025 · 15 min · Zelina
Cover image

Fraud, Trimmed and Tagged: How Dual-Granularity Prompts Sharpen LLMs for Graph Detection

TL;DR for operators Fraud teams already know the problem: the suspicious review, shop, seller, or account is rarely suspicious in isolation. The useful evidence is scattered across neighbours — same user, same product, same rating pattern, same time window, same commercial ecosystem. The less useful evidence is also scattered there. At scale, that second pile is larger. How inconvenient. ...

July 30, 2025 · 15 min · Zelina