Cover image

Agents, Not Tasks: Rethinking Business Processes in the Age of AI

A mechanism-first reading of how agentic AI changes business process design from fixed task flows to goal-driven coordination.

July 30, 2025 · 19 min · Zelina
Cover image

Beyond Words: Teaching AI to See and Fix Charts with ChartM3

ChartM3 shows why chart editing needs more than language: models must connect what users point at to the code that actually controls the visual object.

July 30, 2025 · 18 min · Zelina
Cover image

Circuits of Understanding: A Formal Path to Transformer Interpretability

A clearer operator-focused reading of how circuit-level interpretability turns transformer explanations from attractive diagrams into testable engineering claims.

July 30, 2025 · 14 min · Zelina
Cover image

Fraud, Trimmed and Tagged: How Dual-Granularity Prompts Sharpen LLMs for Graph Detection

A practical reading of DGP, a graph-enhanced LLM framework that improves fraud detection by preserving target-node detail while compressing noisy neighbourhood context.

July 30, 2025 · 15 min · Zelina
Cover image

OneShield Against the Storm: A Smarter Firewall for LLM Risks

IBM’s OneShield paper shows why enterprise LLM safety is becoming a configurable control plane, not merely a better moderation classifier.

July 30, 2025 · 18 min · Zelina
Cover image

The User Is Present: Why Smart Agents Still Don't Get You

A close reading of UserBench and what it reveals about the gap between tool-using agents and genuinely user-centred assistants.

July 30, 2025 · 17 min · Zelina
Cover image

Too Nice to Be True? The Reliability Trade-off in Warm Language Models

Warmth in AI assistants may improve the user experience, but this paper shows it can also make models less reliable and more sycophantic.

July 30, 2025 · 17 min · Zelina
Cover image

Don't Trust. Verify: Fighting Financial Hallucinations with FRED

A practical look at FRED, a finance-focused framework for detecting and editing hallucinations in retrieval-grounded LLM outputs.

July 29, 2025 · 17 min · Zelina
Cover image

From Molecule to Mock Human: Why Programmable Virtual Humans Could Rewrite Drug Discovery

A practical reading of programmable virtual humans as a proposed bridge between molecular AI, digital twins, and physiology-first drug discovery.

July 29, 2025 · 18 min · Zelina
Cover image

Mirage Agents: When LLMs Act on Illusions

MIRAGE-Bench reframes agent hallucination as context-unfaithful action and gives operators a sharper way to test LLM agents before they touch real systems.

July 29, 2025 · 19 min · Zelina