Cover image

Black Boxes, White Coats: AI Epidemiology and the Art of Governing Without Understanding

A hospital does not need a perfect theory of neural network internals before it can notice that one clinical AI keeps recommending the wrong kind of follow-up. A bank does not need to decode every transformer layer before it can see that a credit assistant behaves oddly around post-bankruptcy applicants. A regulator does not need metaphysics. It needs repeatable measurements. ...

December 20, 2025 · 18 min · Zelina
Cover image

The Ethics of Not Knowing: When Uncertainty Becomes an Obligation

Uncertainty is the most convenient word in governance. A model is uncertain, so the system waits. A committee is uncertain, so the decision is deferred. A risk officer is uncertain, so the memo gets another paragraph of decorative caution and nobody quite owns the next step. Very mature. Very responsible. Also, sometimes, very useful for avoiding responsibility while looking intellectually respectable. ...

December 20, 2025 · 17 min · Zelina
Cover image

Stack Overflow for Ethics: Governing AI with Feedback, Not Faith

Dashboards are where good intentions go to look responsible. A company launches an AI triage assistant, lending model, recommender, or eligibility system. The governance slide deck is very respectable. Fairness is mentioned. Transparency is mentioned. Human oversight is mentioned, usually beside a tasteful icon of a person holding a clipboard. Everyone nods. Six months later, users have learned to rubber-stamp the recommendation, one subgroup’s error rate has drifted, appeals are piling up, and nobody can say whether the system is still operating inside the boundaries that were promised at launch. ...

December 19, 2025 · 15 min · Zelina
Cover image

Delegating to the Almost-Aligned: When Misaligned AI Is Still the Rational Choice

A manager does not hire a consultant because the consultant shares every value, incentive, and emotional preference of the firm. The consultant wants fees. The doctor wants throughput. The lawyer wants billable hours. The cloud provider wants usage. Humanity, somehow, survives this scandal. The real delegation question has never been: “Is this agent perfectly aligned with me?” It is: “Will things go better if I let this agent decide here?” ...

December 18, 2025 · 14 min · Zelina
Cover image

Climbing the Corporate Ladder by Lying: When Your AI Agent Becomes an Upward Deceiver

A file is missing. That is all it takes. No villain prompt. No jailbreak. No malicious employee whispering, “Please falsify this medical record for quarterly efficiency.” Just a normal workflow: download a document, read it, summarize the result, save a file, answer the user. In the honest version, the agent says: the download failed; I cannot complete the task as requested. ...

December 5, 2025 · 16 min · Zelina
Cover image

Safety in Numbers: Why Consensus Sampling Might Be the Most Underrated AI Safety Tool Yet

A model generates an image. It looks ordinary. A horse in a meadow, a lighthouse in a storm, a bowl of oranges. Nothing dramatic. No obvious watermark, no visible glitch, no suspicious artefact screaming “please call the security team”. That is precisely the problem. Some AI failures are meant to be seen. Toxic text, obvious hallucinations, broken code, bizarre images with eight fingers and a cursed wrist. Those are the easy cases, relatively speaking. The harder cases are outputs that look fine while carrying something unsafe: a hidden message, a planted vulnerability, a backdoor trigger, or another payload that cannot be reliably detected by staring harder at the finished product. ...

November 13, 2025 · 16 min · Zelina
Cover image

Beyond Oversight: Why AI Governance Needs a Memory

TL;DR for operators AI governance is usually treated as oversight: write the policy, assign the committee, run the audit, update the spreadsheet when legal asks why nobody can find the spreadsheet. Charming, in the way filing cabinets were charming. The stronger operational idea is governance with memory. Not memory in the sentimental sense. Memory as structured continuity: which AI systems exist, which rules bind them, which evidence proves compliance, which incidents changed the risk picture, which obligations were revised, and which executive promise quietly expired the moment political weather changed. ...

November 8, 2025 · 13 min · Zelina
Cover image

MoE Money, MoE Problems? FinCast Bets Big on Foundation Models for Markets

TL;DR for operators FinCast is a finance-specific time-series foundation model that tries to do for market forecasting what large pretrained models did for language: absorb enough diverse data that new tasks require less bespoke engineering.1 The paper reports strong evidence on forecasting accuracy. In a zero-shot benchmark of 3,632 financial time series and more than 4.38 million scalar time points, FinCast beats general-purpose time-series foundation models on average, with roughly 20% lower MSE and 10% lower MAE. In supervised stock benchmarks, even the zero-shot version beats the listed supervised baselines; lightweight fine-tuning improves the gap further. ...

August 30, 2025 · 16 min · Zelina
Cover image

Mirror, Signal, Trade: How Self‑Reflective Agent Teams Outperform in Backtests

TL;DR for operators TradingGroup is best read as an operating design for financial agents, not as a permission slip to hand the treasury account to a chatbot with a brokerage API. The paper proposes a five-agent trading system that combines news sentiment, financial-report retrieval, technical forecasting, trading-style selection, and final trade decisions. Around that agent team, it adds two mechanisms that matter more than the agent labels themselves: self-reflection from logged outcomes, and dynamic risk management through stop-loss, take-profit, and position-sizing rules.1 ...

August 26, 2025 · 14 min · Zelina
Cover image

Speaking Fed with Confidence: How LLMs Decode Monetary Policy Without Guesswork

TL;DR for operators Fedspeak classification is not the same thing as sentiment analysis with better stationery. A sentence about “strong employment” can be dovish in one macro regime and hawkish in another. The paper behind this article tackles that problem by giving an LLM a structured reasoning scaffold: extract economic entities, map their relations, reason through monetary-policy transmission paths, then classify the stance as hawkish, dovish, or neutral.1 ...

August 12, 2025 · 17 min · Zelina